Arbitrary Style Transfer Method and Device Incorporating an Interactive Attention Mechanism
By combining the interactive attention mechanism of Transformer encoder and reversible neural network, the problem of content deviation and insufficient detail expression in existing image style transfer is solved, and a more natural stylized image generation is achieved, which is suitable for any style image transfer.
Patent Information
- Application Number
- CN202410397562.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-03
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-04-03
AI Technical Summary
The existing image style transfer methods have shortcomings in processing high-level semantic information and local details, resulting in content deviation and artifact problems, especially when faced with stylized images with irregular patterns.
The style transfer method based on the interactive attention mechanism is adopted, and the global features of the content image and the style image are extracted using the Transformer encoder, detailed features are extracted in combination with the reversible neural network, and feature fusion is performed through the interactive attention mechanism of the channel and space, adaptive interpolation is performed using the space-aware interpolation module to finally generate a stylized image.
It effectively improves the global and detailed information retention of image style transfer, and the generated images are more natural and beautiful, can adapt to any style images, reduce content deviation and artifacts, and improve the generalization ability of the model.
Smart Images

Figure CN118052706B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image style transfer, and particularly to an arbitrary style transfer method and device incorporating an interactive attention mechanism. Background Art
[0002] Image style conversion is an interesting and practical technology with wide applications in art, entertainment, design, and marketing. The purpose of image style transfer is to create an image with the structure of the content image and the style pattern of the style image while given a content image and a style image. Traditional image style transfer is mainly based on non-parametric algorithms, which resample the style image to maintain a realistic natural effect. However, this method has obvious limitations, that is, it can only focus on the low-level features of the image and cannot effectively process high-level semantic information, and these methods usually cannot decouple the content from the style.
[0003] Since Gatys' pioneering work, research on neural network style transfer has been carried out in multiple fields. The research mainly focuses on artistic style conversion, fidelity conversion, and video conversion. Current style transfer methods can be divided into three categories: single style transfer, multiple style transfer, and arbitrary style transfer. In particular, trained single style transfer methods only generate one style. For example, Johnson considered using a feedforward network to directly generate rendered images and achieve real-time style transfer. Wu et al. introduced a direction-aware loss function to produce more accurate and detailed style transfer results. Ulyanov et al. proposed a neural network method that can generate high-quality texture images. To enable a single network to achieve more style transfers, multiple style transfer models have been proposed. The classic Dumoulin et al. encoded the style as a low-dimensional vector representation, enabling a single network to achieve multiple style transfers. Cheng et al. developed an explicit representation for neural image style transfer. Zhang et al. encoded the content image and different artistic styles into a shared latent space to generate multiple styles. Since most multiple style transfer methods represent a style through learnable small networks, the types of styles that can be transferred are limited by the number of learnable small networks. To meet arbitrary style transfer, arbitrary style transfer models have emerged. Some scholars' pioneering work proposed a simple and effective method AdaIN, which transfers the global mean and variance of the style image to the content image in the feature space to support arbitrary input style images. Li et al. proposed a general image style transfer method based on feature transformation to achieve arbitrary style transfer. Cheng et al. designed a fast block-based style transfer method for arbitrary style images. Sheng et al. designed a multi-scale zero style transfer method based on feature decoration. Jing et al. combined the arbitrary style transfer method with dynamic instance normalization. However, these methods ignore local details, and their local stylization performance is greatly reduced.
[0004] To address the above problems, people have proposed local block stylization methods based on the attention mechanism. For example, SANet first uses the self-attention mechanism in artistic style transfer to achieve semantic mapping. Attention is used to calculate regions with similar content and style to achieve a mage with higher details. MANet improves the input of attention so that content and style features can be better decoupled. In addition, a new disentanglement loss function is proposed to enable the network to extract the main style patterns and precise content structures. AdaAttN adopts a fusion strategy that combines the attention mechanism and a method based on global feature statistics. In addition, to maintain the structure, a new reconstruction loss is proposed. PAMA uses attention operations and spatially aware interpolation to align the content and style spaces. In addition, StyTr2 introduces Transformer and content-aware positional encoding. However, when they face style images with irregular patterns, special repetitive artifacts will appear, resulting in abnormal stylization results. On this basis, Master elaborates on the problems of the attention mechanism in style transfer in detail and uses learnable residual connections to solve the problem of abnormal stylization results generated by the model when facing style images with irregular patterns. Some methods such as IEcontraAST, CAST, and AesUST explore contrastive learning and adversarial learning, but they still have abnormal patterns and fail to express detailed textures.
[0005] Deficiencies of the prior art:
[0006] 1. The encoder is responsible for encoding the input image into a feature representation in style transfer. However, the encoders used in previous studies have deficiencies in the ability to extract global and local information.
[0007] 2. Currently, attention-based style transfer networks are prone to content bias.
[0008] Since the attention mechanism can generate detailed images, it is often used as a style transfer technique. As a pioneering method using the attention mechanism, SANet calculates the attention map from style features and content features, and then uses the attention map to adjust the style features to obtain weighted style features. Finally, SANet fuses them into the content features to obtain stylized content features. However, the distribution of features is ignored when fusing the weighted style features and content features. Therefore, the generated image is easily affected by strong styles and exhibits content bias. Summary of the Invention
[0009] Aiming at the deficiencies of the existing technology, the present invention proposes an arbitrary style transfer method based on an interactive attention mechanism. The style transfer method uses the Transformer encoder in the joint feature encoder to extract the global features of the content image and the style image, and uses a reversible neural network to extract the detailed features of the content image and the style image. The global and detailed features of the content image and the style image are respectively sent into the channel and spatial interactive attention for fusion to obtain the global stylized features and the detailed stylized features. Then, a spatial-aware interpolation module is used for adaptive interpolation fusion, and finally, the stylized image is decoded. Specifically, it includes:
[0010] Step 1: Prepare the dataset, including the MS-COCO dataset as the content image and the WiKiArt dataset as the style image;
[0011] Step 2: Preprocess the dataset respectively. First, scale it to a size of 512X512, and then randomly crop it to a size of 256X256;
[0012] Step 3: Construct and initialize the style transfer network. The style transfer network includes a joint feature encoder, a style conversion module, a spatial-aware interpolation module, a decoder, and a discriminator. Among them,
[0013] The joint feature encoder of the style transfer network is composed of two parallel independent classification backbone networks. The independent classification backbone network includes a Transformer encoder and a reversible neural network. The branch where the Transformer encoder is located is used as the global branch, and the branch where the reversible neural network is located is used as the detailed branch. The style conversion module includes two identical first branches and second branches. The first branch is used for the fusion of global style and content features, and the second branch is used for the fusion of detailed style and content features. Both the first branch and the second branch include a channel-spatial attention module and a spatial-channel attention module. The spatial-aware interpolation module is used for fusing the stylized global features and the stylized detailed features. The decoder decodes the fused features into a stylized image, and the discriminator is used to discriminate the authenticity of the generated stylized image and the style image;
[0014] Step 4: Input the training data processed in Step 2 into the style transfer network constructed in Step 3 to train the network. Specifically, it includes:
[0015] Step 41: Input the content image I C and the style image I S in the training set into the joint feature encoder respectively to extract feature information. The content image I C obtains the global content feature T Cand the detailed content feature D C wherein the style image I S respectively obtains the global style feature T S and the detailed style feature D S after passing through the Transformer encoder and the reversible neural network in the joint feature encoder;
[0016] Step 42: Input the feature map extracted in Step 41 into the style conversion module. Specifically, input the global content feature T C and the global style feature T S into the first branch of the style conversion module to obtain the global stylized feature T CS , input the detailed content feature D C and the detailed style feature D S into the second branch of the style conversion module to obtain the detailed stylized feature D CS ;
[0017] Step 43: Input the global stylized feature T CS and the detailed stylized feature D CS into the spatial-aware interpolation module for fusion to obtain the fused feature F CS ;
[0018] Step 44: Use the decoder to decode the fused feature F CS to obtain the stylized image I CS ;
[0019] Step 5: Calculate the total loss of the style transfer network, including at least matrix matching loss, perceptual loss, color consistency loss, and adversarial loss. Specifically:
[0020] Step 51: Use the pre-trained VGG network to extract the features from layer 1 to layer 5, calculate the matrix matching loss and perceptual loss using the features from layer 1 to layer 5, and calculate the self-similarity loss and rEMD loss using the features from layer 3 to layer 5;
[0021] Step 52: Calculate the color consistency loss between the stylized image I CS and the style image I S ;
[0022] Step 53: Input the stylized image I CS and the style image I S into the discriminator to determine whether it is a style image, so as to calculate the adversarial loss;
[0023] Step 6: Steps 4 and 5 are sequentially passed through a set total number of training times. After each fixed number of training times, the model weights are saved. Then, the test set is passed into the trained image migration network for testing, and it is calculated whether the current test metric of the style migration network is the highest. If so, the training ends; if not, the weight parameters of the model loss are adjusted and training is restarted.
[0024] According to a preferred implementation, the processing process of the first branch in step 42 includes:
[0025] Step 421: First, the global content feature and the global style feature T S are input into the channel spatial attention module for processing. Three 1X1 convolutions are used to adjust the input features. One convolution processes the global content feature T C to obtain the feature map Q, and the other two convolutions process the global style feature T S to obtain the feature maps K and V. Then, the feature map Q is multiplied by the feature map K to obtain the semantic relationship map between the content and the style. Then, the Softmax activation function is used to map it, and then it is multiplied by the feature map V to obtain the weighted style feature M;
[0026] Step 422: Use a 3X3 convolution to process the global content feature T C to obtain the feature map T * C . The style feature M and the feature map T * C are input into the channel spatial interaction module for fusion. First, channel attention is used to process the feature map T * C to obtain the channel attention coefficient CA. Then, spatial attention is used to process the style feature M to obtain the spatial attention coefficient SA. Subsequently, the channel attention coefficient CA is multiplied by the style feature M to dynamically adjust its features in the channel dimension. The spatial attention coefficient SA is multiplied by the feature map T * C to dynamically adjust its features in the spatial dimension. Finally, the adjusted style feature M and the adjusted T * C are added to obtain the content feature E CS ;
[0027] Step 423: Then, the content feature E CS and the global style feature T S are input into the spatial channel attention module, and the weighted style feature M2 is obtained in the same way as in step 421;
[0028] Step 424: Use a 3X3 convolution to process the content feature ECS Process to obtain the feature map T * CS , and combine the style feature M2 with the feature map T * CS Fuse them through a spatial-channel interaction module. First, use channel attention to process the style feature M2 to obtain the channel attention coefficient CA, and then use spatial attention to process the feature map T * CS to obtain the spatial attention coefficient SA. Subsequently, multiply the channel attention coefficient CA by the feature map T * CS to dynamically adjust its features in the channel dimension, multiply the spatial attention coefficient SA by the style feature M2 to dynamically adjust its features in the spatial dimension, and finally add the adjusted style feature M2 and the feature map T * CS to obtain the global stylized feature T CS .
[0029] According to a preferred embodiment, the specific operation of the spatial perception interpolation module in step 43 is as follows:
[0030] Step 431: First, for the input global stylized feature T CS and the detailed stylized feature D CS , splice them on the channel, then use a convolution with a convolution kernel of 1X1 for channel splicing, and then use convolutions with convolution kernels of 1X1, 3X3, and 5X5 respectively. After convolution, use the Sigmoid activation function to map them to obtain three mapping values, and finally add the three obtained mapping values and divide by three to obtain the weight α;
[0031] Step 432: For the input global stylized feature T CS and the detailed stylized feature D CS , adjust them respectively through two convolutions with a convolution kernel of 1X1, and then multiply the weight α obtained in step 431 by the global stylized feature T CS to adjust the features, use the weight 1-α to multiply the detailed stylized feature D CS to adjust it, and finally add them to obtain the final fused feature F CS .
[0032] An arbitrary style transfer device integrating an interactive attention mechanism, the arbitrary style transfer device includes an image preprocessing module, a joint feature encoder, a style conversion module, a spatial perception interpolation module, a decoder, and a discriminator, where
[0033] The image preprocessing module is used to uniformly scale the style images and content images in the dataset and crop them to the same size;
[0034] The joint feature encoder includes a Transformer encoder and a reversible neural network. The Transformer encoder is used to extract the global features of the style image and the content image, and the reversible neural network is used to extract the detailed features of the style image and the content image;
[0035] The style conversion module is used to fuse the extracted global features and detailed features. Specifically, the first branch fuses the global style and content features, and the second branch fuses the detailed style and content features;
[0036] The spatial perception interpolation module is used to fuse the stylized global features and the stylized detailed features;
[0037] The decoder decodes the fused features into a stylized image;
[0038] The discriminator is used to distinguish the authenticity of the generated stylized image and the input style image, and guide the update of the model parameters.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0040] 1. The present invention designs a new style transfer framework for arbitrary style transfer. The input image can be any picture, and a stylized result with good preservation of the structure and details of the input content image can be generated.
[0041] 2. To enhance the global and detailed information of the stylized image, a joint feature extraction module of a reversible neural network and a Transformer encoder is proposed. The Transformer encoder performs well in dealing with long-distance dependencies and sequence data, while the CNN is effective in dealing with local patterns and extracting features. By combining them, the encoder can better capture the long-distance dependencies and local patterns in the sequence data, thereby improving the performance of the model. This combination can also improve the generalization ability of the model, making it applicable to a wider range of tasks and datasets.
[0042] 3. Aiming at the problem of strong style dominance deviation, channel-spatial attention and spatial-channel attention are designed through adaptive channel-spatial interaction. This design enables the style to better consider the expression of the structure of the original content image in space and channels during the transfer, better maintaining the original structure of the image and making the generated image more natural and beautiful. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic structural diagram of the style transfer network generator of the present invention;
[0044] Figure 2It is a schematic structural diagram of the channel spatial attention module and the spatial channel attention module of the present invention;
[0045] Figure 3 It is a schematic structural diagram of the reversible neural network and the spatial perception interpolation module of the present invention;
[0046] Figure 4 It is a qualitative comparison result diagram of different methods;
[0047] Figure 5 It is a user survey result diagram of different methods;
[0048] Figure 6 It is a structural ablation experiment diagram of the present invention;
[0049] Figure 7 It is a loss ablation experiment diagram of the present invention;
[0050] Figure 8 It is a schematic structural diagram of the style transfer device of the present invention. Detailed implementation manners
[0051] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the specific implementation manners and with reference to the accompanying drawings. It should be understood that these descriptions are exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0052] The present invention relates to the field of image style transfer, and mainly solves the problems of image content deviation generated by current methods and the lack of expression in implementation details of current global-statistics-based style transfer methods. With the development of machine learning, attention-based methods are proposed to capture more detailed style transfer effects, but the use of attention makes the generated pictures have artifacts, content deviation and abnormal patterns.
[0053] CSA in the present invention represents the channel spatial attention module, and SCA represents the spatial channel attention module.
[0054] Aiming at the deficiencies of existing solutions, the present invention proposes an arbitrary style transfer method integrating channel and spatial interaction attention mechanisms. The core and main innovation of the image transfer network of the present invention lies in the joint feature encoding module, style conversion module, and spatial perception interpolation module. Each joint feature encoder consists of a Transformer encoder and a reversible neural network, and is arranged in parallel to form a feature extractor, which can effectively extract the global and detailed information of the picture. It also includes a style conversion module for fusing style features and content features. The style conversion module includes two branches, one branch for fusing detailed style and content features, and one branch for fusing global style and content features; the style transfer network also includes a spatial perception interpolation module for fusing global and detailed stylized features, and then decoding through a decoder to output a generated image.
[0055] Figure 1 FIG. is a schematic structural diagram of the generator of the style transfer network of the present invention. The generator includes a joint feature encoder, a style conversion module, a spatial perception interpolation module, and a decoder. Its input data is a style image and a content image, and the output is a stylized image. Among them, I C represents the content image, with a resolution of C×W×H, where C represents the number of channels, H represents the height of the image, and W represents the width of the image. I S represents the style image, with a resolution of also C×W×H.
[0056] Figure 2 (a) is the channel-spatial attention module CSA in the style conversion module, Figure 2 and (b) is the spatial-channel attention module SCA in the style conversion module. These two modules adopt the channel-spatial interaction module and spatial-channel interaction module designed by the present invention to adaptively fuse and weight the style features and content features in the channel and space respectively.
[0057] Specifically, the channel-spatial attention module (CSA) consists of two parts: an attention module and a channel-spatial interaction module. The attention module is usually used to measure the similarity between content features and style features, aiming to find the correspondence between relevant content and style semantic regions. First, the content features and style features are normalized, and then they are embedded into the calculation of the attention map. Then, the attention map is used as an affine transformation to spatially rearrange the style features to obtain weighted style features. The channel-spatial interaction module adaptively adjusts the weights of the structural information or style information in the features in the channel and space according to the distribution of the content features and the weighted style features in the channel and space to achieve better feature fusion.
[0058] More specifically, two operations are designed in the channel - space interaction module: spatial attention (SA) and channel attention (CA). Both spatial attention and channel attention consist of convolutional networks, which are used to generate spatial attention maps and channel attention maps. In the channel - space interaction module, spatial attention calculates the spatial attention map and adaptively adjusts the weights of features according to the attention map to better understand and utilize the spatial structure of features. Similarly, channel attention calculates the channel attention map and weights different channels according to the attention map, enabling the model to better focus on and utilize the information of different channels. Using single - layer attention may not capture all relevant information, resulting in the lack of style patterns in the generated images.
[0059] The present invention designs the spatial - channel attention module SCA as the second attention, which consists of two parts: an attention module and a spatial - channel interaction module. Based on the previous CSA, the fusion module is replaced by a spatial - channel interaction module, and spatial and channel information is added again to enhance the effect of style transfer.
[0060] The joint feature encoding module takes the content image and the style image as inputs and inputs them into the Transformer encoder and the invertible neural network respectively. Benefiting from its self - attention mechanism and multi - head attention mechanism, the Transformer encoder can establish correlations throughout the input sequence, pay attention to different parts simultaneously, and capture global features.
[0061] Figure 3 (a) is the schematic structural diagram of the invertible neural network. The invertible neural network module can effectively extract the detailed features in the input data due to its characteristics of ensuring information integrity through reversibility, providing multi - layer abstractions with deep structures, and optimizing detailed features through backpropagation.
[0062] Figure 3 (b) is the schematic structural diagram of the spatial - aware interpolation module, which is used to fuse the global and local stylized features from the Transformer branch and the invertible neural network branch. However, ordinary additive fusion is prone to color anomalies and halation phenomena. To better fuse features, a spatial - aware interpolation module is designed for feature fusion. It adaptively inserts regional information between the stylized global features and the stylized detailed features. Channel - dense operations apply convolutional kernels of different scales on the cascaded features to fuse multi - scale regional information. The features of the two branches are connected on the channel, and convolutional kernels of different scales are used to fuse multi - scale regional information.
[0063] The method proposed by the present invention constructs an image style transfer network. The MS-COCO dataset is used as the content image and the WiKiArt dataset is used as the style image for learning. The constructed style transfer network includes a joint feature encoder, a style conversion module, a spatial perception interpolation module, a decoder, and a discriminator. The joint feature encoder is used to extract the global and detailed features of the image. The style conversion module is used to transfer the style of the style features to the content image. The spatial perception interpolation module is used to fuse the global stylized features and the detailed stylized features. The decoder is used to decode the features into an image. The discriminator is used to distinguish the authenticity of the style image and the generated image, specifically including:
[0064] Step 1: Prepare the dataset, including the MS-COCO dataset as the content image and the WiKiArt dataset as the style image.
[0065] Step 2: Preprocess the datasets obtained in Step 1. First, scale them to a size of 512X512, and then randomly crop them to a size of 256X256.
[0066] Step 3: Construct and initialize the style transfer network. The style transfer network includes a joint feature encoder, a style conversion module, a spatial perception interpolation module, a decoder, and a discriminator. Among them,
[0067] The joint feature encoder of the style transfer network is composed of two parallel independent classification backbone networks. The independent classification backbone network includes a Transformer encoder and a reversible neural network. The branch where the Transformer encoder is located serves as the global branch, and the branch where the reversible neural network is located serves as the detailed branch. The style conversion module includes two first branches and second branches with the same structure. The first branch is used for the fusion of global style and content features, and the second branch is used for the fusion of detailed style and content features. Both the first branch and the second branch include a channel-spatial attention module and a spatial-channel attention module. The spatial perception interpolation module is used to fuse the stylized global features and the stylized detailed features. The decoder decodes the fused features into a stylized image, and the discriminator is used to distinguish the authenticity of the generated stylized image and the style image.
[0068] Step 4: Input the training data processed in Step 2 into the style transfer network constructed in Step 3 to train the network, specifically including:
[0069] Step 41: Input the content image I in the training set C and the style image I S into the joint feature encoder respectively to extract feature information. After passing through the Transformer encoder and the reversible neural network in the joint feature encoder, the content image I C obtains the global content feature T respectivelyC and the detailed content feature D C , the style image I S After passing through the Transformer encoder and the invertible neural network in the joint encoder, the global style feature T and the detailed style feature D are obtained respectively S and the detailed style feature D S .
[0070] The data processing of the joint feature extraction module specifically includes:
[0071] Step 411: First, use a convolution with a convolution kernel of 4×4 to adjust the channels and size of the input image, and convert an image with a size of (3, 256, 256) into data with a size of (256, 64, 64).
[0072] Step 412: Then, input the converted data into the Transformer encoder for extracting global information and the invertible neural network for extracting detailed information respectively
[0073] In the extraction process of the Transformer encoder, self-attention calculation is first performed on the input sequence to capture the dependencies between different positions in the input sequence. This step enables the model to simultaneously focus on all positions in the input sequence, rather than being limited to a fixed window size. Using residual connections and layer normalization between the self-attention layer and the feed-forward neural network layer helps to alleviate the problems of gradient vanishing or gradient explosion when training deep neural networks. Further non-linear transformation and feature extraction are performed on the feature representation after self-attention calculation
[0074] In the invertible neural network, the input features are split on the channel dimension to create two branches, enabling the two branches to interact during the generation process, and finally merging the two channels together
[0075] Step 42: Input the feature maps extracted in Step 41 into the style conversion module. Specifically, input the global content feature T C and the global style feature T S into the first branch of the style conversion module to obtain the global stylized feature T CS , input the detailed content feature D C and the detailed style feature D S into the second branch of the style conversion module to obtain the detailed stylized feature D CS .
[0076] Both the first branch and the second branch contain a channel-spatial attention module and a spatial-channel attention module, and have the same network structure. Here, the first branch is used as an example for introduction, and the specific processing process includes:
[0077] Step 421: First, input the global content feature TC and the global style feature T S The input channel spatial attention module processes the input features by using three 1×1 convolutions to adjust the input features. One convolution processes the global content feature T C to obtain the feature map Q, and the other two convolutions process the global style feature T S to obtain the feature maps K and V. Then, the feature map Q is multiplied by the feature map K to obtain the semantic relationship map between the content and the style, and then it is mapped by using the Softmax activation function. Subsequently, it is multiplied by the feature map V to obtain the weighted style feature M.
[0078] Step 422: Use a 3×3 convolution to process the global content feature T C to obtain the feature map T * C , in order to achieve the retention of the structure, the weighted style feature M obtained in step 421 is combined with T * C . In the present invention, a spatial channel interaction module is used for fusion. First, channel attention is used to process the feature map T * C to obtain the channel attention coefficient CA, and then spatial attention is used to process the style feature M to obtain the spatial attention coefficient SA. Subsequently, the channel attention coefficient CA is multiplied by the style feature M to dynamically adjust its features in the channel dimension, and the spatial attention coefficient SA is multiplied by the feature map T * C to dynamically adjust its features in the spatial dimension. Finally, the adjusted style feature M and the adjusted T * C are added together to obtain the content feature E CS .
[0079] Step 423: Input the content feature E CS and the global style feature T S into the spatial channel attention module SCA, and use the same method as in step 421 to obtain the weighted style feature M2.
[0080] Step 424: Then, use a 3×3 convolution to process the content feature E CS to obtain the feature map T * CS , in order to achieve the retention of the structure, the style feature M2 obtained in step 423 is combined with the feature map T * CS . In the present invention, a spatial channel interaction module is used for fusion. First, channel attention is used to process the style feature M2 to obtain the channel attention coefficient CA, and then spatial attention is used to process the feature map T* CS Process to obtain the spatial attention coefficient SA, and then multiply the channel attention coefficient CA with the feature map T * CS Multiply them to dynamically adjust its features in the channel dimension, multiply the spatial attention coefficient SA with the style feature M2 to dynamically adjust its features in the spatial dimension, and finally add the adjusted style feature M2 and the feature map T * CS to obtain the global stylized feature T CS 。
[0081] Similarly, for the detailed content feature D C and the detailed style feature D S First process them through the channel-spatial attention module CSA, and then through the spatial-channel attention module SCA to obtain the detailed stylized feature D CS 。
[0082] Step 43: Input the global stylized feature T CS and the detailed stylized feature D CS into the spatial-aware interpolation module for fusion. Specifically:
[0083] Step 431: First, for the input global stylized feature T CS and the detailed stylized feature D CS , concatenate them in the channel dimension, then use a 1X1 convolution kernel for channel concatenation, and then use 1X1, 3X3, and 5X5 convolution kernels to convolve them respectively. After convolution, use the Sigmoid activation function to map them to obtain three mapping values, and finally add the three obtained mapping values and divide by three to get the weight α.
[0084] Step 432: For the input global stylized feature T CS and the detailed stylized feature D CS Adjust them respectively through two 1X1 convolutions, then use the weight α obtained in Step 431 to multiply the global stylized feature T CS to adjust the features, use the weight 1 - α to multiply the detailed stylized feature D CS to adjust it, and finally add them to obtain the final fusion feature F CS 。
[0085] Step 44: Use the decoder to decode the fusion feature F CS to obtain the stylized image I CS 。
[0086] Step 5: Calculate the total loss of the style transfer network, including at least matrix matching loss, perceptual loss, color consistency loss, and adversarial loss. Specifically:
[0087] Step 51: Use the pre-trained VGG network to extract features from layer 1 to layer 5. Calculate the matrix matching loss and perceptual loss using the features from layer 1 to layer 5, and calculate the self-similarity loss and rEMD loss using the features from layer 3 to layer 5.
[0088] Step 52: Calculate the stylized image I CS and the style image I S for the color consistency loss.
[0089] Step 53: Input the stylized image I CS and the style image I S into the discriminator to determine whether it is a style image, and calculate the adversarial loss accordingly.
[0090] Step 6: Steps 4 and 5 are sequentially passed through the set total number of training times. After each fixed number of trainings, save the model weights, then input the test set into the trained image transfer network for testing, and calculate whether the current test metrics of the style transfer network are the highest. If so, end the training; otherwise, adjust the weight parameters of the model loss and retrain.
[0091] To verify the effectiveness of the method of the present invention, the method of the present invention is compared with other existing methods. For a fair comparison, the pre-trained codes officially released by other methods are used, and all methods are implemented in the same computing environment, and quantitative, qualitative analysis, efficiency analysis, and user surveys are carried out simultaneously. The 10 comparison methods are specifically as follows: The AdaIN method of Method 1 is the most classic image style transfer method, which adds the mean and variance of the style image to the content image for style transfer; the Linear method of Method 2 realizes style transfer through a linear weighting method; the SANet method of Method 3 uses the attention mechanism to realize semantic mapping, and realizes high-detail style transfer by calculating the regions where the content and style are similar in artistic style transfer; the MANet method of Method 4 improves the attention input, decouples the content and style features, and proposes an untangling loss function to extract the main style patterns and precise content structures; the MCCNet method of Method 5 uses multi-channel correlation to capture the style information between video frames and uses this information for style transfer; the ArtFlow method of Method 6 realizes unbiased style transfer through reversible neural flow and can maintain the content and structure of the image during the image style transfer process; the AdaAttN method of Method 7 adopts a fusion strategy that combines the attention mechanism and the method based on global feature statistics; the PAMA method of Method 8 uses attention manipulation and spatial-aware interpolation to align the content and style spaces; in addition, the StyTr2 method of Method 9 introduces Transformer and content-aware position encoding. The UniST method of Method 10 adopts a method of wind transfer jointly trained by video and image, and realizes high content retention.
[0092] Table 1 shows the quantitative comparison results of the average content retention index and the average style retention index of 10 different methods.
[0093] Table 1 Comparison Results of Objective Evaluation Indicators of Different Methods
[0094]
[0095] Among them, the content retention metric is used to calculate the structural similarity between two samples. The smaller the value, the smaller the structural difference between the two images, indicating that the model performs better in retaining the structure. The style retention metric is used to evaluate how much style has been successfully transferred from the style image to the generated image. The smaller the value, the smaller the style difference between the style image and the generated image. The self-similarity loss and the image perception similarity metric LPIPS are also used to evaluate the structural difference between the generated image and the content image. Table 1 shows the corresponding quantitative results. Generally speaking, the method of the present invention ranks second in terms of content retention rate, where UniST is the best. In terms of style difference, the method of the present invention also achieves the lowest value, followed by AdaIN. The method also shows good performance in the two metrics of self-similarity loss and LPIPS. The results indicate that the model achieves a good balance between content and style samples.
[0096] To more intuitively illustrate the effectiveness of the method of the present invention, the migration effects of the existing method and the method of the present invention are compared. Figure 4 The following are the comparative results of the qualitative results, where Figure 4 (a) is the content image and the style image, Figure 4 (b) is the style transfer result of the method of the present invention, Figure 4 (c) is the style transfer result of the method UniST, Figure 4 (d) is the style transfer result of the method StyTr2, Figure 4 (e) is the style transfer result of the method PAMA, Figure 4 (f) is the style transfer result of the method AdaAttN, Figure 4 (g) is the style transfer result of the method ArtFlow, Figure 4 (h) is the style transfer result of the method MCCNet, Figure 4 (i) is the style transfer result of the method MANet, Figure 4 (j) is the style transfer result of the method SANet, Figure 4 (k) is the style transfer result of the method Linear, Figure 4 (l) is the style transfer result of the method AdaIN.
[0097] Due to the simplified arrangement of global information, the result of the AdaIN method lacks sufficient style patterns, resulting in crack artifacts in the generated images and affecting the overall transmission quality. The Linear and MCCNet methods modify features through linear projection and per-channel correlation respectively, both producing relatively clean stylized outputs. However, this method cannot adaptively capture the texture patterns of the style image, resulting in the loss of content details and the retention of the colors of the content image. The SANet and MANet methods use attention mechanisms to centrally transfer style features to deep content features. This leads to damaged content structures and unclear textures. Some style patches are even wrongly transferred directly to the content image. The AdaAttN method aims to solve the problems of previous attention-based methods. Nevertheless, it sacrifices the ability to render prominent style patterns and textures, and the retention of details is not very good. The PAMA method uses manifold alignment for image style conversion. The generated stylized images have consistent colors, but this is not perfect for local stylization. The feature representation ability of the flow-based model is limited, so the results of the ArtFlow method often suffer from insufficient or inaccurate styles. The StyTr2 method is the first model to use the Transformer architecture for style conversion, including many parameters. The generated stylized images have poor content details and inconsistent colors. The UniST method is a unified framework for image and video transmission based on Transformer. Unnecessary textures and color inconsistencies appear in the generated stylized images. In contrast, this method uses networks based on Transformer and INN, with better feature representation to capture the global and detailed features of the input image. In addition, the proposed style conversion module effectively solves the distortion problem caused by the simple fusion of single-layer attention results and the original content features. The research results of the present invention can achieve a good balance between the content structure and the style pattern.
[0098] Table 2 shows the running time performance of the method of the present invention and the feedforward method. The present invention was tested on pictures with sizes of 256 pixels and 512 pixels, and all experiments were carried out using a single GTX 1080Ti GPU.
[0099] Table 2 Analysis of the running time performance of different methods
[0100]
[0101] According to the results shown in Table 2, the method proposed by the present invention can obviously generate stylized images effectively in real time.
[0102] Since it is quite subjective to judge whether the artistic stylization is successful, the present invention conducts a user study to compare the quality of artistic expressions. The present invention samples 20 stylized results from random content and style image pairs and compares them with representative models. Figure 5 It is a graph of the results of user surveys of different methods; the present invention collects 600 responses from 30 subjects, Figure 5 and the results show that the method of the present invention is the most effective and can generate results in real time. Figure 5 The ordinate in [graph name] is the number of votes recognized by users.
[0103] To study and verify the method of the present invention, a structural ablation experiment is conducted on the model. To verify the necessity of the Transformer encoder and the reversible neural network in the image encoder. The Transformer encoder is changed to a reversible neural network, that is, both global features and detailed features are extracted by the reversible neural network. Similarly, features of different modalities are extracted through Transformer blocks. Using only the reversible neural network will result in the failure of local migration of images because it is difficult for the reversible neural network to capture the dependencies between the top and the bottom, resulting in unbalanced migration of structured objects. Using all Transformer encoders will result in unreasonable image migration and color distribution.
[0104] As Figure 6 shown, the present invention presents the results of the ablation study to verify the effectiveness of the spatial-aware interpolation module, where Figure 6 (a) is the style picture, Figure 6 (b) is the content picture, Figure 6 (c) is the result output by the complete model, Figure 6 (d) is the result without the spatial-aware interpolation module. From Figure 6 (e) and Figure 6 (f) comparison, it can be seen that the image generated without the spatial-aware interpolation module has serious halo phenomena, and the background color does not match the style color. The bent part of the faucet shows unreasonable migrated colors. In contrast, using the spatial-aware interpolation module can effectively solve these problems and make the generated image more reasonable visually.
[0105] The present invention also conducts a loss ablation experiment to verify the effectiveness of each loss term used for training the model. The results are as Figure 7 shown, where Figure 7 (a) is the content and style pictures, Figure 7 (b) is the result obtained by the complete model, Figure 7 (c) is the result obtained by training without self-similarity loss, Figure 7 (d) is the result obtained by training without perceptual loss, Figure 7 (e) is the result obtained by training without moment matching loss,Figure 7 (f) is the result obtained by training without the rEMD loss. The following points can be analyzed: (1) Without the self-similarity loss, the image loses the diversity of colors and some small texture details are lost. (2) Without the perceptual loss, the structure of the content image is damaged and the style image is reconstructed through mapping. These findings highlight the importance of the perceptual loss for structure maintenance and show that together with the self-similarity loss, it can enhance content inconsistency and improve the style distribution. (3) Without the rEMD loss, the overall stylization degree is reduced. (4) Without the moment matching loss, the result of style transfer is unacceptable and the transformation of textures and lines is hardly distinguishable. These results indicate that the moment matching loss is necessary for the method of the present invention, and combining the rEMD loss can further enhance the style transformation to achieve better results.
[0106] Figure 8 It is a schematic structural diagram of the style transfer device of the present invention. The present invention also proposes an arbitrary style transfer device integrating an interactive attention mechanism. The arbitrary style transfer device includes an image preprocessing module, a joint feature encoder, a style conversion module, a spatial perception interpolation module, a decoder and a discriminator. Among them, the image preprocessing module is used to uniformly scale the style image and the content image in the dataset and crop them to the same size. The joint feature encoder includes a Transformer encoder and a reversible neural network. The Transformer encoder is used to extract the global features of the style image and the content image, and the reversible neural network is used to extract the detailed features of the style image and the content image. The style conversion module is used to fuse the extracted global features and detailed features. Specifically, the first branch fuses the global style and content features, and the second branch fuses the detailed style and content features. The spatial perception interpolation module is used to fuse the stylized global features and the stylized detailed features. The decoder decodes the fused features into a stylized image. The discriminator is used to distinguish the authenticity of the generated stylized image and the input style image and guide the update of the model parameters.
[0107] It should be noted that the above specific embodiments are exemplary. Those skilled in the art can come up with various solutions inspired by the disclosed content of the present invention, and these solutions also belong to the scope of disclosure of the present invention and fall within the protection scope of the present invention. Those skilled in the art should understand that the description and drawings of the present invention are illustrative and do not constitute a limitation to the claims. The protection scope of the present invention is defined by the claims and their equivalents.
Claims
1. An arbitrary style transfer method based on an interactive attention mechanism, characterized in that The specific style transfer method includes the following: Step 1: Prepare the dataset; Step 2: Preprocess the dataset; Step 3: Construct and initialize the style transfer network, which includes a joint feature encoder, a style conversion module, a spatial-aware interpolation module, a decoder, and a discriminator; Step 4: Input the training data processed in Step 2 into the style transfer network for training, which specifically includes: Step 41: Input the content image I in the training set C and the style image I S into the joint feature encoder respectively to extract feature information. After passing through the Transformer encoder and the reversible neural network in the joint feature encoder, I C obtains the global content feature T C and the detailed content feature D C respectively. After passing through the Transformer encoder and the reversible neural network in the joint feature encoder, I S obtains the global style feature T S and the detailed style feature D S respectively; Step 42: Input the features extracted in Step 41 into the style conversion module. Specifically, input T C and T S into the first branch of the style conversion module to obtain the globally stylized feature T CS , and input D C and D S into the second branch of the style conversion module to obtain the detailed stylized feature D CS ; The processing process of the first branch includes: Step 421: First, input T C and T S into the channel spatial attention module for processing. Use three 1×1 convolutions to adjust the input features. One convolution processes T C to obtain the feature map Q, and the other two convolutions process T S to obtain the feature maps K and V. Then multiply Q by K to obtain the semantic relationship map between content and style, and then use the Softmax activation function to map it. Subsequently, multiply it by V to obtain the weighted style feature M; Step 422: Process T using a 3×3 convolution C to obtain the feature map T * C , and input M and T * C into the channel spatial interaction module for fusion. First, use channel attention to process T * C to obtain the channel attention coefficient CA, then use spatial attention to process M to obtain the spatial attention coefficient SA. Subsequently, multiply CA by M to dynamically adjust its features in the channel dimension, multiply SA by T * C to dynamically adjust the features in the spatial dimension. Finally, add the adjusted M and the adjusted T * C to obtain the content feature E CS ; Step 423: Then, input E CS and T S into the spatial channel attention module, and obtain the weighted style feature M2 in the same way as in Step 421; Step 424: Process E using a 3×3 convolution CS to obtain the feature map T * CS , fuse M2 and T * CS through the spatial-channel interaction module. First, use channel attention to process M2 to obtain CA, then use spatial attention to process T * CS to obtain SA. Subsequently, multiply CA and T * CS to dynamically adjust its features in the channel dimension, multiply SA and M2 to dynamically adjust its features in the spatial dimension, and finally add the adjusted M2 and T * CS to obtain the global stylized feature T CS ; Step 43: Input T CS and D CS into the spatial perception interpolation module for fusion to obtain the fused feature F CS ; Step 44: Use a decoder to decode F CS to obtain a stylized image I CS ; Step 5: Calculate the total loss of the style transfer network, including at least matrix matching loss, perceptual loss, color consistency loss, and adversarial loss; Step 6: Steps 4 and 5 go through a set total number of training times in sequence. After each fixed number of training times, save the model weights, and then input the test set into the trained image transfer network for testing.
2. The style transfer method according to claim 1, wherein The specific operation of the spatial-aware interpolation module in Step 43 is as follows: Step 431: First, for the input global stylized feature T CS and the detail stylized feature D CS , splice them on the channel, then use a convolution with a convolution kernel of 1×1 for channel splicing, and then use convolution kernels of 1×1, 3×3, and 5×5 to convolve them respectively. After convolution, use the Sigmoid activation function to map them to obtain three mapping values. Finally, add the three obtained mapping values and divide by three to get the weight α; Step 432: For the input global stylized feature T CS and the detail stylized feature D CS adjust them respectively through two 1×1 convolutions, then multiply the global stylized feature T by the weight α obtained in Step 431 CS to adjust the feature, multiply the detail stylized feature D by the weight 1 - α CS to adjust it, and finally add them up to obtain the final fused feature F CS .
3. The style transfer method according to claim 2, wherein The calculation of the total loss in Step 5 includes: Step 51: Use the pre-trained VGG network to extract features from layer 1 to layer 5. Calculate the matrix matching loss and perceptual loss using the features from layer 1 to layer 5, and calculate the self-similarity loss and rEMD loss using the features from layer 3 to layer 5; Step 52: Calculate the stylized image I CS and the color consistency loss of the style image I S ; Step 53: Input the stylized image I CS and the style image I S into the discriminator to determine whether it is a style image, and thereby calculate the adversarial loss.
4. An arbitrary style transfer device integrating an interactive attention mechanism, characterized in that Any of the style transfer devices is used to implement the style transfer method described in claim 1, and includes an image preprocessing module, a joint feature encoder, a style conversion module, a spatial-aware interpolation module, a decoder, and a discriminator, where The image preprocessing module is used to uniformly scale the style images and content images in the dataset and crop them to the same size; The joint feature encoder includes a Transformer encoder and a reversible neural network. The Transformer encoder is used to extract the global features of the style images and content images, and the reversible neural network is used to extract the detailed features of the style images and content images; The style conversion module is used to fuse the extracted global features and detailed features. The first branch fuses the global style and content features, and the second branch fuses the detailed style and content features; The spatial-aware interpolation module is used to fuse the stylized global features and stylized detailed features; The decoder decodes the fused features into a stylized image; The discriminator is used to distinguish the authenticity of the generated stylized image and the input style image and guide the update of the model parameters.
Citation Information
Patent Citations
Arbitrary style migration method based on multi-attention network
CN114170066A