Image art style migration algorithm based on deep learning, storage medium and equipment
Through the three-source feature fusion architecture of the TFEST network, the balance problem between content maintenance and style transfer of lightweight image style transfer is solved, and efficient and stable style transfer effect is achieved, which improves the visual quality and artistic expression of the image.
Patent Information
- Application Number
- CN202510434122.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing lightweight image style transfer model is difficult to balance content maintenance and style transfer, which makes it difficult to balance the quality and consistency of style transfer, and is prone to artifacts and block effects when processing complex images, affecting visual quality and artistic expression.
The TFEST network based on deep learning is adopted, including a dual-stream feature extraction module, a style feature supplement and a style perception decoder, extract content and style sequences through independent encoding paths, and build a style parameter supplementator using the scaling dot product attention mechanism and residual blocks to realize the fusion of three-source features, dynamically adjust the application weight of the style parameter to ensure that semantics maintain automatic balance with style conversion.
It achieves the quality and consistency of style transfer in lightweight networks, solves the problems of semantic blur and style inconsistency, has the ability to process complex images, and maintains efficient style transfer effect when computing resources are consumed less.
Smart Images

Figure CN120387923A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and specifically relates to an image art style transfer algorithm, storage medium and device based on deep learning. Background Art
[0002] Image style transfer, a key research direction in computer vision, aims to transform the visual expression of an image into a target style while preserving its semantic content. This technology uses deep learning algorithms to analyze and extract the content features of the source image and the artistic features of the style reference image, achieving an organic fusion of the two, thereby creating visual works that retain the recognizable features of the original image while presenting a new artistic style. In terms of implementation mechanism, mainstream methods rely on multi-level feature representations extracted by convolutional neural networks, finding the optimal output by minimizing the weighted sum of content loss and style loss. Professional artists and content creators can use style transfer technology to quickly implement stylization attempts and explore innovative expressions in traditional art forms such as painting and photography. In the process of digitizing cultural heritage, style transfer provides new technical means for the restoration and reproduction of historical relics and artworks, contributing to the protection and inheritance of traditional culture.
[0003] Existing methods for image style transfer using transformers exist. However, these methods suffer from the following issues: Most lack effective feature decoupling mechanisms, making it difficult to separate and control content and style features, resulting in uneven style application in complex scenes. When processing images with rich details and complex textures, existing algorithms are prone to artifacts and blocking effects, which affect the visual quality and artistic expression of the generated images. Traditional convolutional neural network (CNN)-based style transfer methods have limitations in feature expression and insufficiently capture global context, resulting in blurred semantic structure and inconsistent style in the generated images. Existing methods struggle to strike a balance between content preservation and style transfer: overemphasizing style can lead to loss of content semantics, while overpreserving content can lead to insufficient style expression. Lightweight models, while offering advantages such as low computational complexity and memory usage, also have the disadvantage of a small number of model parameters, making it even more difficult to balance content preservation and style transfer. Furthermore, the quality and consistency of style transfer can be compromised or inconsistent due to the lightweight nature of the model. Summary of the invention
[0004] The purpose of the present invention is to provide an image art style transfer algorithm based on deep learning, which is used to solve the technical problems in the existing technology that lightweight style transfer models have difficulty in striking a balance between content preservation and style transfer, and the quality and consistency of style transfer are difficult to take into account with the requirements of model lightweighting.
[0005] The described image art style transfer algorithm based on deep learning includes the following steps.
[0006] Step S1, construct the TFEST network and perform training optimization;
[0007] The TFEST network includes a two-stream feature extraction module, a style feature supplementer, and a style-aware decoder. The two-stream feature extraction module includes two independent encoding paths, namely a content feature encoder and a style feature encoder;
[0008] Step S2, respectively extract the content sequence and the style sequence from the content image and the style image through the independent encoding paths of the two-stream feature extraction module;
[0009] Step S3, introduce the style feature supplementer to process the style image and extract style parameters;
[0010] Step S4, input the content sequence, the style sequence, and the style parameters into the style-aware decoder to output the corresponding style transfer image.
[0011] Preferably, in step S3, the style parameter supplementer extracts and quantifies the style features from the style image through a multi-layer structure. Each layer of the multi-layer structure contains a scaled dot product attention mechanism and a residual block; after the style image is input into the style parameter supplementer, it is segmented into non-overlapping patch units of equal size in the image chunking layer. The features of each patch are composed of the original pixel RGB values, and after being projected by the linear embedding layer, they are input into the multi-layer structure of the style parameter supplementer, and the style parameters are output after being processed by the multi-layer structure.
[0012] Preferably, in the multi-layer structure of the style parameter supplementer, the calculation formula of the scaled dot product attention mechanism is as follows in sequence:
[0013]
[0014] X = LN(MLP(X') + X') = LN(W2σ(W1X' + b1) + b2 + X'),
[0015]
[0016] where, F s represents the input of the multi-layer structure, H represents the number of heads of the scaled dot product attention mechanism, Q, K, and V respectively represent the query vector, the key vector, and the value vector, and are the parameter matrices of the h-th head of each scaled dot product attention mechanism, d kis the dimension of the key vector, W1 and W2 are the weight matrices of the first and second layers in the MLP layer, b1 and b2 are the bias vectors of the first and second layers in the MLP layer, σ represents the activation function, LN represents the layer normalization operation, X”, X' and X are the outputs of the first, second and third layers of the style parameter supplementer in sequence. The extracted style parameter X is dimensionally reduced to obtain the final style parameter
[0017] Preferably, in step S4, the style-aware decoder receives the content sequence Φ c , the style sequence Φ s and the style parameter These three types of inputs. In the feature fusion stage, first, a scaled dot product attention mechanism calculation operation is performed on each input feature. The operation process is expressed as the following formula:
[0018]
[0019] Q = Φ c , K = V = Φ s , where α and γ are learnable weight parameters, Φ″ cs represents the intermediate feature obtained by calculating the content sequence Φ c and the style parameter through the scaled dot product attention mechanism. Φ' cs represents the result obtained by calculating the intermediate feature Φ″ cs and the style sequence Φ s through the scaled dot product attention mechanism. H is the number of heads of the scaled dot product attention mechanism, and are the parameter matrices of the h-th head of each scaled dot product attention mechanism head, and d k is the dimension of the key vector. The processed feature Φ' cs is further processed by a multi-layer perceptron, and the calculation formula is as follows:
[0020]
[0021] where δ is a learnable weight parameter, LN represents the layer normalization operation, MLP represents the multi-layer perceptron, and Φ cs represents the feature output after the style-aware decoder fuses the three-source features of the content sequence, style sequence, and style parameter.
[0022] Preferably, in step S2, in the image block layer, the input content image and the style image in RGB format are respectively segmented into non-overlapping patch units of equal size. The features of each patch are composed of the concatenated original pixel RGB values, and are respectively input into the content feature encoder and the style feature encoder for processing; in the independent encoding path, the features projected by the linear embedding layer are successively processed by a number of AFEM modules. The AFEM module includes an AdaRF module, a CAM module and an AdaCW module. The calculation formula of the AFEM module is: Φ = z + F AdaCW (F CA (F AdaRF (z))), where Φ represents the output of the AFEM module, z represents the input feature, and F AdaRF represents the application of the AdaRF module for processing, F CA represents the application of the CAM module for processing, F AdaCW represents the application of the AdaCW module for processing. The output of the previous AFEM module is downsampled and input into the next AFEM module, and finally the corresponding feature sequence is output. The feature sequence includes the content sequence Φ c and the style sequence Φ s .
[0023] Preferably, the AdaRF module can dynamically adjust the receptive field size of the network and capture both detailed textures and global structures at the same time; the calculation formula of the AdaRF module is:
[0024] F AdaRF (z) = z + α · Conv(z, Ψ(r(z))), where: z represents the input feature, α is a learnable scaling parameter, Conv represents the convolution operation, Ψ is the kernel generation function, and r(z) is the adaptively calculated receptive field parameter; the calculation formula of r(z) is as follows:
[0025] r(z) = σ(MLP(AvgPool(z))),
[0026] where: σ is the sigmoid activation function, MLP represents the multi-layer perceptron, and AvgPool represents the global average pooling.
[0027] Preferably, the calculation formula of the CAM module is:
[0028] F CA (z) = z ⊙ σ(FC2(ReLU(FC1(AvgPool(z))))),
[0029] where: ⊙ represents element-wise multiplication, AvgPool represents the global average pooling, FC1 and FC2 are learnable weight matrices, and σ is the sigmoid activation function.
[0030] Preferably, the calculation formula of the AdaCW module is:
[0031] F AdaCW (z) = z ⊙ σ(W2 · ReLU(W1 · AvgPool(z))),
[0032] where: W1 is the weight parameter of the first-layer transformation, AvgPool represents global average pooling, ReLU represents the ReLU activation function, and W2 is the weight parameter of the second-layer transformation.
[0033] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of an image art style transfer algorithm based on deep learning as described above.
[0034] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor executes the computer program, it implements the steps of an image art style transfer algorithm based on deep learning as described above.
[0035] The present invention has the following advantages:
[0036] 1. The algorithm proposed by the present invention innovatively designs a three-source feature processing architecture. In terms of the automatic balance mechanism of semantic preservation and style conversion, the dual-stream feature extraction module of the present invention extracts the content sequence and the style sequence respectively; the style parameter supplementer quantifies the style information and extracts style parameters to provide a more abstract style representation. When the style-aware decoder receives both types of information simultaneously, it can dynamically adjust the application weight of the style parameters according to the semantic importance of the content features, thereby forming an automatic balance mechanism. This mechanism relies on the information complementarity and dynamic weight adjustment between the outputs of the dual-stream feature extraction module and the style parameter supplementer. And this three-source feature fusion architecture can process information simultaneously at multiple scales, fuse the extracted content sequence, style sequence, and style parameters, ensure that large-scale features (such as overall composition) and small-scale features (such as local texture) can maintain appropriate style transfer, solve the problem of excessive interference of style information on content information in the prior art, and overcome the technical problems of semantic ambiguity and style inconsistency.
[0037] 2. In the construction of the style parameter supplementer of the present invention, the combination of the scaled dot - product attention mechanism and the residual block is adopted. By introducing the style parameter supplementer with this structure into the lightweight network model, the increased computational complexity and the additional consumption of computing resources are relatively small. After construction, the overall network still conforms to the characteristics of a lightweight network. At the same time, the multi - layer scaled dot - product attention mechanism can extract and quantify more abstract and comprehensive style features, that is, it summarizes complex visual styles into key feature descriptors. Therefore, a balance is achieved between maintaining the lightweight of the network model and improving the quality and consistency of style transfer, and it has a better style transfer effect. In terms of the non - linear optimization of computational complexity, the style parameter supplementer compresses the high - dimensional style sequence into a low - dimensional parameter vector, and the style - aware decoder uses the compressed representation to fuse the three - source features, forming an efficient computing resource utilization chain.
[0038] 3. The style - aware decoder in the present invention realizes a multi - level and progressive feature fusion strategy, and precisely integrates the three - source feature information through the style - aware decoder and the multi - head scaled dot - product attention mechanism. The style - aware decoder performs corresponding fusion processing at different levels, progressively at multiple levels of the network, from shallow - layer texture fusion to deep - layer semantic fusion, ensuring that style elements are fully expressed at different granularities while maintaining the integrity of the content structure. This fusion method also has the ability to dynamically adjust the application method of style features according to the semantic information of content features, realizing differential processing of different regions. For example, it can retain more content information in important semantic regions such as the human face and allow stronger style performance in regions such as the background.
[0039] 4. In terms of enhancing the style generalization ability, the combined design of the variable receptive field of the AFEM module and the abstract expression of the style parameter supplementer enables the system to obtain the ability to process unseen styles. The AFEM provides a flexible feature expression basis, the style parameter supplementer abstracts complex styles into parameter vectors, and the style - aware decoder flexibly applies these parameters. The three cooperate to create a generalization ability that exceeds their respective functions. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is the basic flowchart of an image art style transfer algorithm based on deep learning in the present invention.
[0041] Figure 2 It is the structural diagram of the image style transfer network TFEST in the present invention.
[0042] Figure 3 It is the structural diagram of the two - stream feature extraction module in the present invention.
[0043] Figure 4 It is the structural diagram of the AFEM module in the present invention
[0044] Figure 5This is the structural diagram of the style parameter supplementer in the present invention.
[0045] Figure 6 This is the structural diagram of the style-aware decoder in the present invention.
[0046] Figure 7 This is the result diagram of the style arbitrariness of the present invention.
[0047] Figure 8 This is the comparison effect diagram of the results generated by the present invention and the existing baseline method.
[0048] Figure 9 This is the comparison diagram of the content loss and style loss of the style transfer between the present invention and the prior art.
[0049] Figure 10 This is the comparison diagram of the stylization speed of the style transfer between the present invention and the prior art on different pixels.
[0050] Figure 11 This is the comparison diagram of the style transfer between the present invention and the prior art in two dimensions of time and space complexity. Detailed implementation manners
[0051] The following is a more detailed description of the specific implementation manners of the present invention by referring to the accompanying drawings and describing the embodiments, so as to help those skilled in the art have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.
[0052] Embodiment 1
[0053] As Figures 1-6 shown, the present invention provides a deep learning-based image art style transfer algorithm, including the following steps.
[0054] Step S1, construct a TFEST network (Three-source Feature Extraction Image Style Transfer Network) and perform training optimization.
[0055] The TFEST network includes a two-stream feature extraction module, a style feature supplementer, and a style-aware decoder. The two-stream feature extraction module includes two independent encoding paths, namely a content feature encoder and a style feature encoder. As Figure 2 shown.
[0056] Step S2, respectively extract a content sequence and a style sequence from the content image and the style image through the independent encoding paths of the two-stream feature extraction module.
[0057] The structure of the two-stream feature extraction module is as Figure 3As shown. After the image is input into the independent encoding path, the image patch embedding technology is adopted. At the image block layer, the input content image and style image in RGB format are respectively segmented into non-overlapping patch units of equal size. Each patch unit is regarded as a "token", and its feature is composed of the concatenation of the original pixel RGB values (3 channels), with a size of 4×4. Therefore, the feature dimension of each patch is 4×4×3 = 48. Then, the features of the patch units generated by the content image and the style image are respectively input into the content feature encoder and the style feature encoder for processing.
[0058] In the above two feature encoders, a linear embedding layer is applied to the input original features (the original features of the patch units) to project them into an arbitrary dimension (denoted as C), and the corresponding features (content information and style information) are obtained for subsequent processing.
[0059] Both independent encoding paths adopt the AFEM module (Adaptive Feature Extraction Module) to make full use of the advantages of the AFEM module in modeling long-range dependencies and effectively extract global context information. The projected features are successively processed by several AFEM modules. The output of the previous AFEM module is downsampled and then input into the next AFEM module, and finally, the corresponding feature sequences, such as the content sequence or the style sequence, are output. The AFEM module is one of the core components of the TFEST network and consists of three key sub-modules working together, and the structure is as Figure 4 shown. The three key sub-modules are the AdaRF module (Adaptive Receptive Field Module, as shown in Figure 4 (a)), the CAM module (Channel Attention Module, as shown in Figure 4 (b)), and the AdaCW module (Adaptive Channel Weighting Module, as shown in Figure 4(as shown in (c)). The feature information input to the AFEM module is processed by the AdaRF module, the CAM module, and the AdaCW module. The AdaRF module can dynamically adjust the receptive field size of the network, solving the limitation of the fixed receptive field of traditional convolutional networks that cannot adapt to features of different scales, enabling the network to capture both detailed textures and global structures simultaneously. The CAM module learns the dependency relationships and relative importance among channels through global pooling, FC1+ReLU, and FC2+Sigmoid components, and learns the general inter-channel feature mapping through the FC layer, which is suitable for capturing the global correlation among channels and enhancing the effectiveness of feature representation. The AdaCW module can further optimize the adaptive allocation strategy of channel weights, highlighting key features and suppressing secondary features, while keeping the number of tokens unchanged (H / 4, W / 4), achieving a finer adjustment of features, which is beneficial to solving the fusion and matching problem of content features and style features in specific style transfer tasks.
[0060] During the feature extraction process, the calculation formula of the AdaRF module is:
[0061] F AdaRF (z) = z + α · Conv(z, Ψ(r(z))), where: z represents the input feature, α is a learnable scaling parameter, Conv represents the convolution operation, Ψ is the kernel generation function, and r(z) is the receptive field parameter calculated adaptively. The calculation formula of r(z) is as follows:
[0062] r(z) = σ(MLP(AvgPool(z))),
[0063] where: σ is the sigmoid activation function, MLP represents the multi-layer perceptron, and AvgPool represents global average pooling.
[0064] The calculation formula of the CAM module is:
[0065] F CA (z) = z ⊙ σ(FC2(ReLU(FC1(AvgPool(z))))),
[0066] where: ⊙ represents element-wise multiplication, AvgPool represents global average pooling, FC1 and FC2 are learnable weight matrices, and σ is the sigmoid activation function.
[0067] The calculation formula of the AdaCW module is:
[0068] F AdaCW (z) = z ⊙ σ(W2 · ReLU(W1 · AvgPool(z))),
[0069] Where: W1 is the weight parameter of the first - layer transformation, which maps the pooled feature map to an intermediate feature space, then performs a non - linear transformation through the ReLU activation function, and then performs the second - layer transformation. W2 is the weight parameter of the second - layer transformation, and AvgPool represents global average pooling.
[0070] The calculation formula of the AFEM module is: Φ = z+F AdaCW (F CA (F AdaRF (z))), where Φ represents the output of the AFEM module, z represents the input feature, and F AdaRF represents the application of the AdaRF module for processing, and F CA represents the application of the CAM module for processing, and F AdaCW represents the application of the AdaCW module for processing.
[0071] In the overall style transfer network, the AFEM module plays a dual role of feature extraction and feature enhancement. As the network depth increases, through multi - stage processing (4 in the embodiment) implemented by cascading several AFEM modules, the network forms a multi - level structure (from "stage 1" to "stage 4"), gradually reducing the spatial resolution (from H / 8, W / 8 to H / 32, W / 32), and being able to increase the feature dimension (from C to 2C), realizing the conversion from low - level visual features to high - level semantic features. This hierarchical design enables the method to fuse style and content at different abstraction levels, generating more natural and higher - quality style transfer results.
[0072] After being processed by several AFEM modules, the output content sequence Φ of the content feature encoder c , and the output style sequence Φ of the style feature encoder s .
[0073] Step S3, introduce a style feature supplementer to process the style image and extract style parameters.
[0074] The style parameter supplementer extracts and quantifies style features from the style image through a multi - layer structure. Each layer of the multi - layer structure contains a scaled dot - product attention mechanism and a residual block, and the structure is as Figure 5 shown. Similar to step S2, when the style image is input into the style parameter supplementer, the input content image in RGB format is now segmented into non - overlapping patch units of equal size in the image block layer. Each patch unit is regarded as a "token", and its feature is composed of the original pixel RGB values (3 channels) spliced together, with a size of 4×4. Then, a linear embedding layer is applied to the original feature of the patch unit to project it into an arbitrary dimension (denoted as C). Then, the style parameters are obtained after being processed by the above - mentioned multi - layer structure in turn.
[0075] In the multi-layer structure of the style parameter supplementer, the calculation formula of the scaled dot product attention mechanism is as follows:
[0076] X = LN(MLP(X') + X') = LN(W2σ(W1X' + b1) + b2 + X'),
[0077]
[0078] where F s represents the result after the style image is processed by the image block layer and the linear embedding layer, that is, the input of the multi-layer structure. H represents the number of heads of the scaled dot product attention mechanism. This parameter determines the number of feature spaces that the model can focus on simultaneously. The more heads there are, the stronger the model's ability to capture multi-dimensional features. Q, K, and V represent the query vector, key vector, and value vector respectively. and are the parameter matrices of the h-th head of each scaled dot product attention mechanism. d k is the dimension of the key vector. W1 and W2 are the weight matrices of the first and second layers in the MLP (multi-layer perceptron) layer, b1 and b2 are the bias vectors of the first and second layers in the MLP layer, σ represents the activation function, LN represents the layer normalization operation, and X'', X', and X are the outputs of the first, second, and third layers of the style parameter supplementer respectively. This operation stabilizes the training process and accelerates convergence by normalizing the output features of each layer. To improve the interpretability and efficiency of the model, dimensionality reduction processing is performed on the extracted style parameter X to obtain the final style parameter Here U = 100, indicating the dimension of the style parameter vector after dimensionality reduction.
[0079] By calculating the importance weights of each channel of the feature map, the scaled dot product attention mechanism can highlight the feature channels that are most important for the current style transfer task, while suppressing irrelevant channels, achieving adaptive adjustment for different content-style combinations. In the spatial dimension, the scaled dot product attention mechanism can automatically identify the key regions and secondary regions in the image, and apply different intensities of style transfer to different regions. Through this multi-level and multi-dimensional application of the scaled dot product attention mechanism, the network can establish an accurate feature mapping relationship between the content domain and the style domain, such as mapping specific brushstroke patterns in the style image to the corresponding texture regions of the content image, or appropriately applying the color scheme of the style image to different parts of the content image, thereby generating a more natural, harmonious, and artistically expressive stylized image.
[0080] Step S4, input the content sequence, style sequence, and style parameters into the style-aware decoder to output the corresponding style transfer image.
[0081] The feature reconstruction of the style-aware decoder uses an autoregressive prediction strategy that first receives the content sequence Φ c , the style sequence Φ s , and the style parameters as three types of inputs. The content sequence Φ c , the style sequence Φ s , and the style parameters are fused through multiple levels. The structure of the style-aware decoder is as shown in Figure 6 . This step sequentially generates decoupled features through a specific calculation mechanism to ensure the semantic integrity of the original content while integrating the target style features. This design enables the system to effectively capture and transform style features at different scales and finally output the feature Φ cs ∈R L×C .
[0082] Specifically, the features at different levels extracted first are fused to form the input F fus of the style-aware decoder, and the corresponding formula is expressed as: Concat represents concatenation. In the feature fusion stage, first, a scaled dot-product attention mechanism calculation operation is performed on each input feature, and the corresponding operation process is expressed as the following formula:
[0083]
[0084] Q = Φ c , K = V = Φ s ,
[0085] where α and γ are learnable weight parameters, Φ″ cs represents the intermediate feature obtained by calculating the content sequence Φ c and the style parameters through the scaled dot-product attention mechanism, Φ' cs represents the result obtained by calculating the intermediate feature Φ″ cs and the style sequence Φ s through the scaled dot-product attention mechanism, H is the number of heads of the scaled dot-product attention mechanism in the style-aware decoder, and are the parameter matrices of the h-th head of each scaled dot-product attention mechanism, and d k is the dimension of the key vector. The processed feature Φ' cs is further processed by a multi-layer perceptron to form the final fused feature, and the calculation formula is as follows:
[0086]
[0087] where δ is a learnable weight parameter, LN represents layer normalization operation, MLP represents multi-layer perceptron, and Φ csIndicates the features output by the style-aware decoder after fusing the three-source features of the content sequence, style sequence, and style parameters.
[0088] Features Φ that fuse the target style cs Then, through the linear projection layer and the VGG-19 decoder processing in sequence, the finally output style transfer image can be obtained to achieve the conversion of the artistic style.
[0089] In the above steps, steps S2 and S3 are not in sequence and can also be carried out simultaneously.
[0090] Next, the specific technical effects of the above image artistic style transfer algorithm based on deep learning will be described in combination with specific experiments.
[0091] Such as Figure 7 shown, in the embodiment, through the full training and iterative optimization of the network, the style parameter supplementer integrated in the network exhibits excellent feature extraction capabilities. This extractor can accurately identify and extract key style features from any input style encoding and convert them into quantifiable parameter vectors. To test the generalization performance of the present invention for unknown styles, the present invention randomly selects 6 works from the WikiArt dataset and 5 photos from the MSCOCO dataset for style transfer experiments.
[0092] Such as Figure 8 shown, among them, these baseline methods are divided into two categories: one is the classical convolutional neural network-based style transfer method represented by AdaIN; the other is the method using the vision Transformer architecture, including StyTr2, WCT, STTR, and S2WAT. The selection of these methods covers traditional convolutional architectures and modern Transformer architectures, and can provide a comprehensive performance comparison benchmark.
[0093] Such as Figure 9 shown, to evaluate the stylization quality of the method proposed by the present invention, the present invention uses content loss and style loss as the core evaluation indicators. These two indicators respectively measure the degree of difference between the generated image and the original input image in terms of content and style, and the lower the value, the better the transfer effect. In the experimental evaluation framework, the present invention constructs a large-scale test dataset: randomly select 30 content images from the MSCOCO test set, and carefully select 30 artworks representing different genres and having different style features from the WikiArt test set. This combination generates 900 content-style image pairs, providing sufficient test samples for the present invention. The present invention calculates the average content loss and average style loss of each comparison method on all image pairs according to the method in the literature. The method proposed by the present invention obtains the minimum average loss values on these two evaluation indicators. This experimental result strongly verifies the superiority of the method of the present invention over the existing representative methods.
[0094] As Figure 10 shown, the average time for each arbitrary style transfer method to stylize 256x256 pixel and 512x512 pixel images, and all method models are run in the same experimental environment. Although the method of the present invention is slower than the AdaIN method, it is faster than the StyTr2, WCT, STTR, and S2WAT methods in terms of stylization speed. Therefore, the overall stylization rate of the network of the present invention is within an acceptable range. When the image size increases to 512x512 pixels, the stylization speed is basically the same as that of 256x256 pixels, and it also has advantages when compared with other benchmark methods. At the same time, the results also show the stability of the network of the present invention in style transfer.
[0095] As Figure 11 shown, through comparative analysis, it is found that although the iterative model based on CNN has a smaller number of parameters, its floating-point operation count is larger. When using a vision Transformer for style transfer, under the condition of ensuring the same and sufficient data volume, the pure vision Transformer model shows better computational performance, but the number of its parameters increases compared with the CNN method. The image decoder and method based on three-source feature fusion of the present invention have significant advantages in both of these two metrics: reducing the floating-point operation count and significantly improving the number of parameters compared with the ordinary vision Transformer model. This result fully demonstrates the superiority of the method of the present invention in terms of computational efficiency.
[0096] Embodiment 2.
[0097] Corresponding to Embodiment 1 of the present invention, Embodiment 2 of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the following steps are implemented according to the method of Embodiment 1.
[0098] Step S1, construct the TFEST network and perform training optimization.
[0099] Step S2, respectively extract a content sequence and a style sequence from the content image and the style image through the independent encoding path of the dual-stream feature extraction module.
[0100] Step S3, introduce a style feature supplementer to process the style image and extract style parameters.
[0101] Step S4, input the content sequence, the style sequence, and the style parameters into the style-aware decoder to output the corresponding style transfer image.
[0102] In the above steps, Step S2 and Step S3 are not in a specific order and can also be carried out simultaneously.
[0103] The above storage medium includes various media that can store program codes, such as USB flash drives, external hard drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), optical discs, etc.
[0104] For the specific limitations on the steps implemented after the program in the computer-readable storage medium, reference can be made to Embodiment 1, and details will not be elaborated here.
[0105] Embodiment 3.
[0106] Corresponding to Embodiment 1 of the present invention, Embodiment 3 of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor executes the computer program, the following steps are implemented according to the method of Embodiment 1.
[0107] Step S1: Construct a TFEST network and perform training optimization.
[0108] Step S2: Respectively extract a content sequence and a style sequence from a content image and a style image through the independent encoding path of the two-stream feature extraction module.
[0109] Step S3: Introduce a style feature supplementer to process the style image and extract style parameters.
[0110] Step S4: Input the content sequence, the style sequence, and the style parameters into a style-aware decoder to output a corresponding style transfer image.
[0111] In the above steps, Step S2 and Step S3 are not in a specific order and can also be carried out simultaneously.
[0112] For the specific limitations on the steps implemented by the computer device, reference can be made to Embodiment 1, and details will not be elaborated here.
[0113] It should be noted that each block in the block diagrams and / or flowcharts in the accompanying drawings of the present invention, and the combination of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs specified functions or actions, or can be implemented by a combination of dedicated hardware and machine instructions obtained.
[0114] The present invention has been described exemplarily above in conjunction with the accompanying drawings. Obviously, the specific implementation of the present invention is not limited by the above methods. As long as various non-substantive improvements are made by adopting the inventive concept and technical solutions of the present invention, or the inventive concept and technical solutions of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.
Claims
1. An image art style transfer algorithm based on deep learning, characterized in that: It includes the following steps: Step S1, construct a TFEST network and perform training optimization; The TFEST network includes a two-stream feature extraction module, a style feature supplementer, and a style-aware decoder. The two-stream feature extraction module includes two independent encoding paths, namely a content feature encoder and a style feature encoder; Step S2, respectively extract a content sequence and a style sequence from a content image and a style image through the independent encoding paths of the two-stream feature extraction module; Step S3, introduce a style feature supplementer to process the style image and extract style parameters; Step S4, input the content sequence, the style sequence, and the style parameters into the style-aware decoder to output a corresponding style transfer image.
2. The image art style transfer algorithm based on deep learning according to claim 1, characterized in that: In Step S3, the style parameter supplementer extracts and quantifies style features from the style image through a multi-layer structure. Each layer of the multi-layer structure contains a scaled dot-product attention mechanism and a residual block; after the style image is input into the style parameter supplementer, it is segmented into non-overlapping patch units of equal size in the image chunking layer. The features of each patch are composed of the original pixel RGB values stitched together, and after being projected by a linear embedding layer, they are input into the multi-layer structure of the style parameter supplementer, and style parameters are output after being processed by the multi-layer structure.
3. The image art style transfer algorithm based on deep learning according to claim 2, wherein: In the multi-layer structure of the style parameter supplementer, the calculation formula of the scaled dot-product attention mechanism is as follows: X = LN(MLP(X') + X') = LN(W2σ(W1X' + b1) + b2 + X'), Among them, F s represents the input of the multi-layer structure, H represents the number of heads of the scaled dot-product attention mechanism, Q, K, and V represent the query vector, key vector, and value vector respectively, and are the parameter matrices of the h-th head of each scaled dot-product attention mechanism, d k is the dimension of the key vector, W1 and W2 are the weight matrices of the first and second layers in the MLP layer, b1 and b2 are the bias vectors of the first and second layers in the MLP layer, σ represents the activation function, LN represents the layer normalization operation, X”, X', and X are the outputs of the first, second, and third layers of the style parameter supplementer in sequence, and the extracted style parameter X is dimensionally reduced to obtain the final style parameter 4. The image art style transfer algorithm based on deep learning according to claim 1, characterized in that: In step S4, the style-aware decoder receives the content sequence Φ c , the style sequence Φ s and the style parameters as three types of inputs. In the feature fusion stage, first, a scaled dot-product attention mechanism calculation operation is performed on each input feature. The operation process is expressed by the following formula: Q = Φ c , K = V = Φ s , where α and γ are learnable weight parameters, and Φ” cs represents the content sequence Φ c and the style parameter are intermediate features calculated by the scaled dot - product attention mechanism, and Φ' cs represents the intermediate feature Φ” cs and the style sequence Φ s are the results calculated by the scaled dot - product attention mechanism, H is the number of heads of the scaled dot - product attention mechanism, and are the parameter matrices of the h - th head of each scaled dot - product attention mechanism, and d k is the dimension of the key vector. The processed feature Φ' cs is further processed by a multi - layer perceptron, and the calculation formula is as follows: Among them, δ is a learnable weight parameter, LN represents layer normalization operation, MLP represents multi-layer perceptron, and Φ cs represents the feature output after the style-aware decoder fuses the three-source features of the content sequence, style sequence, and style parameters.
5. A deep learning-based image art style transfer algorithm according to claim 1, characterized in that: In step S2, in the image patch layer, the input content image and style image in RGB format are respectively segmented into non-overlapping patch units of equal size. The features of each patch are formed by concatenating the original pixel RGB values, and are respectively input into the content feature encoder and style feature encoder for processing; in the independent encoding path, the features projected by the linear embedding layer are successively processed by a number of AFEM modules. The AFEM module includes an AdaRF module, a CAM module, and an AdaCW module. The calculation formula of the AFEM module is: Φ = z + F AdaCW (F CA (F AdaRF (z))), where Φ represents the output of the AFEM module, z represents the input feature, and F AdaRF represents the application of the AdaRF module for processing, F CA represents the application of the CAM module for processing, F AdaCW represents the application of the AdaCW module for processing. The output of the previous AFEM module is downsampled and input into the next AFEM module, and finally the corresponding feature sequences are output. The feature sequences include the content sequence Φ c and the style sequence Φ s .
6. The image art style transfer algorithm based on deep learning according to claim 5, characterized in that: The AdaRF module can dynamically adjust the receptive field size of the network and capture detailed textures and global structures at the same time; the calculation formula of the AdaRF module is: F AdaRF (z) = z + α·Conv(z, Ψ(r(z))), where: z represents the input feature, α is a learnable scaling parameter, Conv represents a convolution operation, Ψ is a kernel generation function, and r(z) is an adaptively calculated receptive field parameter; the calculation formula of r(z) is as follows: r(z) = σ(MLP(AvgPool(z))), where: σ is the sigmoid activation function, MLP represents a multi-layer perceptron, and AvgPool represents global average pooling.
7. A deep learning-based image art style transfer algorithm according to claim 5, characterized in that: The calculation formula of the CAM module is: F CA (z) = z ⊙ σ(FC2(ReLU(FC1(AvgPool(z))))) where: ⊙ represents element-wise multiplication, AvgPool represents global average pooling, FC1 and FC2 are learnable weight matrices, and σ is the sigmoid activation function.
8. A deep learning-based image art style transfer algorithm according to claim 5, characterized in that: The calculation formula of the AdaCW module is: F AdaCW (z) = z ⊙ σ(W2 · ReLU(W1 · AvgPool(z))), where: W1 is the weight parameter of the first layer transformation, AvgPool represents global average pooling, ReLU represents the ReLU activation function, and W2 is the weight parameter of the second layer transformation.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the steps of a deep learning-based image art style transfer algorithm as described in any one of claims 1-8.
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of a deep learning-based image art style transfer algorithm as described in any one of claims 1-8.
Citation Information
Patent Citations
Multi-domain image fusion method and system based on deep learning
CN116229229A
Style migration method based on comparative learning and attention mechanism, computer equipment, readable storage medium and program product
CN118014822A
Audio and video dual-mode emotion recognition method and system based on adapter fusion
CN120411863A
Image style transfer method and apparatus, and image style transfer model training method and apparatus
WO2022048182A1
Cited By
Human portrait stylization system and method based on attention mechanism and diffusion model
CN121437253A
AIGC cross-medium-based meta-universe scene dynamic generation method
CN121614035A
Aigc cross-medium based metaverse scene dynamic generation method
CN121614035B