Image art style migration algorithm, storage medium and equipment

The algorithm addresses the challenges of image style transfer by integrating CSAFM and SimAM with a multi-level perceptual loss, enhancing feature fusion and detail preservation, resulting in superior image quality and natural style transfer.

CN120318059AActive Publication Date: 2025-07-15ANHUI POLYTECHNIC UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510506535.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-15
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

In the prior art, the image style transfer algorithm has poor quality during the feature fusion stage, it is difficult to retain the details and structural information of the image, and the local details are insufficient, and the computing resource requirements are large.

Method used

The channel-space adaptive fusion module (CSAFM) and parameterless attention module (SimAM) are used to enhance feature expression, and combined with multi-level perceived loss function, the architectural design of generator and discriminator is improved through adaptive feature selection and multi-scale feature fusion.

Benefits of technology

It improves the quality and nature of image style transfer, enhances the perception ability of local features, maintains the integrity of image structure, and reduces the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318059A_ABST
    Figure CN120318059A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to an image artistic style migration algorithm, which comprises the following steps: S1, constructing a generative network G, the basic framework of which comprises an encoder and a decoder; s2, constructing a discrimination network D which adopts a Markov discriminator and is composed of a series of continuous convolution modules; and S3, adversarial training is carried out on the generative network G and the discrimination network D, and the trained generative network G is used for generating a target style migration image. According to the method, image structure information and style features can be effectively reserved in the feature fusion process, the cross-scale feature fusion effect is greatly improved, and the generated image has richer details and more natural style transition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to an image art style transfer algorithm, a storage medium, and a device. Background Art

[0002] As an important research direction in the field of computer vision, image style transfer has broad application prospects in fields such as art creation, visual design, and media entertainment. This technology is essentially a method for image conversion between different visual domains, aiming to transfer the style features of the source image to the content of the target image while maintaining the integrity of the content structure. However, this technology faces challenges such as the balance problem between content preservation and style conversion, poor local detail processing quality, and high computational resource requirements.

[0003] For example, in the prior art, the generator structure of style transfer lacks an effective feature selection mechanism in the downsampling stage, resulting in difficulties in accurately capturing local significant features, and in the upsampling stage, a simple feature splicing method is used, which cannot make full use of multi-scale feature information and other problems. In the feature fusion stage, on the one hand, due to the aforementioned problems in the prior art, the feature information is not rich enough and the accuracy is insufficient. On the other hand, the existing conventional feature fusion methods also cannot make the generated results meet the satisfactory standard in terms of comprehensiveness and depth. There are problems such as insufficient local detail features and excessive noise points in the transferred image, and at the same time, some important feature information cannot be recognized and highlighted. Therefore, how to improve the quality and effect of feature fusion and retain sufficient detail and structure information in the image has become a technical problem to be solved in the prior art. Summary of the Invention

[0004] The purpose of the present invention is to provide an image art style transfer algorithm to solve the technical problems of poor quality and effect of feature fusion in the prior art and difficulty in repeatedly retaining sufficient detail and structure information in the image.

[0005] The described image art style transfer algorithm includes the following steps.

[0006] Step S1, construct a generation network G. The basic framework of the generation network G includes an encoder and a decoder; the decoder uses a channel-space adaptive fusion module to fuse the upsampling features with the features of the downsampling module of the encoder; the channel-space adaptive fusion module includes a channel attention mechanism, a spatial attention mechanism, and an adaptive feature selection stage; the adaptive feature selection stage dynamically integrates and selects the features enhanced by dual attention, then fuses based on the relationship between different channel features in the adaptively selected features, and finally generates an output result;

[0007] Step S2: Construct a discriminant network D. The discriminant network D adopts a Markov discriminator and consists of a series of consecutive convolutional modules;

[0008] Step S3: Conduct adversarial training on the generator network G and the discriminant network D. The trained generator network G is used to generate target style transfer images.

[0009] Preferably, the adaptive feature selection stage accepts the feature map processed by the channel attention mechanism and the spatial attention mechanism as input, establishes associations between channels through 1×1 convolution, and captures the interdependence between different channel features; subsequently, batch normalization is applied to stabilize the training process and accelerate convergence, while reducing internal covariate shift; finally, a threshold transfer unit is introduced to introduce a non-linear transformation, enhance the model's expressive ability, and filter out irrelevant information; the expression of adaptive feature selection is:

[0010]

[0011] where, is the feature after adaptive selection of output channel a, represents the linear transformation matrix that maps the input to output channel a, b a is the bias term of output channel a, μ a is the channel mean vector of output channel a, β a is the learnable offset parameter of output channel a, is the feature variance of output channel a, γ a is the learnable scaling parameter of output channel a, ∈ is the numerical stability constant.

[0012] Preferably, the adaptively selected feature F a is linearly transformed in the channel dimension through 1×1 convolution operation to establish the connection between different channel features, and then batch normalization operation is applied. By calculating the mean μ B,c and variance of each channel, the feature is normalized to zero mean and unit variance, and then adjusted by the learnable scaling parameter γ c and offset parameter β c to stabilize the training process effectively and accelerate convergence; finally, a threshold transfer unit is introduced to introduce a non-linear transformation, enhance the model's expressive ability and suppress negative value responses, and filter out irrelevant information to complete the final feature fusion; the linear transformation performed in the channel dimension is expressed as W f ·F a +b f where W f is the weight matrix, b f is the bias term; the complete mathematical expression of the final feature fusion is:

[0013]

[0014] Among them, O c,h,w is the value of the final output feature at position (c, h, w), γ c is the scaling parameter of the c-th output channel, W f,c,i is the weight from input channel i to output channel c, F a,i,h,w is the value of the feature after adaptive selection at position (i, h, w), b f,c is the bias term of output channel c, μ B,c is the statistical mean of output channel c, is the statistical variance of output channel c, ∈ is the numerical stability constant, β c is the offset parameter of output channel c.

[0015] Preferably, the processing flow of the channel-spatial adaptive fusion module includes: first, concatenating the upsampled feature and the skip connection feature in the channel dimension, and then introducing the channel attention mechanism; first, compressing the spatial dimension through adaptive global average pooling to extract channel-level features, and then generating channel attention weights through the non-linear transformation of dimensionality reduction convolution, threshold transfer unit, and dimensionality increase convolution; further introducing the spatial attention mechanism, which first calculates the average value and maximum value of the feature map in the channel dimension, concatenates the average value and the maximum value, and generates a spatial attention map through convolution and non-linear activation units, that is, the feature map input to the adaptive feature selection stage.

[0016] Preferably, an attention-free self-attention module is added after each downsampling module of the encoder to enhance feature representation; the processing flow of the attention-free self-attention module includes: first calculating the spatial mean ω of the feature Figure X and then calculating the squared difference (X - ω) between each feature position and the mean 2 , then calculating the normalized variance ρ of the entire feature map 2 , and further calculating the preliminary attention score based on the squared difference of each position. The calculation formula is:

[0017] In the formula, a is the scaling factor, ε is a small constant used to prevent division by zero error, b is the offset constant; finally, the preliminary attention score y is transformed into a weight coefficient between 0 and 1 through the non-linear activation function ρ(y), and these weights are applied to the original feature. The calculation formula is: X out = X × ρ(y), where X out is the weighted output feature map, and × is the element-wise multiplication.

[0018] Preferably, the encoder is input through an initial module, which is sequentially connected to a reflection padding layer, a convolutional layer, a normalization layer, and a threshold transfer unit; then it is connected to two downsampling modules, each of which consists of a convolutional layer, a normalization layer, and a threshold transfer unit, and a parameter-free self-attention module is added after each downsampling module to enhance the feature representation, thus forming the encoder; the decoder includes two upsampling modules, each of which includes a transposed convolutional layer, a normalization layer, and a threshold transfer unit connected in sequence; then a channel-spatial adaptive fusion module is used to fuse the upsampled features with the features of the downsampling modules of the encoder; finally, it is connected to a reflection padding layer, a convolutional layer, and a hyperbolic tangent non-linear layer to form the decoder.

[0019] Preferably, in the convolutional module of the discriminative network D, the first convolutional module includes a convolutional layer and a ramp rectified linear unit; each of the intermediate convolutional modules includes a convolutional layer, an instance normalization layer, and a ramp rectified linear unit, and multi-scale feature representations of the input image are extracted through layer-by-layer convolution and downsampling operations; the last convolutional module only includes a convolutional operation to generate the final discriminative feature map.

[0020] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of an image art style transfer algorithm as described above.

[0021] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor, and when the processor executes the computer program, it implements the steps of an image art style transfer algorithm as described above.

[0022] The present invention has the following advantages:

[0023] 1. The present invention designs a CSAFM channel-spatial adaptive fusion module and integrates it into the upsampling stage of the generator, solving the limitations of simple feature stitching in traditional style transfer. This module adopts a dual mechanism of channel attention and spatial attention to enhance feature expression from the channel and spatial dimensions respectively, and realizes the intelligent fusion of multi-scale features through an adaptive feature selector. Compared with the existing AFF (Adaptive Feature Fusion) technology that uses a recursive fusion method, in the architecture design, on the one hand, the CSAFM is integrated into the upsampling stage of the generator, focusing on the effective fusion of upsampled features and downsampled features, and performing adaptive selection first in the fusion stage, that is, performing various attention mechanisms and feature selection processing, and then improving the final fusion method to achieve a more complex processing flow, enhancing the comprehensiveness and depth of feature fusion, and is particularly suitable for detail retention and structure reconstruction in image generation tasks.

[0024] 2. The present invention is also improved by introducing the SimAM parameter-free attention module. This module generates attention weights by calculating the local self-similarity of feature maps during the downsampling stage of the generator, without introducing additional parameters. Compared with traditional channel attention mechanisms, SimAM operates directly in the original feature space, more effectively capturing the spatial correlation of local features and significantly enhancing the network's perception ability of local salient features.

[0025] 3. The present invention constructs a multi-level perceptual loss based on the VGG16 network, making up for the deficiency of the original loss function that only relies on adversarial loss and cycle consistency loss. This multi-level perceptual loss mechanism realizes hierarchical supervision from low-level visual features to high-level semantics by matching shallow, middle, and deep features and respectively constraining texture details, local structures, and high-level semantic information. The design of the loss function provides richer gradient information for the generator, enabling the model to simultaneously focus on local details and global semantics, effectively improving the quality and naturalness of style transfer and achieving a more delicate style transfer effect while maintaining the integrity of the image structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is the basic flowchart of the image art style transfer algorithm in the present invention.

[0027] Figure 2 is the structural diagram of the overall network in the present invention.

[0028] Figure 3 is the structural diagram of the generator network G in the present invention.

[0029] Figure 4 is the structural diagram of the CSAFM module in the present invention.

[0030] Figure 5 is the structural diagram of the discriminator network D in the present invention.

[0031] Figure 6 is the comparison diagram after marking the defect areas of the generation results of the present invention and the existing style transfer models.

[0032] Figure 7 is the comparison diagram of the generation results of the present invention and the existing style transfer models.

[0033] Figures 8 - 12 are the effect diagrams of the generation results of the present invention and the existing style transfer models on five indicators of FID, SSIM, LPIPS, PSNR, and MSE in sequence. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The following will, with reference to the accompanying drawings, further elaborate on the specific embodiments of the present invention through the description of embodiments, so as to help those skilled in the art have a more complete, accurate and in-depth understanding of the inventive concept and technical solutions of the present invention.

[0035] Embodiment 1.

[0036] As Figures 1 - 5 shown, the present invention provides an image art style transfer algorithm, including the following steps.

[0037] Step S1, construct a generation network G. The basic framework of the generation network G includes an encoder and a decoder.

[0038] The basic framework is mainly composed of functional modules such as standard convolutional blocks, downsampling modules, residual modules, and upsampling modules. Specifically, the encoder is input through an initial module, which is sequentially connected to a reflection padding layer, a 7×7 convolutional layer, a normalization layer, and a threshold transfer unit; then two downsampling modules are connected, and each downsampling module consists of a 3×3 convolutional layer, a normalization layer, and a threshold transfer unit. Compared with the prior art, the present invention adds a parameter-free self-attention module (SimAM) after each downsampling module to enhance feature representation, forming the encoder.

[0039] The decoder mainly includes two upsampling modules, and each module is sequentially connected to a transposed convolutional layer, a normalization layer, and a threshold transfer unit; compared with the prior art, the present invention uses a channel-spatial adaptive fusion module (CSAFM) to fuse the upsampled features with the downsampled features of the encoder, and finally connects a reflection padding layer, a 7×7 convolutional layer, and a hyperbolic tangent nonlinear layer to form the decoder.

[0040] 1) The processing flow of SimAM includes: first, calculate the spatial mean of the feature map :

[0041] In the formula, B is the batch size, representing the number of samples in a batch of data, C is the number of channels, representing the depth of the feature map, H is the height of the feature map, representing the number of pixels in the vertical direction, W is the width of the feature map, representing the number of pixels in the horizontal direction, and X i,j is the feature value at position (i, j).

[0042] Then calculate the squared difference (X - ω) of each feature position from the mean 2 , and then calculate the normalized variance of the entire feature map:

[0043]

[0044] In the formula, X i,j-ω is the squared difference between the eigenvalue and the mean at position (i,j), and H×W - 1 is the degree of freedom for unbiased variance calculation in statistics.

[0045] Furthermore, divide the squared difference at each position by four times the normalized variance plus a small constant ε, and then add 0.5 to obtain the preliminary attention score. The calculation formula is:

[0046]

[0047] In the formula, a is the scaling factor, which is taken as 4 in the embodiment to control the influence amplitude of variance differences; ε is the small constant for numerical stability to prevent division-by-zero errors; b is the offset constant, which is taken as 0.5 in the embodiment to ensure the baseline value of the attention score.

[0048] Finally, convert the preliminary attention score y into a weight coefficient between 0 and 1 through the non-linear activation function ρ(y), and apply these weights to the original features. The calculation formula is: X out = X × ρ(y), where X out is the weighted output feature map, and × is the element-wise multiplication.

[0049] This mechanism requires no additional training parameters. By adaptively adjusting the feature weights, it effectively enhances the feature discrimination ability and receptive field, improves the network's ability to capture key information, and solves the problem of information loss in the traditional downsampling process.

[0050] 2) CSAFM includes a channel attention mechanism, a spatial attention mechanism, and an adaptive feature selection stage. The processing flow of CSAFM is as follows: First, concatenate the upsampled feature and the skip connection feature in the channel dimension to obtain the feature concatenation result F concat , and the calculation formula is: In the formula, F up is the upsampled feature, F down is the downsampled feature, C = C1 + C2 is the total number of channels after concatenation, H is the height of the feature map, representing the number of pixels in the vertical direction, W is the width of the feature map, representing the number of pixels in the horizontal direction, and [·,·] is the concatenation operation in the channel dimension.

[0051] After that, introduce the channel attention mechanism. The core calculation process of the channel attention mechanism is divided into two key steps: First, compress the spatial dimension through adaptive global average pooling to extract channel-level features, and then generate channel attention weights through non-linear transformations of dimensionality reduction convolution, threshold transfer unit, and dimensionality increase convolution. This mechanism can significantly reduce the number of parameters while capturing important relationships between channels, achieving efficient channel feature enhancement. The feature after channel attention enhancement is calculated using the following formula:

[0052]

[0053] In the formula, F c ∈R C×H×W is the feature after channel attention enhancement, ⊙ is the Hadamard product (i.e., element-wise multiplication), and σ is the Sigmoid function; while is the dimensionality reduction transformation matrix, r is the dimensionality reduction ratio, is the dimensionality increase transformation matrix, max(0, x) is the ReLU function, which returns when x > 0 and returns 0 otherwise, is the global average pooling operation.

[0054] Based on the feature with channel attention enhancement, CSAFM further introduces a spatial attention mechanism to learn the importance distribution of the feature map in the spatial dimension. Specifically as follows: The spatial attention mechanism first calculates the average value and the maximum value of the feature map in the channel dimension, concatenates them, and generates a spatial attention map through a 7×7 convolution and a non-linear activation unit.

[0055] The feature after spatial attention enhancement is calculated using the following formula:

[0056]

[0057] In the formula, F s ∈R C×H×W is the feature after spatial attention enhancement, f 7×7 ∈R 1×2×7×7 is the convolution transformation, which receives inputs of 2 channels and outputs 1 channel, is to calculate the average value along the channel dimension, is to calculate the maximum value along the channel dimension, (;) represents the feature concatenation operation between two features.

[0058] The adaptive feature selection stage dynamically integrates and selects the features after dual attention enhancement; this stage receives the feature map after channel-spatial dual attention processing as input, establishes associations between channels through 1×1 convolution, and captures the interdependent relationships between different channel features; subsequently, batch normalization is applied to stabilize the training process and accelerate convergence, while reducing internal covariate shift; finally, a non-linear transformation is introduced through a threshold transfer unit to enhance the model's expressive ability and filter out irrelevant information. The expression for adaptive feature selection is:

[0059]

[0060] Among them, is the feature after adaptive selection of output channel a, represents the linear transformation matrix that maps the input to output channel a, b a is the bias term of output channel a, μ ais the channel mean vector of output channel a, β a is the learnable offset parameter of output channel a, is the feature variance of output channel a, γ a is the learnable scaling parameter of output channel a, and ∈ is the numerical stability constant.

[0061] to receive the feature map F processed in the adaptive feature selection stage a as input, perform a linear transformation in the channel dimension through a 1×1 convolution operation to establish the connection between features of different channels. This transformation can be expressed as W f ·F a +b f , where W f is the weight matrix, and b f is the bias term; subsequently, apply the batch normalization operation to normalize the features to zero mean and unit variance by calculating the mean μ B,c and variance of each channel, and then adjust the distribution through the learnable scaling parameter γ c and offset parameter β c to effectively stabilize the training process and accelerate convergence; finally, introduce a non-linear transformation through the threshold transfer unit to enhance the model's expressive ability and suppress negative value responses, filtering out irrelevant information to complete the final feature fusion. The complete mathematical expression for the final feature fusion is:

[0062]

[0063] where, O c,h,w is the value of the final output feature at position (c, h, w), γ c is the scaling parameter of the c-th output channel, W f,c,i is the weight from input channel i to output channel c, F a,i,h,w is the value of the feature after adaptive selection at position (i, h, w), b f,c is the bias term of output channel c, μ B,c is the statistical mean of output channel c, is the statistical variance of output channel c, ∈ is the numerical stability constant, and β c is the offset parameter of output channel c.

[0064] The above - mentioned generation network adopts a multi - stage processing flow: feature splicing → channel attention → spatial attention → adaptive selection → final fusion. In contrast, the existing AFF (Adaptive Feature Fusion) technology adopts a recursive fusion method. For the generation network applied in this method, in terms of architecture design, on the one hand, the up - sampling stage of the generator integrates CSAFM, focusing on the effective fusion of up - sampled features and down - sampled features, and first performs adaptive selection in the fusion stage, that is, performs multiple attention mechanisms and feature selection processing, and then improves the final fusion method to achieve a more complex processing flow, enhancing the comprehensiveness and depth of feature fusion, and is particularly suitable for detail retention and structure reconstruction in image generation tasks.

[0065] Step S2, construct a discriminator network D. The discriminator network D adopts a Markov discriminator and is composed of a series of consecutive convolutional modules.

[0066] In this step, the discriminator network D provides an accurate adversarial training signal for the generator by distinguishing real images from the images synthesized by the generator. Among them, the first module contains a convolutional layer and a ramp - type rectified linear unit; the middle modules contain a convolutional layer, an instance normalization layer, and a ramp - type rectified linear unit, and extract multi - scale feature representations of the input image through layer - by - layer convolution and down - sampling operations; the last module only contains a convolutional operation to generate the final discriminant feature map.

[0067] This discriminator is designed as a fully convolutional network. Each point in the finally output feature map corresponds to a receptive field of about 70×70 pixels in the original image, realizing the independent discrimination of local regions of the image. The discriminant strategy based on image patches not only breaks through the limitation of traditional discriminators for overall evaluation of the entire image, reduces the number of parameters and computational complexity, but also can more accurately perceive and evaluate the detailed features and texture information of local regions, which is beneficial for the generator to learn better local structures and texture details.

[0068] Step S3, perform adversarial training on the generator network G and the discriminator network D. The trained generator network G is used to generate target style - transferred images.

[0069] In this step, the original style - transfer model mainly relies on adversarial loss and cycle - consistency loss to achieve image style conversion, but this design has limitations in feature expression and semantic understanding: the adversarial loss only provides supervision in the pixel space and it is difficult to ensure the authenticity of high - level features; the cycle - consistency loss is based on a simple L1 distance metric and cannot effectively express hierarchical semantic information.

[0070] To address these problems, the present invention introduces a multi - level perceptual loss mechanism based on pre - trained VGG16 to construct a feature matching framework from low - level to high - level. Specifically as follows:

[0071] (1) Shallow feature matching is responsible for capturing local visual features such as textures and edges, ensuring the authenticity of the generated image at the detail level;

[0072] (2) Middle-level feature matching focuses on the medium-scale structure and local semantic information of the image, enhancing the coherence of the content;

[0073] (3) High-level feature matching extracts high-level semantic representations, ensuring the consistency of the generated image at the overall semantic level.

[0074] The multi-level perceptual loss provides rich gradient information for the generator, enabling it to simultaneously focus on local details and global semantics, effectively compensating for the deficiency that it is difficult to maintain the consistency of deep features only relying on pixel-level loss, thus significantly improving the quality and naturalness of style transfer.

[0075] The loss function of the present invention consists of four parts: adversarial loss, cycle-consistency loss, identity loss, and perceptual loss.

[0076] First is the adversarial loss, which is used to ensure that the generated image looks real and conforms to the feature distribution of the target domain. Its adversarial loss can be expressed as:

[0077]

[0078] To ensure the bidirectional consistency of image conversion and prevent mode collapse, cycle-consistency loss is introduced:

[0079]

[0080] The identity loss is used to maintain the color consistency of the source-domain image:

[0081]

[0082] The perceptual loss is obtained by weighted accumulation of the mean square error between the feature maps of three different levels of the VGG16 network, and is used to measure the difference between the generated image and the real image in terms of high-level semantic features. Its formula is expressed as:

[0083]

[0084] Among them, Φ i (x) represents the feature extraction function, which extracts image features from different abstraction levels, x represents the real image, G(y) represents the image generated by the generator, and λ is the weight coefficient of each feature level. represents the L2 norm (mean square error).

[0085] Therefore, the improved total loss function can be expressed as:

[0086]

[0087] Among them, λ GAN 、λ cyc 、λ identity 、λ perceptual are the corresponding weight coefficients in turn. The combined design of multiple loss functions improves the overall performance of the model by optimizing complementary objective functions, enabling the generated images to not only maintain authenticity in texture details but also maintain coherence and rationality in semantic features, providing a more perfect quality assurance mechanism for the image generation task.

[0088] Next, a process of an image art style transfer algorithm, storage medium, and device will be described in combination with specific experiments.

[0089] As Figure 6 shown, from left to right are the input image, CycleGAN algorithm, DualGAN algorithm, DiscoGAN algorithm, CUT algorithm, and the algorithm of the present invention (hereinafter denoted as AFST-GAN).

[0090] When the adaptive module is not used, the traditional feature fusion method has the following defects:

[0091] 1. Local detail loss: As shown in the area marked by the red box, other methods such as CycleGAN, DualGAN, and DiscoGAN cannot well retain important detail information in the converted image.

[0092] 2. Noise and artifacts: The comparison methods introduce unnecessary noise and visual artifacts during the processing, which is particularly obvious in areas with rich textures (such as forest areas and rock surfaces).

[0093] 3. Insufficient recognition of feature information: Traditional methods are difficult to recognize and highlight important feature information, resulting in the converted image lacking key visual elements.

[0094] As Figures 7 - 12 shown, for the description experiment of verifying semantic descriptors in similar scenarios, an objective and subjective evaluation scheme was designed to verify the performance of the proposed AFST-GAN model.

[0095] ​In the objective evaluation stage, four representative benchmark models were selected for comparison: CycleGAN, DualGAN, DiscoGAN and CUT. The images generated by each model were evaluated overall using five indicators: SSIM (Structural similarity index measure), PSNR (Peak signal-to-noise ratio), MSE (Mean squared error), LPIPS (Learned perceptualimage patch similarity) and FID (Frechet inception distance). Figures 8 - 12 The results show that the proposed method outperforms the comparison method in overall performance. In terms of quantitative analysis of objective evaluation indicators, AFST-GAN achieved leading performance on the summer2winter style transfer dataset: the FID value was as low as 51.63, which was 21.32%, 45.30%, 40.32% and 9.37% lower than CycleGAN, DualGAN, DiscoGAN and CUT, respectively, indicating that the quality and authenticity of the generated images were significantly improved; the SSIM value reached 0.90, which was 7.14%, 20.00%, 28.57% and 36.36% higher than the other four methods, respectively, proving that The structural similarity is significantly improved; the PNSR value is as high as 25.03, which is 5.43%, 47.58%, 74.18% and 29.35% better than the comparison methods, respectively, reflecting the improvement of image reconstruction quality; the LPIPS value is only 0.12, and the lower perceptual distance is reduced by 25.00%, 64.71%, 52.00% and 57.14%, respectively, indicating that the generated image has a higher perceptual consistency with the target style; the MSE value is reduced to 69.02, which is reduced by 8.51%, 22.18%, 32.39% and 22.14%, respectively, proving the effective reduction of pixel-level errors.

[0096] Based on this, AFST-GAN has achieved the best level in all five evaluation indicators, especially in FID and LPIPS, the key indicators for measuring perceptual quality. These objective data fully verify that the AFST-GAN algorithm proposed in this paper can not only maintain the structural integrity and detail information of the image, but also achieve a more natural and harmonious style transfer effect, which fully proves the effectiveness and advancement of the improved method from multiple dimensions.

[0097] In subjective qualitative analysis, Figure 7The results show that CycleGAN has obvious texture distortion and detail loss problems in landscape image translation. Especially when dealing with complex terrains, the generated images often appear blurred and artifacted. In addition, CycleGAN shows instability in maintaining the original scene structure and sometimes produces unnatural color transitions.

[0098] DualGAN shows serious color problems during image style translation. Some images have obvious purple and gold color tone deviations, and this unnatural color conversion significantly reduces the realism of the images. In addition, when dealing with shadow areas, the ability to maintain details is insufficient, resulting in a weakened sense of hierarchy in the images.

[0099] Although DiscoGAN can achieve basic style translation, it has obvious defects in maintaining the integrity of the image structure. Especially when dealing with complex scenes, problems of structural distortion are likely to occur. At the same time, this model performs poorly in dealing with local details of images, and the generated images often show overly smooth features, resulting in serious loss of detail information and reducing the visual quality of the generated images.

[0100] The winter images generated by CUT maintain a relatively good original structure and have a unique style in light processing. However, some images show the characteristics of being overly bright, and the winter features are not obvious enough. It is worth noting that CUT performs well in the first and fourth column images, successfully converting summer landscapes into relatively realistic winter scenes, but the overall consistency is not as stable as other models.

[0101] AFST-GAN demonstrates better overall performance: it is superior in maintaining the integrity of the original scene structure, especially in dealing with complex terrains, it can better maintain the landform features; it is more excellent in maintaining image texture details; the color translation is more natural and harmonious, avoiding common color distortion problems in other algorithms; it shows stronger adaptability in dealing with light changes, can better balance the details of bright and dark areas, making the images more similar to real winter scenes.

[0102] Embodiment 2.

[0103] Corresponding to Embodiment 1 of the present invention, Embodiment 2 of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the following steps are implemented according to the method of Embodiment 1.

[0104] Step S1, construct a generation network G, and the basic framework of the generation network G includes an encoder and a decoder.

[0105] Step S2, construct a discriminant network D, and the discriminant network D adopts a Markov discriminator, which is composed of a series of consecutive convolutional modules.

[0106] Step S3: Conduct adversarial training on the generation network G and the discriminant network D. The trained generation network G is used to generate target style transfer images.

[0107] The above storage medium includes various media that can store program codes, such as USB flash drives, external hard drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), optical discs, etc.

[0108] For the specific limitations on the steps implemented after the program in the above computer-readable storage medium, reference can be made to Embodiment 1, and details will not be elaborated here.

[0109] Embodiment 3.

[0110] Corresponding to Embodiment 1 of the present invention, Embodiment 3 of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor executes the computer program, the following steps are implemented according to the method of Embodiment 1.

[0111] Step S1: Construct a generation network G. The basic framework of the generation network G includes an encoder and a decoder.

[0112] Step S2: Construct a discriminant network D. The discriminant network D adopts a Markov discriminator and is composed of a series of consecutive convolutional modules.

[0113] Step S3: Conduct adversarial training on the generation network G and the discriminant network D. The trained generation network G is used to generate target style transfer images.

[0114] For the specific limitations on the steps implemented by the above computer device, reference can be made to Embodiment 1, and details will not be elaborated here.

[0115] It should be noted that each block in the block diagram and / or flowchart in the accompanying drawings of the present invention, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and machine instructions obtained.

[0116] The present invention has been described exemplarily above with reference to the accompanying drawings. Obviously, the specific implementation of the present invention is not limited by the above methods. As long as various non-substantive improvements are made by adopting the inventive concept and technical solutions of the present invention, or the inventive concept and technical solutions of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.

Claims

1. An image artistic style transfer algorithm, characterized in that: It includes the following steps: Step S1: Construct a generation network G. The basic framework of the generation network G includes an encoder and a decoder. The decoder uses a channel-spatial adaptive fusion module to fuse the upsampled features with the features of the downsampling module of the encoder. The channel-spatial adaptive fusion module includes a channel attention mechanism, a spatial attention mechanism, and an adaptive feature selection stage; The adaptive feature selection stage dynamically integrates and selects the features enhanced by the dual attention, then fuses based on the connections between different channel features in the adaptively selected features, and finally generates the output result; Step S2: Construct a discriminant network D. The discriminant network D uses a Markov discriminator and is composed of a series of consecutive convolutional modules; Step S3: Conduct adversarial training on the generation network G and the discriminant network D. The trained generation network G is used to generate target style transfer images.

2. The image art style transfer algorithm according to claim 1, wherein: The adaptive feature selection stage accepts the feature map processed by the channel attention mechanism and the spatial attention mechanism as input, establishes associations between channels through 1×1 convolution, and captures the interdependence between different channel features; subsequently, batch normalization is applied to stabilize the training process and accelerate convergence, while reducing internal covariate shift; finally, a threshold transfer unit is introduced to introduce a non-linear transformation to enhance the model's expressive power and filter out irrelevant information; The expression of the adaptive feature selection is: Among them, is the feature after adaptive selection of output channel a, represents the linear transformation matrix from the input to output channel a, b a is the bias term of output channel a, μ a is the channel mean vector of output channel a, β a is the learnable offset parameter of output channel a, is the feature variance of output channel a, γ a is the learnable scaling parameter of output channel a, and is the numerical stability constant.

3. An image artistic style transfer algorithm according to claim 2, characterized in that: The feature F after adaptive selection a Perform a linear transformation in the channel dimension through a 1×1 convolution operation to establish the connection between features of different channels. Subsequently, apply the batch normalization operation by calculating the mean μ of each channel B,c and variance Normalize the features to zero mean and unit variance, and then adjust the distribution through the learnable scaling parameter γ c and the offset parameter β c to effectively stabilize the training process and accelerate convergence. Finally, introduce a non-linear transformation through the threshold transfer unit to enhance the model's expressive ability and suppress negative value responses, filtering out irrelevant information to complete the final feature fusion; The linear transformation performed on the channel dimension is represented as W f ·F a +b f , where W f is the weight matrix, and b f is the bias term; the complete mathematical expression for the final feature fusion is: Among them, O c,h,w is the value of the final output feature at position (c, h, w), γ c is the scaling parameter of the c-th output channel, W f,c,i is the weight from input channel i to output channel c, F a,i,h,w is the value of the feature after adaptive selection at position (i, h, w), b f,c is the bias term of output channel c, μ B,c is the statistical mean of output channel c, is the statistical variance of output channel c, which is a numerical stability constant, β c is the offset parameter of output channel c.

4. The image art style transfer algorithm according to claim 1, wherein: The processing flow of the channel-spatial adaptive fusion module includes: First, the upsampled features and the skip connection features are concatenated in the channel dimension, and then the channel attention mechanism is introduced. First, the spatial dimension is compressed through adaptive global average pooling to extract channel-level features, and then the channel attention weights are generated through non-linear transformations of dimensionality reduction convolution, threshold transfer unit, and dimensionality increase convolution; further, the spatial attention mechanism is introduced. The spatial attention mechanism first calculates the average value and the maximum value of the feature map in the channel dimension, concatenates the average value and the maximum value, and generates a spatial attention map through convolution and non-linear activation units, that is, the feature map input to the adaptive feature selection stage.

5. An image artistic style transfer algorithm according to claim 1, characterized in that: An attention-free self-attention module is added after each downsampling module of the encoder to enhance feature representation; The processing flow of the parameterless self-attention module includes: first, calculating the spatial mean ω of the feature map X, and then calculating the squared difference (X - ω) from the mean for each feature position 2 , then calculating the normalized variance ρ of the entire feature map 2 , further calculating the preliminary attention scores based on the squared differences at each position, and the calculation formula is: where a is a scaling factor, ε is a small constant used to prevent division-by-zero errors, and b is an offset constant; finally, the preliminary attention score y is transformed into a weight coefficient between 0 and 1 through the non-linear activation function ρ(y), and these weights are applied to the original features, and the calculation formula is: X out = X × ρ(y), where X out is the weighted output feature map, and × is the element-wise multiplication.

6. The image art style transfer algorithm according to claim 1, characterized in that: The encoder is input through an initial module. The initial module is sequentially connected to a reflection padding layer, a convolutional layer, a normalization layer, and a threshold transfer unit; then two downsampling modules are connected. Each downsampling module consists of a convolutional layer, a normalization layer, and a threshold transfer unit, and an attention-free self-attention module is added after each downsampling module to enhance feature representation to form the encoder. The decoder includes two upsampling modules. Each upsampling module includes a transposed convolutional layer, a normalization layer, and a threshold transfer unit connected in sequence; then the channel-spatial adaptive fusion module is used to fuse the upsampled features with the features of the downsampling module of the encoder; finally, a reflection padding layer, a convolutional layer, and a hyperbolic tangent non-linear layer are connected to form the decoder.

7. An image artistic style transfer algorithm according to claim 1, characterized in that: In the convolutional module of the discriminative network D, the first convolutional module includes a convolutional layer and a ramp rectified linear unit; each of the intermediate convolutional modules includes a convolutional layer, an instance normalization layer, and a ramp rectified linear unit, and extracts multi-scale feature representations of the input image through layer-by-layer convolution and downsampling operations; the last convolutional module only includes a convolutional operation to generate the final discriminative feature map.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the steps of an image art style transfer algorithm as described in any one of claims 1-8.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of an image art style transfer algorithm as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Attention circulation adversarial network-based style migration system, method and device

    CN114493991A

  • Lightweight infrared and visible light image fusion method, system, equipment and medium

    CN116363034A

  • Sketch generation method and system based on cyclic generative adversarial network

    CN116503499A

  • Arbitrary style migration method and device fused with interactive attention mechanism

    CN118052706A

  • Three-dimensional seismic fault recognition model based on cross-channel interaction strategy and training method

    CN118644772A