High-fidelity image style migration method based on potential diffusion model
By performing image style transfer in the potential diffusion model, combining Transformer and multi-head self-attention mechanism, the high computational cost and detail loss of high-resolution image style transfer in the existing technology is solved, and efficient and flexible style transfer effect is achieved.
Patent Information
- Application Number
- CN202510634059.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-19
AI Technical Summary
The existing image style migration technology is cost-effective and slow when processing high-resolution and detailed images, making it difficult to meet real-time processing needs, and there are problems of detail retention and style consistency, making it difficult to achieve the best effect between style expression and content retention.
Using a method based on the potential diffusion model, by transferring the image style in the latent space, using the stepwise noise addition and denoising process of the diffusion model, combining the Transformer structure and the multi-head self-attention mechanism, we optimize image features, generate stylized images, and introduce content loss, style loss and denoising loss functions to balance style expression and content retention.
It improves the detail fidelity and style consistency of style transfer, reduces calculation costs, is suitable for real-time style transfer of high-resolution images, enhances the flexibility and controllability of style transfer, and can operate effectively in resource-limited environments.
Smart Images

Figure CN120510489A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and in particular relates to a high-fidelity image style transfer method based on a latent diffusion model. Background Art
[0002] Image style transfer technology is currently widely used in fields such as image editing, video processing, and artistic creation. Since Gatys et al. proposed using a deep neural network (DNN) combined with a VGG network for style transfer, this method has achieved a separation between image style and content through a combination of content loss and style loss, significantly improving the visual quality of style transfer. Subsequently, techniques such as generative adversarial networks (GANs) have also been introduced to style transfer, further enhancing the effectiveness and stability of style transfer. These methods rely on the powerful feature extraction capabilities of convolutional neural networks (CNNs). By progressively extracting and combining features from low-level to high-level layers, they significantly improve the overall style and detail consistency of the generated stylized images, gradually developing towards detailed and realistic style transfer effects.
[0003] Despite rapid developments in style transfer technology, existing methods still face several bottlenecks. Existing deep learning models typically have high computational costs and slow speeds when processing high-resolution and detail-rich images, making it difficult to meet real-time processing requirements. Furthermore, detail preservation in style transfer is a prominent issue, with existing methods lacking in finesse, contrast, and edge clarity, particularly in high-resolution scenes, resulting in blurring or structural distortion. Furthermore, traditional style transfer models struggle to balance style and content features, often requiring a trade-off between style expression and content preservation, a balance that often fails to achieve optimal results. Furthermore, due to the complexity of integrating multi-level features, the model is prone to vanishing gradients or instability during training, making it difficult to stabilize the quality of style transfer.
[0004] like Figure 2 As shown in the figure, global style transfer methods mainly focus on reconciling the characteristics of style and content on the whole, emphasizing a unified style effect. AdaIN proposed by Huang and Belongie achieves fast and efficient global style transfer by adjusting the mean and variance of the content image to match the statistical characteristics of the style image. This method is particularly suitable for real-time style transfer due to its simplicity and efficiency. WCT developed by Li et al. uses more complex statistical matching techniques (whitening and colorization) to globally adjust content features to match the covariance of style features. Although this method can maintain a high degree of content structural integrity, it has a high computational cost. StyTr2 is based on the transformer architecture and attention mechanism technology. By optimizing the interaction between style and content features, it not only improves the detail fidelity and visual effects, but also enhances the adaptability and flexibility of the model in dealing with complex styles and content.
[0005] In recent years, diffusion models, as an emerging generative model, have gradually attracted widespread attention from both academia and industry. Diffusion models generate images through a progressive process of adding noise (forward process) and removing noise (reverse process), demonstrating good generation stability and strong detail preservation capabilities. Compared to traditional generative adversarial networks (GANs), diffusion models can gradually restore image details through a gradual noise removal process, typically resulting in higher visual quality in generated images. They are particularly effective in preserving image structure and texture when processing detail-rich images.
[0006] Diffusion models were initially applied to image generation and have evolved into a variety of variants. Latent space-based diffusion models model images by mapping them from a high-dimensional pixel space to a low-dimensional latent space, significantly improving generation efficiency. By shifting the generation process to the latent space, latent diffusion models avoid the bottleneck of high-dimensional computation in pixel space, thereby improving computational efficiency and reducing resource consumption. However, existing diffusion models mostly focus on image generation tasks and have not been fully applied to image style transfer. Despite their ability to generate high-quality images, applying diffusion models to style transfer tasks while maintaining style consistency and detail fidelity remains a technical challenge. Summary of the Invention
[0007] The present invention addresses the technical problems encountered in the prior art. The present invention proposes a high-fidelity image style transfer method based on a latent diffusion model. By introducing a diffusion model for image style transfer, the method not only effectively improves the detail fidelity and style consistency of style transfer, but also avoids the style distortion and detail loss common in existing methods. Compared with traditional methods, the present invention offers the following innovations: by performing style transfer in latent space, the computational bottleneck of high-dimensional pixel space is avoided, effectively preserving image details while maintaining style consistency; by utilizing the gradual denoising process of the diffusion model, the present invention can gradually optimize the style and details of the image at each stage of style transfer, avoiding over-stylization or detail loss; and by modeling in latent space, the present invention significantly improves computational efficiency, making it suitable for real-time style transfer tasks for high-resolution images.
[0008] The technical solution adopted by the method of the present invention is as follows: a high-fidelity image style transfer method based on a latent diffusion model, comprising the following steps:
[0009] S1: The content image is input to the image encoder for feature extraction. The image encoder uses the VGG network to extract content image features, while the text encoder uses the CLIP model to extract semantic features of the style text, providing effective feature representation for subsequent image style transfer processing.
[0010] S2: The content image and style text description are each subjected to multi-level feature extraction through multiple convolutional layers, capturing both low-level features (such as texture and edges) and high-level features (such as semantic information) of the image. Feature maps of the content image and style text description are extracted through different convolutional layers of the image encoder.
[0011] S3: The extracted image features and style text features are fed into the latent diffusion model for further processing. The latent space optimizes the image features through a gradual process of denoising and adding noise, ultimately generating a stylized image.
[0012] S4: Transformer is introduced into the space. Image features and style text features are calculated through linear layers to obtain query vectors, key vectors, and value vectors. The relationship between image and text features is captured through a multi-head self-attention mechanism to generate style guidance signals and fuse them with image features.
[0013] S5: The style guidance signal is fused with the image features through a 1x1 convolution to generate stylized image features. This process ensures that the style guidance signal can accurately adjust the image features, thereby achieving style transfer while maintaining the fidelity of image details and structure.
[0014] S6: The stylized image features are fed into the denoising module, which gradually removes noise and restores image details, ultimately generating a stylized image. The denoising network's gradual denoising process gradually sharpens the details of the generated image, ultimately achieving the desired style transfer effect.
[0015] S7: Use the Adam optimizer to train the model. By dynamically adjusting the loss weight, we find the optimal balance between style expression and content preservation, thereby improving the detail fidelity and structural consistency of the generated images.
[0016] Preferably, in step S1, in order to pre-process the content image and style text and ensure their consistency in size and resolution, this study performs standardization on the input image to improve the adaptability and effect of the model in the style transfer process. The specific method is as follows:
[0017] S101: Scale the content image and style image uniformly to 512×512 pixels to preserve the global information of the image while ensuring the consistency of the resolution of images from different sources, providing rich input information for the model.
[0018] S102: Content image I c It is input to the image encoder for feature extraction. The image encoder uses the VGG network to extract the content image features F c :
[0019] F c =ε img (I c )#(1)
[0020] Among them, I c is the content image, ε img Represents an image encoder, which is used to extract image features.
[0021] Preferably, in step S2, in order to improve the accuracy and quality of image style transfer, the present invention captures deep information of the content image and style text through multi-level feature extraction.
[0022] S201: Content image I c Multi-level feature extraction is performed through the image encoder. A pre-trained convolutional neural network (such as VGG-19) is used to extract image features, mainly including low-level features (such as texture and edges) and high-level features (such as semantic information). The image encoder extracts the feature map of the image through multiple convolutional layers (such as Relu_3_1, Relu_4_1, Relu_5_1), which is specifically expressed by the following formula:
[0023]
[0024] in, Represents the image features extracted by different layers. V encoder (,) represents the image encoder's feature extraction process. Here, the image encoder extracts image features through different convolutional layers. Relu_3_1, Relu_4_1, and Relu_5_1 are feature extraction layers in the VGG network. Relu_3_1, Relu_4_1, and Relu_5_1 are used to extract low-level, mid-level, and high-level visual features, respectively. These layers extract feature maps at different levels of the image, which are used for subsequent style transfer.
[0025] S202: Style text description T s Extract the semantic features of style f through the text encoder (based on the CLIP model) s The text encoder converts the style text into feature representation in the semantic space, providing text feature support for subsequent style transfer. This process is expressed by the following formula:
[0026] f s =ε text (T s )#(5)
[0027] Where: T s is a text description of the style. text Represents a text encoder, which is used to extract semantic features of style text.
[0028] Preferably, in step S3, the image features and the style text features are processed by the encoder to obtain a low-dimensional representation, and these features can be further optimized in the latent space to achieve high-quality style transfer. In order to improve the efficiency and effect of image style transfer, the present invention designs a latent diffusion model to perform style transfer. By processing image features by adding noise and denoising in a low-dimensional latent space, high-dimensional calculations in the pixel space are avoided, significantly improving computational efficiency, while effectively retaining the details and style consistency of the image. The specific steps are as follows:
[0029] S301: In this step, the multi-layer image features and style text features extracted by the image encoder and text encoder are input into the latent diffusion model for further processing. Specifically, the image features and style features f s As input, it enters the latent space and, through a process of denoising and adding noise, gradually optimizes the image features. The main advantage of the latent diffusion model is that it performs calculations in a low-dimensional latent space, avoiding high-dimensional calculations in pixel space, significantly improving computational efficiency while effectively preserving image details and style consistency.
[0030] S302: The potential feature z0 of the image is transformed into noise z through a gradual noise addition process T This process simulates the evolution of the image during diffusion by adding noise to the latent representation. Specifically, at step t, the latent feature z t is added until it becomes pure noise z T The noise addition process can be expressed by the following formula:
[0031]
[0032] Among them, z t represents the potential representation of the t-th step, μ θ (z t-1 ,t) is the mean predicted by the denoising network, is the variance of the noise.
[0033] S303: In the denoising process, a denoising network (UNet) is used to represent the noise potential z T The denoising process gradually removes noise and restores the image's latent features, thereby generating a stylized image. This process can be expressed as follows:
[0034]
[0035] Among them, z t-1 represents the potential representation of the t-1th step, is the mean predicted by the denoising network, is μ used for denoising θ (z t ,t-1) noise variance.
[0036] Preferably, in step S4, to achieve cross-modal style consistency and effectively fuse image features with text style features, the present invention employs a Transformer architecture, which uses a self-attention mechanism to establish deep connections between image and text features. In this step, by introducing the Transformer, image style can be more accurately transferred in the latent space while ensuring consistency in the image's content structure and style.
[0037] S401: In order to improve cross-modal style consistency, the Transformer structure is introduced in the LDM diffusion process. Image features and style text features f s The query vector Q, key vector K, and value vector V are calculated through the linear layer. These vectors will serve as input for subsequent self-attention calculations. Specifically, the calculation of query, key, and value vectors is expressed by the following formula:
[0038]
[0039] K=Linear(f s )#(9)
[0040] V=Linear(f s )#(10)
[0041] in, is the image feature, f s It is the style text feature, and the linear transformation converts the image features and text features into corresponding query, keys and values.
[0042] S402: Through the multi-head self-attention mechanism, Transformer can capture the relationship between image and text features and generate style guidance signal f att Specifically, the multi-head self-attention mechanism first calculates the similarity between the query and the key, and then performs a weighted average of the values to generate the final style guidance signal. The calculation formula is as follows:
[0043]
[0044] here, It is a scaling factor for the dimensions of Q and K. The softmax function calculates the similarity between Q and K, thereby weighting V to obtain the final style guidance signal.
[0045] Preferably, in step S5, to enhance image detail and style consistency, the present invention uses a feedforward network (FFN) to process the denoised image features. The feedforward network performs nonlinear transformations on image features through fully connected layers, thereby enhancing the image's expressiveness. This operation helps preserve detail during the style transfer process and ensures that the image's style details better match the original content structure.
[0046] S501: In the Transformer, the denoised image features are processed by a feed-forward network (FFN). The feed-forward network contains two fully connected layers to further process and enhance the representation of the features:
[0047] FFN(x)=max(0,xW1+b1)W2+b2#(12)
[0048] S502: Style guidance signal f att After being generated by the aforementioned cross-modal attention mechanism, it is injected into the image features to transfer the style information to the image. The style guidance signal is fused with the image features through a 1x1 convolution. This step is expressed by the following formula:
[0049] The style guidance signal f generated by the Transformer module att Injected into the potential features of the image. The style guidance signal is fused with the image features through 1x1 convolution to generate stylized image features.
[0050]
[0051] Among them, W cs It is the learned weight matrix used to adjust the influence of style features on image features.
[0052] Preferably, in step S6, in order to accurately restore image details and ensure the quality of the final stylized image, the present invention inputs the stylized image features into a latent space denoising module for stepwise denoising to restore image details.
[0053] S601: Stylized image features It is input to the denoising module and the details of the image are restored by gradual denoising. The denoising process gradually removes the noise and finally generates the potential representation F of the stylized image. out .
[0054] The optimization formula of the denoising process is as follows:
[0055]
[0056] Finally, the denoised latent feature F out is passed to the decoder to generate a stylized image I cs
[0057] Preferably, in step S7, to ensure that the generated image achieves optimal results in terms of detail, style consistency, and content preservation, the present invention optimizes the style transfer process by introducing multiple loss functions. These loss functions include content loss, style loss, and denoising loss, which optimize the image from different dimensions to ensure that the image performance is optimal in all aspects. The specific steps are as follows:
[0058] 701: Content loss is used to ensure that the generated image remains semantically consistent with the original content image. Specifically, the content loss measures the content image I c and generate image I gen The deep feature difference between them. The formula of content loss is as follows:
[0059]
[0060] in, and Represents the content image I c and generate image I gen High-level features are computed by a deep feature extraction network (e.g., VGG network), where N is the dimension of the feature and ||·|| represents the Euclidean distance. In this way, the loss not only reflects pixel differences but also captures higher-level semantic structures.
[0061] 702: Style loss ensures that the generated image is consistent with the target style image in style, especially in terms of visual features such as color and texture. The style loss formula is as follows:
[0062]
[0063] in, and Represents the generated image and content image in the first i Feature histogram of the layer, N is the number of layers, ||·|| represents the L1 norm, which measures the difference in color or texture distribution between the two.
[0064] 703: Denoising loss directly affects the detail recovery and style consistency of the image. The present invention gradually denoises the noise in the latent space through a denoising network (such as UNet), so that the image can maintain style consistency and refine image details during the generation process. Denoising loss can be measured by calculating the difference between the latent features of the generated image and the latent features of the target stylized image. z0 is the latent feature of the denoised image, z gen To generate the potential features of the image, the denoising loss L d It can be expressed as:
[0065]
[0066] 704: Comprehensive loss function L t Combining content loss, style loss and denoising loss, it can be expressed as:
[0067] L t =λ c L c +λ s L s +λ d L d #(18)
[0068] Among them, λ c ,λ s ,λ d is the weight coefficient used to balance the influence of each loss function during the training process.
[0069] Compared with the prior art, the present invention has the following beneficial effects:
[0070] (1) Higher style consistency and detail preservation: By combining a cross-modal attention mechanism with a latent diffusion model, this paper can more accurately fuse the target style with the original image content, ensuring a more natural image style transfer while preserving more details. Compared to traditional methods, this paper avoids the common problems of detail loss and uneven style transitions during style transfer.
[0071] (2) Improved computational efficiency: By performing computations in latent space, latent space models significantly reduce the computational resources and time required for traditional pixel-space computations. Compared to other image style transfer methods, this method improves computational efficiency and is applicable to a wider range of hardware devices, especially in resource-constrained environments.
[0072] (3) Enhanced flexibility and controllability of style transfer: The cross-modal attention mechanism of this invention makes the style transfer process more flexible. Users can precisely control the style intensity of different regions through style text, thereby achieving fine-grained style transfer. This flexibility increases the diversity of image style transfer, making the generated images more suitable for different needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0074] Figure 1 Schematic diagram of the method flow of the present invention;
[0075] Figure 2 This is a diagram of the existing cross-modal style transfer framework;
[0076] Figure 3 is a potential space framework diagram of the present invention;
[0077] Figure 4 This is a graph showing the application results of the image style transfer of the present invention in artistic creation. DETAILED DESCRIPTION
[0078] The present invention will be further explained below in detail with reference to the accompanying drawings so that those skilled in the art can have a deeper understanding of the present invention and be able to implement it. However, the following reference examples are only used to explain the present invention and are not intended to limit the present invention.
[0079] like Figure 1 As shown in FIG, a high-fidelity image style transfer method based on a latent diffusion model includes the following steps:
[0080] S1: The content image is input to the image encoder for feature extraction. The image encoder uses the VGG network to extract content image features, while the text encoder uses the CLIP model to extract semantic features of the style text, providing effective feature representation for subsequent image style transfer processing.
[0081] S2: The content image and style text description are each subjected to multi-level feature extraction through multiple convolutional layers, capturing both low-level features (such as texture and edges) and high-level features (such as semantic information) of the image. Feature maps of the content image and style text description are extracted through different convolutional layers of the image encoder.
[0082] S3: The extracted image features and style text features are fed into the latent diffusion model for further processing. The latent space optimizes the image features through a gradual process of denoising and adding noise, ultimately generating a stylized image.
[0083] S4: Transformer is introduced into the space. Image features and style text features are calculated through linear layers to obtain query vectors, key vectors, and value vectors. The relationship between image and text features is captured through a multi-head self-attention mechanism to generate style guidance signals and fuse them with image features.
[0084] S5: The style guidance signal is fused with the image features through a 1x1 convolution to generate stylized image features. This process ensures that the style guidance signal can accurately adjust the image features, thereby achieving style transfer while maintaining the fidelity of image details and structure.
[0085] S6: The stylized image features are fed into the denoising module, which gradually removes noise and restores image details, ultimately generating a stylized image. The denoising network's gradual denoising process gradually sharpens the details of the generated image, ultimately achieving the desired style transfer effect.
[0086] S7: Use the Adam optimizer to train the model. By dynamically adjusting the loss weight, we find the optimal balance between style expression and content preservation, thereby improving the detail fidelity and structural consistency of the generated images.
[0087] In step S1, to pre-process the content image and style text and ensure their consistency in size and resolution, this study normalized the input image to improve the adaptability and effectiveness of the model during style transfer. The specific method is as follows:
[0088] S101: Scale the content image and style image uniformly to 512×512 pixels to preserve the global information of the image while ensuring the consistency of the resolution of images from different sources, providing rich input information for the model.
[0089] S102: Content image I c It is input to the image encoder for feature extraction. The image encoder uses the VGG network to extract the content image features F c :
[0090] F c =ε img (I c )#(1)
[0091] Among them, I c is the content image, ε img Represents an image encoder, which is used to extract image features.
[0092] In step S2, in order to improve the accuracy and quality of image style transfer, the present invention captures the deep information of the content image and style text through multi-level feature extraction.
[0093] S201: Content image I c Multi-level feature extraction is performed through the image encoder. A pre-trained convolutional neural network (such as VGG-19) is used to extract image features, mainly including low-level features (such as texture and edges) and high-level features (such as semantic information). The image encoder extracts the feature map of the image through multiple convolutional layers (such as Relu_3_1, Relu_4_1, Relu_5_1), which is specifically expressed by the following formula:
[0094]
[0095]
[0096]
[0097] in, Represents the image features extracted by different layers. V encoder (,) represents the image encoder's feature extraction process. Here, the image encoder extracts image features through different convolutional layers. Relu_3_1, Relu_4_1, and Relu_5_1 are feature extraction layers in the VGG network. Relu_3_1, Relu_4_1, and Relu_5_1 are used to extract low-level, mid-level, and high-level visual features, respectively. These layers extract feature maps at different levels of the image, which are used for subsequent style transfer.
[0098] S202: Style text description T s Extract the semantic features of style f through the text encoder (based on the CLIP model) s The text encoder converts the style text into feature representation in the semantic space, providing text feature support for subsequent style transfer. This process is expressed by the following formula:
[0099] f s =ε text (T s )#(5)
[0100] Where: T s is a text description of the style. text Represents a text encoder, which is used to extract semantic features of style text.
[0101] like Figure 3As shown, in step S3, the image features and style text features are processed by the encoder to obtain low-dimensional representations, and these features can be further optimized in the latent space to achieve high-quality style transfer. In order to improve the efficiency and effect of image style transfer, the present invention designs a latent diffusion model to perform style transfer. By processing image features by adding noise and denoising in a low-dimensional latent space, high-dimensional calculations in the pixel space are avoided, significantly improving computational efficiency while effectively preserving the details and style consistency of the image. The specific steps are as follows:
[0102] S301: In this step, the multi-layer image features and style text features extracted by the image encoder and text encoder are input into the latent diffusion model for further processing. Specifically, the image features and style features f s As input, it enters the latent space and, through a process of denoising and adding noise, gradually optimizes the image features. The main advantage of the latent diffusion model is that it performs calculations in a low-dimensional latent space, avoiding high-dimensional calculations in pixel space, significantly improving computational efficiency while effectively preserving image details and style consistency.
[0103] S302: The potential feature z0 of the image is transformed into noise z through a gradual noise addition process T This process simulates the evolution of the image during diffusion by adding noise to the latent representation. Specifically, at step t, the latent feature z t is added until it becomes pure noise z T The noise addition process can be expressed by the following formula:
[0104]
[0105] Among them, z t represents the potential representation of the t-th step, μ θ (z t-1 ,t) is the mean predicted by the denoising network, is the variance of the noise.
[0106] S303: In the denoising process, a denoising network (UNet) is used to represent the noise potential z T The denoising process gradually removes noise and restores the image's latent features, thereby generating a stylized image. This process can be expressed as follows:
[0107]
[0108] Among them, z t-1 represents the potential representation of the t-1th step, is the mean predicted by the denoising network, is μ used for denoising θ (z t ,t-1) noise variance.
[0109] In step S4, to achieve cross-modal style consistency and effectively fuse image and text style features, the present invention employs a Transformer architecture, which leverages a self-attention mechanism to establish deep connections between image and text features. Introducing the Transformer in this step allows for more precise image style transfer in the latent space while maintaining consistency in the image's content structure and style.
[0110] S401: In order to improve cross-modal style consistency, the Transformer structure is introduced in the LDM diffusion process. Image features and style text features f s The query vector Q, key vector K, and value vector V are calculated through the linear layer. These vectors will serve as input for subsequent self-attention calculations. Specifically, the calculation of query, key, and value vectors is expressed by the following formula:
[0111]
[0112] K=Linear(f s )#(9)
[0113] V=Linear(f s )#(10)
[0114] in, is the image feature, f s It is the style text feature, and the linear transformation converts the image features and text features into corresponding query, keys and values.
[0115] S402: Through the multi-head self-attention mechanism, Transformer can capture the relationship between image and text features and generate style guidance signal f att Specifically, the multi-head self-attention mechanism first calculates the similarity between the query and the key, and then performs a weighted average of the values to generate the final style guidance signal. The calculation formula is as follows:
[0116]
[0117] here, It is a scaling factor for the dimensions of Q and K. The softmax function calculates the similarity between Q and K, thereby weighting V to obtain the final style guidance signal.
[0118] In step S5, to enhance image detail and style consistency, the present invention uses a feedforward network (FFN) to process the denoised image features. The FFN performs nonlinear transformations on image features through fully connected layers, thereby enhancing the image's expressiveness. This operation helps preserve detail during the style transfer process and ensures that the image's style details better match the original content structure.
[0119] S501: In the Transformer, the denoised image features are processed by a feed-forward network (FFN). The feed-forward network contains two fully connected layers to further process and enhance the representation of the features:
[0120] FFN(x)=max(0,xW1+b1)W2+b2#(12)
[0121] S502: Style guidance signal f att After being generated by the aforementioned cross-modal attention mechanism, it is injected into the image features to transfer the style information to the image. The style guidance signal is fused with the image features through a 1x1 convolution. This step is expressed by the following formula:
[0122] The style guidance signal f generated by the Transformer module att Injected into the potential features of the image. The style guidance signal is fused with the image features through 1x1 convolution to generate stylized image features.
[0123]
[0124] Among them, W cs It is the learned weight matrix used to adjust the influence of style features on image features.
[0125] In step S6, in order to accurately restore image details and ensure the quality of the final stylized image, the present invention inputs the stylized image features into the latent space denoising module for step-by-step denoising to restore image details.
[0126] S601: Stylized image features It is input to the denoising module and the details of the image are restored by gradual denoising. The denoising process gradually removes the noise and finally generates the potential representation F of the stylized image. out .
[0127] The optimization formula of the denoising process is as follows:
[0128]
[0129] Finally, the denoised latent feature F out is passed to the decoder to generate a stylized image Ics
[0130] In step S7, to ensure that the generated image achieves optimal results in terms of detail, style consistency, and content preservation, the present invention optimizes the style transfer process by introducing multiple loss functions. These loss functions include content loss, style loss, and denoising loss, which optimize the image from different dimensions to ensure that the image performance is optimal in all aspects. The specific steps are as follows:
[0131] 701: Content loss is used to ensure that the generated image remains semantically consistent with the original content image. Specifically, the content loss measures the content image I c and generate image I gen The deep feature difference between them. The formula of content loss is as follows:
[0132]
[0133] in, and Represents the content image I c and generate image I gen High-level features are computed by a deep feature extraction network (e.g., VGG network), where N is the dimension of the feature and ||·|| represents the Euclidean distance. In this way, the loss not only reflects pixel differences but also captures higher-level semantic structures.
[0134] 702: Style loss ensures that the generated image is consistent with the target style image in style, especially in terms of visual features such as color and texture. The style loss formula is as follows:
[0135]
[0136] in, and Represents the generated image and content image in the first i Feature histogram of the layer, N is the number of layers, ||·|| represents the L1 norm, which measures the difference in color or texture distribution between the two.
[0137] 703: Denoising loss directly affects the detail recovery and style consistency of the image. The present invention gradually denoises the noise in the latent space through a denoising network (such as UNet), so that the image can maintain style consistency and refine image details during the generation process. Denoising loss can be measured by calculating the difference between the latent features of the generated image and the latent features of the target stylized image. z0 is the latent feature of the denoised image, z gen To generate the potential features of the image, the denoising loss L d It can be expressed as:
[0138]
[0139] 704: Comprehensive loss function L t Combining content loss, style loss and denoising loss, it can be expressed as:
[0140] L t =λ c L c +λ s L s +λ d L d #(18)
[0141] Among them, λ c ,λ s ,λ d is the weight coefficient used to balance the influence of each loss function during the training process.
[0142] In specific implementation, Example 1: Cross-dataset comparative test of image style transfer effect
[0143] This example compares the style transfer accuracy of our high-fidelity image style transfer method based on the Latent Diffusion Model (LDM) with traditional style transfer methods. Large-scale testing on multiple style transfer datasets demonstrates our advantages in improving the representation of style details and preserving content.
[0144] We randomly selected and combined 8,000 images from multiple standard style transfer datasets, including the VGG19 style image set, WikiArt, and COCO, as the test set. These images cover a variety of styles (such as oil painting, impressionism, and abstract art) and include different resolutions and scenes to ensure the diversity and representativeness of the dataset.
[0145] This experiment uses PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index), and Style Detail Retention (SDR) as evaluation metrics to calculate the difference between the generated image and the style image. To ensure fairness, the style transfer method based on the latent diffusion model of this invention is compared with the traditional style transfer method based on VGG, using the same evaluation criteria and parameter settings.
[0146] Each test image is input into the method of the present invention and divided into content and style feature streams after data preprocessing. The feature extraction module gradually optimizes the contrast, color expression and detail retention of the generated image through the latent diffusion model to ensure the optimal balance between style and detail.
[0147] Experimental results show that our style transfer method based on the latent diffusion model achieves a PSNR of 34.2dB on the test set, a 3.4dB improvement over the traditional method (30.8dB). The SSIM value is 0.91, compared to 0.83 for the traditional method, demonstrating significant improvements in structural consistency of the generated images. The style detail retention (SDR) is also improved by 15%, particularly in complex style images, where our method is able to better capture details and textures.
[0148] Example 2: Application of Text-based Image Style Transfer in Digital Art Creation
[0149] This example demonstrates the application of the present invention's high-fidelity image style transfer method based on cross-modal attention and latent diffusion models in digital art creation, particularly its performance in transferring artistic styles through text descriptions. With the advancement of digital art creation and text-driven technologies, artists can easily transform images into works with specific artistic styles through text descriptions. Using this method, artists can quickly generate artistically stylized images based on the style of text descriptions, expanding the boundaries of digital art creation.
[0150] To verify the effectiveness of the method in digital art creation, this example selected a dataset of text descriptions and image pairs from several classic art styles. These styles include Impressionism, Post-Impressionism, Cubism, and Modern Art. Natural landscapes, portraits, and cityscapes were also selected as content images, resulting in a total of 5,000 images for testing.
[0151] The text descriptions in the experiment cover the key features of various artistic styles. For example, "soft tones" and "wavy brushstrokes" represent the Impressionist style, "geometric shapes" and "symmetrical layout" represent the Cubist style, and so on. The content image is input together with the corresponding style text description to generate a stylized image.
[0152] In the experiment, we first extract features of the content image using an image encoder (such as VGG-19), followed by extracting semantic features of the style text using a CLIP-based text encoder. These image and text features are then fed into a latent diffusion model (LDM) for processing. LDM optimizes style transfer within a low-dimensional latent space through a gradual process of denoising and adding noise. This effectively avoids computations in a high-dimensional pixel space, significantly improving computational efficiency while preserving image detail and style consistency.
[0153] To ensure the fairness of the experiment, the method of the present invention was compared with the traditional style transfer method based on convolutional neural network. All methods used the same image preprocessing, parameter settings and style transfer objectives.
[0154] PSNR: During the generation of stylized images, the PSNR of the proposed method is 37.3dB, which is 4.1dB higher than that of the traditional method (33.2dB), indicating that the quality of the generated images is significantly improved. SSIM: The SSIM value of the proposed method is 0.91, which shows better structural consistency than the traditional method (0.83), and the generated images retain more content structure. Style Detail Retention (SDR): The style detail retention is improved by 20%, especially in complex style images (such as modern art and abstract art). The LDM method can better preserve style details and avoid style distortion.
[0155] like Figure 4 As shown, this example demonstrates the practical application of the high-fidelity image style transfer method based on cross-modal attention and latent diffusion models in digital art creation. Experimental results show that the method can accurately transfer artistic style through text descriptions, generating high-quality, stylized images with consistent style. In artistic creation, the method of the present invention has significant advantages over traditional methods, providing artists with an efficient and flexible creative tool, and is particularly suitable for artistic style transfer based on text style guidance.
[0156] Example 3: Application of Image Style Transfer in Social Media Content Creation
[0157] This example demonstrates the application of the present invention's high-fidelity image style transfer method based on a latent diffusion model and a cross-modal attention mechanism in social media content creation. With the continuous development of social media platforms, users' demand for stylized image and video content continues to increase. Many users want to quickly transform their images or videos into works that conform to a specific style. The present method provides an efficient and convenient solution to this need.
[0158] To validate the effectiveness of our method in social media content creation, experiments used a public image dataset from social media platforms. This dataset includes user-generated content (UGC) from various social platforms, encompassing a wide range of image types, such as selfies, landscapes, pets, and food. To enhance diversity, a test set of 10,000 images was selected, each accompanied by a relevant style description.
[0159] Style descriptions include "bright colors" and "dreamy lighting effects" for modern art, "black and white" and "vintage images" for retro, and "natural landscapes" and "fresh tones" for natural scenery. The experiment used different style descriptions to transfer image styles and generate stylized social media content.
[0160] In the experiment, each image is first passed through an image encoder (such as VGG-19) to extract features. Then, a text encoder based on the CLIP model extracts semantic features of the style text. Next, the image and text features are fed into a latent diffusion model for further processing. The image is gradually optimized through a process of denoising and adding noise, ultimately generating an image with the target style. This method can quickly and efficiently generate stylized images that meet user requirements.
[0161] To ensure the fairness of the experiment, this experiment compares the proposed method with the traditional style transfer method, using the same image preprocessing and style description to ensure the consistency of the test environment.
[0162] This example demonstrates the effectiveness of the image style transfer method based on the latent diffusion model (LDM) and cross-modal attention mechanism in social media content creation. Experimental results show that the method can quickly and with high quality convert content images into images that conform to the style of a specific text description, especially in the stylization requirements of user-generated content (UGC). The method not only improves image quality and style consistency, but also significantly enhances the user experience, making it a powerful tool for content creators on social media platforms.
[0163] The specific implementation scheme described above further illustrates in detail the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above is only a specific implementation scheme of the present invention and is not intended to limit the scope of the present invention. Any equivalent changes and modifications made by any technician in this field without departing from the concept and principle of the present invention should fall within the scope of protection of the present invention.
Claims
1. A high-fidelity image style transfer method based on a latent diffusion model, characterized in that: The steps include: Step S1: Take the content image as input and input it into the image encoder for feature extraction; Step S2: The content image and style text description are respectively subjected to multi-level feature extraction through multiple convolutional layers to capture low-level and high-level features of the image. Feature maps of the content image and style text description are extracted through different convolutional layers of the image encoder. Step S3: The extracted image features and style text features are input into the latent diffusion model for further processing. The latent space optimizes the image features through a process of gradual denoising and denoising, and finally generates a stylized image. Step S4: Transformer is introduced into the space. The image features and style text features are calculated through linear layers to obtain the query vector, key vector, and value vector. The relationship between image and text features is captured through a multi-head self-attention mechanism to generate a style guidance signal and fuse it with the image features. Step S5: The style guidance signal is fused with the image features through 1x1 convolution to generate stylized image features; Step S6: The stylized image features are input into the denoising module, which gradually denoises and restores the image details, and finally generates a stylized image; Step S7: Use the Adam optimizer to train the model. By dynamically adjusting the loss weight, we can find the optimal balance between style expression and content preservation, thereby improving the detail fidelity and structural consistency of the generated images.
2. The high-fidelity image style transfer method based on the latent diffusion model according to claim 1, characterized in that In step S1, in order to pre-process the content image and ensure its consistency in size and resolution, the input image is normalized to improve the adaptability and effect of the model in the style transfer process. The specific method is as follows: S101: Scale the content and style images to 512×512 pixels to preserve the global information of the image while ensuring that the resolution of images from different sources is consistent, providing rich input information for the model. S102: Content image I c It is input into the image encoder for feature extraction. The image encoder uses the VGG network to extract the content image feature F c : F c =e img (I c ) (1) Among them, I c is the content image, ε img Represents an image encoder, which is used to extract image features.
3. The high-fidelity image style transfer method based on the latent diffusion model according to claim 2, characterized in that In step S2, the specific method is as follows: S201: Content image I c Multi-level feature extraction is performed through the image encoder. The pre-trained convolutional neural network VGG-19 is used to extract image features, mainly including low-level features and high-level features. The image encoder extracts the feature map of the image through multiple convolutional layers, which is specifically expressed by the following formula: in, Represents image features extracted by different layers; V encoder (,) represents the process of image encoder extracting image features; S202: Style text description T s Through the text encoder, the semantic features of style are extracted based on the CLIP model. s , the text encoder converts the style text into feature representation in the semantic space, providing text feature support for subsequent style migration. This process is expressed by the following formula: f s =e text (T s ) (5) Where: T s is the style text description; ε text Represents a text encoder, which is used to extract semantic features of style text.
4. The high-fidelity image style transfer method based on the latent diffusion model according to claim 3, characterized in that In step S3, the image features and the style text features are processed by the encoder to obtain low-dimensional representations, and these features can be further optimized in the latent space to achieve high-quality style transfer. The specific steps are as follows: S301: In this step, the multi-layer image features and style text features extracted by the image encoder and text encoder are input into the latent diffusion model for further processing. Specifically, the image features and style features f s As input, it enters the latent space and gradually optimizes the image features through the process of adding and denoising. S302: The potential feature z0 of the image is transformed into noise z through a gradual noise addition process T ; This process simulates the evolution of the image during diffusion by adding noise to the latent representation. Specifically, at step t, the latent feature z t is added until it becomes pure noise z T ; The noise addition process is expressed by the following formula: Among them, z t represents the potential representation of the t-th step, μ θ (z t-1 ,t) is the mean predicted by the denoising network, is the variance of the noise; S303: In the denoising process, the denoising network UNet is used to represent the noise potential z T The potential feature z0 of the image is restored in ; the denoising process gradually removes the noise and restores the potential features of the image, thereby generating a stylized image; this process is expressed by the following formula: Among them, z t-1 represents the potential representation of the t-1th step, is the mean predicted by the denoising network, is μ used for denoising θ (z t ,t-1) noise variance.
5. The high-fidelity image style transfer method based on the latent diffusion model according to claim 4, characterized in that In step S4, in order to achieve cross-modal style consistency and effectively fuse image features with text style features, a Transformer structure is adopted. This structure uses the self-attention mechanism to establish a deep connection between image and text features. The specific steps are as follows: S401: In order to improve cross-modal style consistency, the Transformer structure is introduced in the LDM diffusion process; image features and style text features f s The query vector Q, key vector K, and value vector V are calculated through the linear layer respectively; these vectors will be used as input for subsequent self-attention calculations. Specifically, the calculation of query, key, and value vectors is expressed by the following formula: K=Linear(f s ) (9) V=Linear(f s ) 10) in, is the image feature, f s It is the style text feature. The linear transformation converts the image features and text features into corresponding queries, keys and values; S402: Through the multi-head self-attention mechanism, Transformer captures the relationship between image and text features and generates a style guidance signal f att Specifically, the multi-head self-attention mechanism first calculates the similarity between the query and the key, and then performs a weighted average of the values to generate the final style guidance signal; the calculation formula is as follows: in, It is a scaling factor for the dimensions of Q and K. The softmax function calculates the similarity between Q and K, thereby weighting V to obtain the final style guidance signal.
6. The high-fidelity image style transfer method based on the latent diffusion model according to claim 5, characterized in that In step S5, in order to enhance the detail expression and style consistency of the image, the denoised image features are processed using a feedforward network (FFN). The specific steps are as follows: S501: In the Transformer, the denoised image features are processed by a feed-forward network (FFN). The feed-forward network contains two fully connected layers to further process and enhance the representation of the features: FFN(x)=max(0,xW1+b1)W2+b2(12) S502: Style guidance signal f att After being generated by the aforementioned cross-modal attention mechanism, it will be injected into the image features to transfer the style information to the image. The style guidance signal is fused with the image features through 1x1 convolution; This step is expressed by the following formula: The style guidance signal f generated by the Transformer module att is injected into the latent features of the image; The style guidance signal and image features are fused through 1x1 convolution to generate stylized image features. Among them, W cs It is the learned weight matrix used to adjust the influence of style features on image features.
7. The high-fidelity image style transfer method based on the latent diffusion model according to claim 6, characterized in that In step S6, the specific steps are as follows: S601: Stylized image features It is input into the denoising module and the details of the image are restored by gradual denoising; the denoising process gradually removes the noise and finally generates the potential representation F of the stylized image. out ; The optimization formula of the denoising process is as follows: Finally, the denoised latent feature F out is passed to the decoder to generate a stylized image I cs .
8. The high-fidelity image style transfer method based on the latent diffusion model according to claim 7, characterized in that: In step S7, in order to ensure that the generated image can achieve the best effect in terms of details, style consistency, and content preservation, the style transfer process is optimized by introducing multiple loss functions. The specific steps are as follows: 701: The content loss is used to ensure that the generated image remains semantically consistent with the original content image; specifically, the content loss measures the content image I c and generate image I gen The deep feature difference between them; the formula of content loss is as follows: in, and Represents the content image I c and generate image I gen High-level features calculated by a deep feature extraction network VGG network, N is the dimension of the feature, ||·|| represents the Euclidean distance; 702: Style loss ensures that the generated image is consistent with the target style image in style, especially in terms of visual features such as color and texture; the style loss formula is as follows: in, and Represents the generated image and content image in the first i Feature histogram of the layer, N is the number of layers, ||·|| represents the L1 norm, which measures the difference in color or texture distribution between the two; 703: Denoising loss directly affects the detail recovery and style consistency of the image. The noise in the latent space is gradually denoised by the denoising network UNet, so that the image can maintain style consistency and refine image details during the generation process. The denoising loss is measured by calculating the difference between the latent features of the generated image and the latent features of the target stylized image. z0 is the latent feature of the denoised image, z gen To generate the potential features of the image, the denoising loss L d Expressed as: 704: Comprehensive loss function L t Combining content loss, style loss and denoising loss, it is expressed as: L t =λ c L c +λ s L s +λ d L d #(18) Among them, λ c ,λ s ,λ d is the weight coefficient used to balance the influence of each loss function during the training process.
Citation Information
Patent Citations
Image style migration method based on diffusion model
CN117689532A
Decoupling type image style migration method based on diffusion model
CN119599861A
Image generation method based on style feature injection
CN119648517A
Cited By
Bidirectional fusion image style migration method and system based on content and style decoupling
CN120765451A
Diffusion style migration method and system based on subsurface distribution reanchoring dynamic injection
CN121481829A
Generative style watermark protection method and system based on submerged space collaborative embedding
CN121582050A
A method and system for generating a style watermark based on latent space collaborative embedding
CN121582050B