Image style conversion method based on content and style analyzer
By introducing content and style parsers into the image style conversion method, the problems of resource density and interweaving of style and content in the prior art are solved, and high-quality style conversion and content retention are achieved.
Patent Information
- Application Number
- CN202510005208.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-02
AI Technical Summary
The existing image style conversion methods rely on diffusion models, training resources are intensive and style and content are intertwined in feature extraction and sampling, resulting in the inclusion of unidirectional deviation and the inclusion of irrelevant content in stylized images.
The image style conversion method based on content and style parser is adopted, and the image content features are learned through the stable diffusion model, the style features are extracted using the visual deformer, and the content and style information are accurately injected into different blocks of U-Net through the content controller and style parser to achieve the decoupling of style and content.
It effectively avoids the problem of resource-intensiveness, reduces the one-way deviation between style and content, improves image quality and style similarity, and ensures the integrity of content and the accurate expression of style.
Smart Images

Figure CN119941491A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence computer vision technology, and in particular to an image style conversion method based on content and style parser. Background Art
[0002] In the field of image style transfer, the current main method is to use diffusion models. Current image style transfer mostly relies on diffusion models. Although they have excellent generative power, they require massive resources to train. For this reason, the industry has gradually focused on fine-tuning the cross-attention mechanism and its weights to reduce costs. However, in this process, style and content are intertwined in feature extraction and sampling, often leading to one-way bias results. Diffusion model training is resource-intensive. When fine-tuning the cross-attention mechanism and weights, style and content are entangled in feature extraction and sampling, resulting in one-way bias results. For example, the diffusion model method is easy to mix irrelevant content into the stylized image and deviate from the target style. Traditional methods may find it difficult to capture complex styles or damage the integrity of the content, and the poor style-content coupling leads to poor stylization effects.
[0003] It should be noted that the information disclosed in the above background technology section is only used for understanding the background of the present application, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the invention
[0004] The main purpose of the present invention is to overcome the defects existing in the above-mentioned background technology and provide an image style conversion method based on content and style parser.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for image style transfer based on content and style parser, comprising the following steps:
[0007] S1. Building a basic model: Using the stable diffusion model as the basic architecture, the learning and encoding of image content features are achieved through the process of adding noise and denoising;
[0008] S2. Build a style parser and extract style features: Use the visual deformer ViT as a style parser to extract style features from the reference image and generate style embedding through a multi-layer perceptron;
[0009] S3. Build a content parser and extract content features: Build a content parser, process content images hierarchically, generate multiple levels of latent representations as content embeddings, and integrate content embeddings into the model through residual addition to stabilize the content architecture and preserve content features during style transfer;
[0010] S4. Build content controller and inject content: Process the content image through the content controller to capture the basic structure and minimize the interference of style features;
[0011] S5. Injecting style and content: Injecting style embedding into the upsampling block of the model to enhance style expression, while injecting content embedding into the downsampling block of the model to ensure content preservation;
[0012] S6. Reasoning: In the reasoning stage, the constructed content parser and style parser are used to decouple style from content, capture and reproduce complex styles while maintaining the integrity of the content.
[0013] Furthermore, the processing of the style parser specifically includes:
[0014] The image is divided into blocks and linearly embedded using the visual deformer ViT;
[0015] Calculate query, key and value matrices through self-attention mechanism, update input feature representation and generate style features;
[0016] By normalizing the output of the self-attention mechanism, the training process is stabilized and the model generalization ability is improved;
[0017] Style embedding is further extracted from the normalized features through a multi-layer perceptron (MLP).
[0018] Furthermore, the training process of the style parser specifically includes:
[0019] During training, we use the Gram loss to ensure style alignment by comparing the Gram matrices of features of the generated and target images.
[0020] Keep the visual features of the generated image and the target image aligned through perceptual loss;
[0021] The original style loss, perceptual loss and Gram loss are weighted summed to get the total style loss;
[0022] Image conditions are randomly dropped during training, enabling no classifier bootstrapping during inference.
[0023] Furthermore, the style parser includes the following style injection strategies:
[0024] Leveraging the hierarchical nature of convolutional neural networks, we identify and specifically use upsampling blocks to capture and inject style information;
[0025] A criss-cross attention mechanism is implemented to inject style information into the upsampling block guided by image conditions, thereby enhancing style expression.
[0026] Furthermore, the operations of the content controller specifically include:
[0027] A tiling control strategy is adopted to preserve content information without affecting the style transfer process;
[0028] Use ControlNet to manage spatial information within the U-Net architecture, capture the basic structure of the content image, and minimize the interference of style features;
[0029] The content fusion encoder integrates the content to generate multiple levels of latent representations to form content embedding;
[0030] Integrate content embedding into U-Net via residual addition;
[0031] Through the ControlNet mechanism, content embedding is accurately injected into the upsampling block of the diffusion model to achieve the separation and conversion of content and style.
[0032] Furthermore, the training process of the content analyzer specifically includes:
[0033] Keep the parameters of the pre-trained diffusion model unchanged during the training phase;
[0034] Trained using image-text pairs to achieve effective correspondence between content and text descriptions;
[0035] Use mean squared error loss to measure and minimize the difference between the model's predicted noise and the actual noise;
[0036] Use adversarial loss to train the discriminator to enhance its ability to distinguish between real and generated images;
[0037] The content loss is formed by weighted combination of mean squared error loss and adversarial loss to balance the contribution of the two losses to content preservation quality.
[0038] Furthermore, the content embedding derived from the content parser is injected into the downsampling block of the model.
[0039] Further, in step S6, the reasoning includes a reversal operation for content retention, specifically including:
[0040] During the sampling process, the input image is encoded into the latent space, and the ReNoise technique reverses the sampling process to generate the latent noise representation directly from the real image, thus eliminating the need for additional noise injection;
[0041] The noise is updated in an iterative manner, enhancing the approximation of the expected position in the forward diffusion progression by averaging the predictions.
[0042] Furthermore, step S6 specifically includes:
[0043] Using the disentangled property of the criss-cross attention mechanism, the image conditional weights for content and style conditions are independently adjusted during the inference phase.
[0044] Construct a latent representation of the final image by weightedly combining the attention results of content and style conditions;
[0045] Set the text condition weight factor to zero;
[0046] When the weight factors of both content and style conditions are zero, the model reverts to the original text-to-image diffusion model.
[0047] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the image style conversion method based on content and style parser is implemented.
[0048] The present invention has the following beneficial effects:
[0049] The present invention proposes an image style transfer method based on content and style parser, which can accurately decouple style and content, effectively avoid the dependence of existing diffusion models on massive resources during training, and solve the one-way deviation problem caused by the interweaving and entanglement of style and content in feature extraction and sampling. Through the specially constructed style and content parser, the present invention can accurately inject style and content information into different blocks of U-Net, so that the style transfer process is flexible and diverse without damaging the integrity of the content. For example, in qualitative comparison, compared with methods such as CSGO, the present invention can fully retain the structure and details of buildings, characters, etc. while accurately migrating the style. In addition, the present invention comprehensively improves the image quality and style similarity through multi-loss collaborative training of style parser and multi-loss fine optimization of content parser, and performs excellently in quantitative evaluation of CSD and CAS indicators, demonstrating the high efficiency of style and content control. User research also shows that the present invention leads in the choice of style and content scoring and overall quality, and the generated images achieve a delicate balance between quality, content similarity and style similarity, significantly optimizing the user's visual experience.
[0050] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 The figure is an overall flow chart of an image style conversion method based on a content and style parser according to an embodiment of the present invention.
[0052] Figure 2 The figure is an algorithm architecture diagram of an image style conversion method based on a content and style parser according to an embodiment of the present invention.
[0053] Figure 3The stylized result images generated by the embodiments of the present invention under different style conditions. The first image is used as the content image, and the illustration represents the style image.
[0054] Figure 4 This is a qualitative comparison between the embodiment of the present invention and other advanced methods. Other methods produce unilateral expression bias: the diffusion model-based methods ((b)-(h)) tend to express style, while the traditional style transfer methods ((i)-(k)) tend to express content.
[0055] Figure 5 These are the results of extensive qualitative comparison experiments between the embodiments of the present invention and the diffusion model-based method (ah column) and the traditional style transfer method (ik column).
[0056] Figure 6 The result is the percentage of users’ preference for the overall effect.
[0057] Figure 7 Ablation experiments in which the content parser of an embodiment of the present invention is embedded in different U-Net blocks.
[0058] Figure 8 Ablation experiments in which the style parser of an embodiment of the present invention is embedded in different U-Net blocks.
[0059] Fig. 9 This is the influence of the style analyzer strength on the result according to the embodiment of the present invention.
[0060] Fig.10 Ablation experiments of (a) stylized representation for style detail control and (b) ReNoise inversion for content enhancement control according to embodiments of the present invention.
[0061] Fig.11 This is an ablation study of an embodiment of the present invention using Canny edge detection as a control condition to replace a content controller.
[0062] Fig.12 This is the influence of the content controller strength on the result according to the embodiment of the present invention.
[0063] Fig.13 This is the influence of the content parser strength on the results of the embodiment of the present invention. DETAILED DESCRIPTION
[0064] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope and application of the present invention.
[0065] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0066] See also Figure 1 and Figure 2 The embodiment of the present invention provides an image style conversion method based on content and style parser, comprising the following steps:
[0067] Step S1. Building a basic model: Using a stable diffusion model as the basic architecture, learning and encoding image content features through the process of adding noise and denoising;
[0068] In a preferred embodiment, the diffusion model uses U-Net to implement the denoising process and is trained through mean square error loss.
[0069] Step S2. Build a style parser and extract style features: Use the visual deformer ViT as a style parser, extract style features from the reference image, and generate style embedding through a multi-layer perceptron.
[0070] In a preferred embodiment, the processing of the style parser specifically includes: performing block division and linear embedding processing on the image through a visual deformer ViT; calculating query, key and value matrices through a self-attention mechanism, updating the input feature representation, and generating style features; stabilizing the training process and improving the generalization ability of the model by normalizing the output of the self-attention mechanism; and further extracting style embedding from the normalized features through a multi-layer perceptron MLP.
[0071] In a preferred embodiment, the training process of the style parser specifically includes: during training, using Gram loss by comparing the Gram matrices of the features of the generated image and the target image to ensure style alignment; maintaining the visual feature alignment of the generated image and the target image through perceptual loss to enhance model performance; the total style loss is a weighted sum of the original style loss, perceptual loss and Gram loss to achieve accurate expression of style and content; randomly discarding image conditions during training to achieve no classifier guidance in reasoning and improve the generalization ability of the model.
[0072] In a preferred embodiment, the style parser includes the following style injection strategy: utilizing the hierarchical characteristics of convolutional neural networks, identifying and exclusively using upsampling blocks to capture and inject style information, to distinguish it from the content information captured by downsampling blocks; implementing a cross-attention mechanism, guided by image conditions, to inject style information into upsampling blocks, thereby enhancing style expression; through a targeted injection strategy, avoiding injecting style information into all network blocks, preventing confusion between style and content features, while ensuring the effective transmission and expression of style information.
[0073] Step S3. Build a content parser and extract content features: Build a content parser, process the content image hierarchically, generate multiple levels of latent representation as content embedding, and integrate the content embedding into the model through residual addition to stabilize the content architecture and ensure that the content features are preserved during the style transfer process.
[0074] In a preferred embodiment, the training process of the content parser specifically includes: keeping the parameters of the pre-trained diffusion model unchanged during the training phase; using image-text pairs for training to achieve effective correspondence between content and text descriptions; using mean square error loss to measure and minimize the difference between model prediction noise and actual noise to improve the quality of generated images; using adversarial loss to train the discriminator to enhance its ability to distinguish between real and generated images, thereby improving the authenticity of generated images; forming content loss by weighted combination of mean square error loss and adversarial loss to balance the contribution of the two losses to content retention quality.
[0075] In a preferred embodiment, the content embedding derived from the content parser is injected into the downsampling block of the model, ensuring that low-level aspects of the content image, such as shapes and edges, are effectively captured and preserved.
[0076] Step S4. Build content controller and inject content: Process the content image through the content controller to capture the basic structure and minimize the interference of style features.
[0077] In a preferred embodiment, the operation of the content controller specifically includes: adopting a tiling control strategy to retain content information without affecting the style transfer process; using ControlNet to manage spatial information within the U-Net architecture to effectively capture the basic structure of the content image and minimize the interference of style features; integrating content through a content fusion encoder to generate multiple levels of potential representations to form content embedding; integrating content embedding into U-Net through residual addition to ensure the stability and integrity of content features during the style transfer process; and accurately injecting content embedding into the upsampling block of the diffusion model through the ControlNet mechanism to achieve effective separation and conversion of content and style.
[0078] Step S5. Injecting style and content: injecting style embedding into the upsampling block of the model to enhance style expression, while injecting content embedding into the downsampling block of the model to ensure content preservation;
[0079] Step S6. Reasoning: In the reasoning stage, the constructed content parser and style parser are used to decouple style from content, allowing complex styles to be captured and reproduced while maintaining the integrity of the content.
[0080] In a preferred embodiment, in step S6, the reasoning includes an inversion operation for content preservation, specifically including: during the sampling process, encoding the input image into a latent space, and reversing the sampling process through the ReNoise technology to directly generate a latent noise representation from the real image, thereby eliminating the need for additional noise injection; using an iterative method to update the noise, enhancing the approximation of the expected position in the forward diffusion progress through average prediction, improving the accuracy of content reconstruction, and minimizing the computational cost; improving the quality of image content reconstruction through an iterative process while maintaining the distribution of the noise latent representation.
[0081] In a preferred embodiment, during the inference process, the separation characteristics of the cross-attention mechanism are utilized to independently adjust the image condition weights of the content and style conditions in the inference stage; the potential representation of the final image is constructed by weighted combination of the attention results of the content and style conditions; the text condition weight factor is set to zero; when the weight factors of the content and style conditions are both zero, the model reverts to the original text-to-image diffusion model.
[0082] This paper proposes an innovative image style transfer method, which achieves accurate decoupling of style and content by constructing content and style parsers. This method uses a stable diffusion model as the basic architecture and learns image content features through the process of denoising and denoising. The style parser is based on the visual deformer ViT to extract style features and generate style embeddings; the content parser processes the content image in layers, generates content embeddings and integrates them into the model through residual addition. In the inference stage, the style embedding is injected into the upsampling block and the content embedding is injected into the downsampling block to ensure style expression and content preservation. This paper effectively overcomes the problem of intensive training resources of traditional diffusion models, avoids the entanglement of style and content in the conversion process, reduces the one-way deviation of style and content, and improves image quality and style similarity. This paper overcomes the balance dilemma of accurate style expression and proper content preservation in traditional methods, and improves the decoupling degree of style-content. User studies show that this paper is overwhelmingly favored in visual experience, and the generated images achieve a delicate balance between quality, content and style similarity, which significantly optimizes the user visual experience.
[0083] The following further describes an algorithm example and experimental verification of a specific embodiment of the present invention.
[0084] Basic Model
[0085] Stable diffusion consists of two main stages: a noise addition stage, in which Gaussian noise ∈ is gradually introduced into the starting data x0 through a Markov chain; and a denoising stage.
[0086] In the denoising stage, the system uses a Gaussian distribution N(0,1) to extract the noise x t Generate samples. This is done through a trainable denoising model ∈ θ (x t ,t,c), the model is parameterized by θ.
[0087] Denoising model ∈ θ (·) Based on the U-Net architecture, and optimized using the mean squared error loss function. This loss function is derived from a simplified version of the variational bound and is expressed as:
[0088]
[0089] Where c represents an optional condition variable.
[0090] Style Parser
[0091] In the present invention, the style parser is very important in extracting and injecting style features. To this end, the present invention makes improvements in three aspects, including style representation, style injection block and training scheme.
[0092] Stylized Representation
[0093] The visual deformer ViT is used to extract style features. The image is divided into blocks and linearly embedded:
[0094] X={x1,x2,…,x N}.
[0095] These embeddings are then processed through ViT layers. In each attention layer, the self-attention weights are calculated as follows:
[0096]
[0097] where Q, K, V are the query, key and value matrices computed as follows:
[0098] Q=XW Q ,K=XW K ,V=XW V .
[0099] Each ViT layer updates the input feature representation to produce the final style feature:
[0100] Z=Norm(X+Attn(Q,K,V)),
[0101] Norm represents the normalization operation, which is used to stabilize the training process and improve the generalization ability of the model.
[0102] The style embedding is extracted as the output of a specific layer:
[0103] E s =MLP(Norm(Z)),
[0104] Where E s represents style embedding and MLP represents multi-layer perceptron, which is used to further extract style features from the normalized feature Z.
[0105] Training Program
[0106] During the training process of the style parser, specific weights are assigned to different loss functions to balance their contributions to the total loss. The training objectives of the style parser include:
[0107]
[0108] To enhance model performance, a perceptual loss is introduced to maintain perceptual fidelity and align the visual features of generated and target images:
[0109] L p =|VGG(g)-VGG(t)| 2 ,
[0110] Where g and t represent the generated image and target image respectively.
[0111] We also use the Gram loss by comparing the Gram matrices of the features of the generated and target images to ensure style alignment:
[0112] L g =|Gram(g)-Gram(t)| 2 .
[0113] The style loss is a weighted sum of all components:
[0114] L style =α s ·L s +β s ·L p +γ s ·L g ,
[0115] The weight α s =0.2,β s =0.4, and γ s =0.4, respectively.
[0116] Additionally, image conditions are randomly dropped during training to enable classifier-free bootstrapping during inference:
[0117]
[0118] If the image condition is dropped, the image embedding will simply be zeroed.
[0119] Style injection block
[0120] It is generally understood that in convolutional neural networks, lower convolutional layers learn low-level aspects such as shape and color, while deeper convolutional layers focus on semantic information.
[0121] Like the text condition, the image condition is generated by injecting guidance through a cross-attention layer. In addition, the inventors observed in experiments that multiple blocks capture style and content information in different ways, such as Figure 8 shown.
[0122] Specifically, the downsampling block tends to capture content information, the upsampling block can capture style information such as color, material, texture, etc., while the intermediate block has no single preference.
[0123] If all blocks are injected, the result will be the generation of wrong content and style features, which is common in this type of methods.
[0124] Therefore, the style information is injected into the upsampling block, such as Figure 2 shown.
[0125] This targeted injection reduces the need for extensive parameter adjustments and enhances stylistic expression.
[0126] Content Parser
[0127] Previous methods usually directly process content image embeddings for style transfer, but may retain inherent style information and affect the performance of complex tasks. Therefore, a content parser is constructed to extract specific structural information for control.
[0128] Content Controller
[0129] Inspired by direct processing methods, tiled control is adopted to ensure that content information is preserved without affecting style. The model processes the content image and effectively captures its basic structure while minimizing style features. Leveraging ControlNet's ability to manage spatial information within the U-Net architecture, the content fusion encoder integrates content and produces multiple levels of latent representations for content embedding f c : in Involving intermediate sample blocks, arrive Contains downsampling blocks, L represents the total number of layers. The content is embedded into f by residual addition. c Integration into U-Net:
[0130]
[0131] f 0 represents the latent features of the intermediate block of the content controller U-Net, and f 1 to f L represents the representation in the upsampling block. Figure 2 As shown, the final injection method follows the standard ControlNet mechanism and is injected into the upsampling block of the diffusion model.
[0132] Training Program
[0133] During the training phase, the present invention only focuses on optimizing the parser without changing the parameters of the pre-trained diffusion model. The training process uses image-text pairs and integrates the following losses.
[0134] In order to measure the difference between the predicted noise and the actual noise, the mean square error loss is introduced, and the formula is as follows:
[0135]
[0136] where p i is the model’s prediction of the noise, a i is the actual noise and n is the number of samples. This loss helps the model predict the noise more accurately, thus improving the quality of the generated images.
[0137] To improve the realism of generated images, an adversarial loss is used as follows:
[0138] L a =-[D(g)log(D(g))+(1-D(g))log(1-D(g))],
[0139] Where D is the probability that the discriminator predicts that the generated image g is real. The adversarial loss trains the discriminator to distinguish between real and generated images, thereby improving the quality of content preservation.
[0140] The content loss is a weighted combination of:
[0141] L content =β c ·L mse +γ c ·L a ,
[0142] where β c = 0.6 and γ c= 0.4 are the weights of the mean squared error loss and the adversarial loss, respectively. These weights are chosen to balance the contribution of each loss to the content loss.
[0143] Content injection block
[0144] Experiments and observations show that the downsampling block is mainly useful for capturing and preserving low-level aspects of the content image (such as shapes and edges, see Figure 7 ) is very important.
[0145] They contribute less to style features but are crucial for content preservation. Therefore, the content embeddings derived from the content parser are injected into the downsampling block (see Figure 2 ).
[0146] Enhancement strategy
[0147] Reverse operation for content preservation
[0148] During the sampling process, the input image is usually encoded into the latent space and noise is added. However, the presence of additional noise may cause content drift, thus deviating from the intended target. ReNoise proposes to reverse the sampling process and generate latent noise representation directly from the real image, eliminating the need for additional noise injection. Although InstantStyle proposes that image inversion may ignore subtle style differences in the image, this limitation does not hinder the content protection task of the present invention. The embodiment of the present invention does not rely solely on a single inversion process, but focuses on average prediction to enhance the approximation of the expected position in the forward diffusion progress, allowing the model to improve the accuracy of content reconstruction while minimizing computational costs. The basic formula for ReNoise inversion is as follows:
[0149]
[0150] This formula updates the noise ∈ t . represents the gradient relative to the noise, ρ t is a scaling factor, z t-1 is the latent representation of the previous step, φ t and ψ t are model parameters.
[0151] This iterative process helps improve the quality of image content reconstruction while maintaining the distribution of the noisy latent representation.
[0152] reasoning
[0153] Since the criss-cross attention is decoupled, the image-conditional weights for content and style conditions can be adjusted independently during inference:
[0154] Z'=λ tAttn(Q,K,V)+λ c Attn(Q,K c ,V c )+λ s Attn(Q,K s ,V s ),
[0155] where λ c and λ s are the weight factors for content and style conditions, respectively. Z’ is the latent representation that combines the construction condition information to produce the final image.
[0156] Text condition weight factor λ t is set to 0.
[0157] On the contrary, if λ c = 0 and λ s = 0, the model will revert to the original text-to-image diffusion model.
[0158] Dataset
[0159] Previous stylization methods often use the Laion-Aesthetics dataset to train style encoders. Laion-Aesthetics contains 92.3% natural images, and this dataset with high aesthetic scores is more suitable for training content parsers. Therefore, an additional 100k art description dataset "Palette" was constructed. The 100,000 art paintings in the "Palette" are mainly selected from WikiArt and LaionArt. Blip-2 is used to query the style details of each work of art to obtain a detailed text description of the painting.
[0160] In the embodiment of the present invention, a style parser is designed, and a visual transformer is used as a tool to extract style features, and style embedding is obtained through multi-step processing. In the training phase, the mean square error loss, perceptual loss and Gram loss are combined, summed according to specific weights, and image conditions are randomly discarded to achieve the effect of no classifier guidance. In view of the differences in the capture characteristics of style and content of different U-Net blocks, the upper block is accurately selected to inject style information, which not only improves the style expression, but also avoids tedious parameter adjustment.
[0161] The content controller generates content embedding through multi-layer integration and injects it into the U-Net block by residual addition to stabilize the content architecture. The model is optimized using the weighted sum of mean square error loss and adversarial loss during training. The Palette dataset is constructed, which contains 100,000 art paintings and text interpretations, and its high proportion of natural images and high-quality aesthetic scores enable content parser training.
[0162] An enhancement strategy is proposed, an inversion procedure is introduced during sampling, noise is iteratively updated according to the denoising formula, the accuracy of content reconstruction is improved, and the noise distribution characteristics are stabilized. The XL version of the stable diffusion model is used as the base model, and the pre-trained visual transformer is used as the image encoder. The image resolution is unified, the learning rate and other training parameters are set, and the denoising steps and memory usage are flexibly allocated according to actual needs during the inference stage to improve processing efficiency.
[0163] Other embodiments
[0164] The weight of the loss function in the style parser can be fine-tuned, such as moderately increasing the weight of the perceptual loss and slightly reducing the weight of the Gram loss, so that the generated image can retain the style characteristics while being closer to the style strength expected by the user. Although the overall effect is slightly inferior to the best solution, the style control is more flexible and resource consumption is slightly increased.
[0165] In scenarios with specific styles or low content complexity, some adjacent U-Net blocks can share style or content injection points to simplify the model architecture and accelerate calculations. Although the accuracy of style-content decoupling is slightly reduced, the processing efficiency is improved, which is still better than traditional methods and can meet the needs of efficiency-sensitive scenarios.
[0166] The present invention has the following remarkable effects:
[0167] Accurately decouple style and content: Style and content parsers are specially constructed based on their essential differences and accurately injected into different blocks of U-Net, cutting off the entanglement between style and content, making style transfer flexible and diverse while preserving the content. For example, in qualitative comparison, compared with CSGO and other methods, the present invention fully retains the structure and details of content such as buildings and characters while accurately migrating the style.
[0168] Improve image quality in all aspects: The style parser is trained collaboratively with multiple losses, the content parser is finely optimized with multi-dimensional losses, and the dataset is accurately adapted according to its characteristics, improving image quality and style similarity from multiple dimensions.
[0169] Deeply optimized user visual experience.
[0170] Experimental Results
[0171] The quantitative evaluation results are shown in Table 1. Compared with the cutting-edge methods, the proposed method stands out in terms of CSD and CAS indicators, demonstrating its excellent effectiveness in style and content control.
[0172] Table 1
[0173]
[0174]
[0175] As shown in Table 2, the user study shows that the present invention leads the way in style and content scoring and overall quality selection, and the generated images achieve a delicate balance between quality, content similarity and style similarity. When users filter by visual appeal and content relevance, the present invention is overwhelmingly favored, which strongly proves its deep optimization of user visual experience.
[0176] Table 2
[0177]
[0178] In summary, the present invention proposes an image style transfer method based on content and style parser. Compared with the traditional technology, the important innovative contributions of the embodiments of the present invention include:
[0179] The style parser uses visual transformers to accurately extract style features. Its unique self-attention mechanism can capture the style associations of distant elements in the image, which is better than the local perception of traditional convolutional layers. When the style is embedded in a specific layer, multiple loss functions are trained collaboratively, adjusted according to different weight balances, and random image conditional discarding enables flexible style control during inference, laying the foundation for accurate style expression.
[0180] The content analyzer uses the expertise of spatial information management to design a content controller to process content images in layers. The multi-layer fusion content is embedded and the residual is injected into the U-Net downsampling block, which uses the basic content-sensitive characteristics of the image shape and contour to stabilize the content. The mean square error loss and adversarial loss are weighted optimized. The former corrects the prediction noise error and improves the generation quality, while the latter uses the discriminator game to ensure that the content is authentic and recognizable. The two balance to prevent content distortion and style pollution to ensure content integrity.
[0181] The application scenarios of the present invention include but are not limited to:
[0182] Video style conversion: Extend image style conversion technology to video streams. Process video images frame by frame, and unify the style according to the video theme or user instructions. For example, in movie special effects production, realistic scenes can be converted into oil painting or ink painting styles with one click, or retro or sci-fi styles can be switched according to the plot atmosphere. By taking advantage of the style-content management of the present invention, the video content can be kept coherent and the character movements can be natural, and the image jitter or content distortion caused by style conversion can be prevented, thereby improving the creation efficiency and artistic effect.
[0183] Virtual reality (VR) / augmented reality (AR) scene construction field: When generating virtual objects and scenes in VR / AR environment, the present invention is used to achieve style diversification and content adaptation. Customize virtual buildings and character styles according to user immersion needs and interactive plots. For example, historical theme VR experience ensures that the style of virtual monuments fits the times, the texture details are authentic, and the style rendering is smoothly switched according to the user's exploration perspective, and the user's immersion and interactive authenticity are enhanced through precise style conversion and content maintenance.
[0184] Digital art creation and design tool field: integrated into drawing and design software to empower artists and designers. For example, when drawing illustrations, you can quickly switch styles to explore creativity, and poster design can intelligently match image styles according to brand styles. By using efficient style conversion and intelligent content retention, it can inspire creative inspiration, shorten the design cycle, improve the quality of works, and transform the digital art creation process and efficiency.
[0185] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.
[0186] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.
[0187] An embodiment of the present invention further provides a processor, wherein the processor executes a computer program and at least executes the method described above.
[0188] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0189] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0190] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0191] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0192] Those skilled in the art can understand that: all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiments; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, etc. Various media that can store program codes.
[0193] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0194] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0195] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0196] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0197] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art of the present invention, several equivalent substitutions or obvious variations can be made without departing from the concept of the present invention, and the performance or use is the same, which should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for image style transfer based on content and style parser, characterized in that: The following steps are involved: S1. Building a basic model: Using the stable diffusion model as the basic architecture, the learning and encoding of image content features are achieved through the process of adding noise and denoising; S2. Build a style parser and extract style features: Use the visual deformer ViT as a style parser to extract style features from the reference image and generate style embedding through a multi-layer perceptron; S3. Build a content parser and extract content features: Build a content parser, process content images hierarchically, generate multiple levels of latent representations as content embeddings, and integrate content embeddings into the model through residual addition to stabilize the content architecture and preserve content features during style transfer; S4. Build content controller and inject content: Process the content image through the content controller to capture the basic structure and minimize the interference of style features; S5. Injecting style and content: Injecting style embedding into the upsampling block of the model to enhance style expression, while injecting content embedding into the downsampling block of the model to ensure content preservation; S6. Reasoning: In the reasoning stage, the constructed content parser and style parser are used to decouple style from content, capture and reproduce complex styles while maintaining the integrity of the content.
2. The image style transfer method based on content and style parser according to claim 1, characterized in that: The processing of the style parser specifically includes: The image is divided into blocks and linearly embedded using the visual deformer ViT; Calculate query, key and value matrices through self-attention mechanism, update input feature representation and generate style features; By normalizing the output of the self-attention mechanism, the training process is stabilized and the model generalization ability is improved; Style embedding is further extracted from the normalized features through a multi-layer perceptron (MLP).
3. The image style transfer method based on content and style parser according to claim 1 or 2, characterized in that: The training process of the style parser specifically includes: During training, we use the Gram loss to ensure style alignment by comparing the Gram matrices of features of the generated and target images. Keep the visual features of the generated image and the target image aligned through perceptual loss; The original style loss, perceptual loss and Gram loss are weighted summed to get the total style loss; Image conditions are randomly dropped during training, enabling no classifier bootstrapping during inference.
4. The image style transfer method based on content and style parser according to any one of claims 1 to 3, characterized in that: The style parser includes the following style injection strategies: Leveraging the hierarchical nature of convolutional neural networks, we identify and specifically use upsampling blocks to capture and inject style information; A criss-cross attention mechanism is implemented to inject style information into the upsampling block guided by image conditions, thereby enhancing style expression.
5. The image style transfer method based on content and style parser according to any one of claims 1 to 4, characterized in that: The operations of the content controller specifically include: A tiling control strategy is adopted to preserve content information without affecting the style transfer process; Use ControlNet to manage spatial information within the U-Net architecture, capture the basic structure of the content image, and minimize the interference of style features; The content fusion encoder integrates the content to generate multiple levels of latent representations to form content embedding; Integrate content embedding into U-Net via residual addition; Through the ControlNet mechanism, content embedding is accurately injected into the upsampling block of the diffusion model to achieve the separation and conversion of content and style.
6. The image style transfer method based on content and style parser according to any one of claims 1 to 5, characterized in that: The training process of the content parser specifically includes: Keep the parameters of the pre-trained diffusion model unchanged during the training phase; Trained using image-text pairs to achieve effective correspondence between content and text descriptions; Use mean squared error loss to measure and minimize the difference between the model's predicted noise and the actual noise; Use adversarial loss to train the discriminator to enhance its ability to distinguish between real and generated images; The content loss is formed by weighted combination of mean squared error loss and adversarial loss to balance the contribution of the two losses to content preservation quality.
7. The image style transfer method based on content and style parser according to any one of claims 1 to 6, characterized in that: The content embeddings derived from the content parser are injected into the downsampling block of the model.
8. The image style transfer method based on content and style parser according to any one of claims 1 to 7, characterized in that: In step S6, the reasoning includes a reversal operation for content preservation, specifically including: During the sampling process, the input image is encoded into the latent space, and the ReNoise technique reverses the sampling process to generate the latent noise representation directly from the real image, thus eliminating the need for additional noise injection; The noise is updated in an iterative manner, enhancing the approximation of the expected position in the forward diffusion progression by averaging the predictions.
9. The image style transfer method based on content and style parser according to any one of claims 1 to 8, characterized in that: Step S6 specifically includes: Using the disentangled property of the criss-cross attention mechanism, the image conditional weights for content and style conditions are independently adjusted during the inference phase. Construct a latent representation of the final image by weightedly combining the attention results of content and style conditions; Set the text condition weight factor to zero; When the weight factors of both content and style conditions are zero, the model reverts to the original text-to-image diffusion model.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the image style conversion method based on content and style parser as described in any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Method for preserving vision-language model pre-trained image and text knowledge in continuous learning
CN118798315A
Video style migration method and device based on diffusion model and electronic equipment
CN118887075A
System and Method for Augmenting Vision Transformers
US20230177662A1
Cited By
Automatic driving image generation method for controllable injection of road traffic conditions
CN120147995A
An image generation method for autonomous driving with controllable injection of road traffic conditions
CN120147995B