Method for quickly generating stylized text to image based on diffusion model
Through the potential consistency model and normalized mixed self-attention mechanism, representative style features are extracted from reference style images, solving the problem of time consumption and poor effect of stylized text generation images in the prior art, and achieving fast and efficient stylized image generation.
Patent Information
- Application Number
- CN202510634237.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The prior art requires fine-tuning of pre-trained large-scale diffusion models or using two-stage methods when generating images in stylized text, resulting in large time consumption and poor results, especially when the reference style image and content image are largely different.
Representative style features are extracted from reference style images using the self-consistent characteristics of the latent consistency model, and a normalized mixed self-attention mechanism is introduced to guide the text image generation process, avoid fine-tuning and inversion operations, and directly generate style-compliant images from the text.
The rapid generation of high-quality stylized images is achieved, reducing inference time and improving efficiency, and the generated images are highly consistent with the reference style image distribution.
Smart Images

Figure CN120543365A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method for quickly generating stylized text into an image based on a diffusion model. Background Art
[0002] Stylized text-based image generation aims to generate an image that matches a desired style based on a small number of style images, representing a new paradigm for image generation. Although similar to neural style transfer, stylized image generation differs fundamentally. Neural style transfer takes a content image and a style image as input and solves an image conversion task, focusing on transferring the artistic style of a style image to the content image. In contrast, stylized image generation generates an image that matches a specific style based on a given text prompt.
[0003] A straightforward approach is to first use a pre-trained text-to-image (T2I) model to generate images based on content cues. State-of-the-art style transfer methods are then used to convert the generated images into the specific style of the reference style image. However, this approach is primarily applicable when there is a certain degree of similarity between the reference style image and the content image. When the two differ significantly, the approach may fail. An alternative approach is to use more reference style images and perform a small amount of training / fine-tuning on the model using LoRA or an adapter, or even fine-tuning the entire T2I model. However, these methods require more time and effort, and are less convenient than simply providing a reference style image.
[0004] The challenges and drawbacks of these technologies are that they require the addition of additional adapters to fully or partially fine-tune pre-trained large-scale diffusion models, which is time-consuming and limits their practical application. Furthermore, two-stage approaches—first generating an image using a text-to-image generative model and then stylizing this image using advanced neural style transfer methods—often involve inverting the diffusion model, which doubles the inference time and often yields suboptimal results. Summary of the Invention
[0005] In view of this, the present invention provides a method for quickly generating stylized text into an image based on a diffusion model, so as to at least solve the above technical problems.
[0006] According to a first aspect of an embodiment of the present invention, a method for rapidly generating stylized text to an image based on a diffusion model is provided, comprising: S1, style feature extraction: inputting a reference style image with a resolution of 512×512×3, and extracting representative style features from the reference style image using the self-consistency of a latent consistency model; S2, generating an image from stylized text: guiding the process of generating an image from text using the extracted representative style features and a normalized hybrid self-attention mechanism.
[0007] Optionally, the method utilizes the self-consistency of the latent consistency model to extract representative style features from the input reference style image, including: mapping the reference style image to a low-dimensional latent space through the encoder of a pre-trained variational autoencoder to obtain an initial latent code; adding Gaussian noise to the initial latent code using a diffusion process noise injection formula to obtain a noisy latent code; inputting the noisy latent code into the noise prediction network of the latent consistency model, extracting the style statistical features of each layer of the Transformer as the representative style features through a single-step denoising operation, and the noise prediction network achieves stable mapping of features of adjacent time steps by minimizing the consistency loss function.
[0008] Optionally, the diffusion process noise injection formula is used to add Gaussian noise to the initial latent code to obtain the latent code after adding noise, which is expressed as:
[0009]
[0010] in, is the potential code after noise addition, t is the diffusion time step, ∈∈N(0,1) is Gaussian noise, z s is the initial latent code.
[0011] Optionally, the process of guiding text to generate images through the extracted representative style features and the normalized hybrid self-attention mechanism includes: in the Transformer feature layer of the diffusion model, fusing the content features with the extracted representative style features to obtain the content features of the intermediate output, and adjusting the content features of the intermediate output through the normalized hybrid self-attention mechanism to ensure that the generated stylized result is highly consistent with the distribution of the reference style image.
[0012] Optionally, the normalized hybrid self-attention mechanism is specifically as follows: in the self-attention module of the Transformer feature layer, the content features of the intermediate output are mapped to query features, key features and value features; the key features and value features are replaced with the key features and value features of the representative style features to obtain the replaced key features and value features, and the attention score matrix is calculated using the query features and the replaced key features and value features; a unified Softmax operation is performed on the attention score matrix to obtain a fused attention weight matrix; according to the fused attention weight matrix, the value features are weighted summed to obtain the stylized content features as the stylization result.
[0013] Optionally, the normalized hybrid self-attention mechanism further includes: before performing a unified Softmax operation on the attention score matrix, performing style distribution normalization processing on the content features, which is expressed as:
[0014]
[0015] in, and Represents content characteristics and representative style features The mean and standard deviation of .
[0016] According to a second aspect of an embodiment of the present invention, a system for rapidly generating stylized text to an image based on a diffusion model is provided, comprising: a style feature extraction module for inputting a reference style image with a resolution of 512×512×3 and extracting representative style features from the reference style image using the self-consistency of a latent consistency model; and a stylized text-to-image generation module for guiding the process of generating an image from text using the extracted representative style features and a normalized hybrid self-attention mechanism.
[0017] According to a third aspect of an embodiment of the present invention, there is provided an electronic device comprising a processor and a memory storing a program, wherein the program comprises instructions that, when executed by the processor, cause the processor to perform the steps of the method according to the first aspect.
[0018] According to a fourth aspect of an embodiment of the present invention, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method of the first aspect is implemented.
[0019] In summary, the present invention utilizes the self-consistency properties of latent consistency models (LCMs) to extract representative style statistics from reference style images to guide the stylization process. In addition, the present invention introduces a normalized mixture of self-attention mechanism, which enables the model to query the most relevant style patterns from these style statistics and use them to adjust the content features of the intermediate output. This mechanism also ensures that the generated stylized results are highly consistent with the distribution of the reference style image. The present invention can generate stylized images from text using only one style reference image, and has faster inference speed and higher efficiency than other methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a flow chart of the steps of the present invention.
[0022] Figure 2 Schematic diagram of CLIP feature similarity of different combinations at different time steps in the present invention.
[0023] Figure 3 Generate stylized text in an image and analyze the fusion of semantic and style information and principal component visualization.
[0024] Figure 4 A comparison of the effects of the normalized hybrid self-attention mechanism in stylized text generation images.
[0025] Figure 5 Schematic diagram of the impact of style distribution normalization.
[0026] Figure 6 For Figure 1 Corresponding overall flow chart of the present invention. DETAILED DESCRIPTION
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0028] The present invention is a method for fast stylized text-to-image generation based on a diffusion model, which can generate high-quality stylized images within six sampling steps based on a given style image. The present invention proposes a novel stylized image generation method, called UniSty, which utilizes a pre-trained large-scale diffusion model without fine-tuning or any additional optimization. Specifically, the present invention utilizes the self-consistency property of latent consistency models (LCMs) to extract representative style statistics from the reference style image to guide the stylization process. In addition, the present invention introduces a normalized mixture of self-attention mechanism, which enables the model to query the most relevant style patterns from these style statistics and use them to adjust the content features of the intermediate output. This mechanism also ensures that the generated stylized results are highly consistent with the distribution of the reference style image.
[0029] See also Figure 1 、 Figure 6 The present invention provides a method for rapidly generating stylized text to image based on a diffusion model, which mainly includes:
[0030] S1. Style feature extraction: Input a reference style image with a resolution of 512×512×3, and use the self-consistency of the latent consistency model to extract representative style features from the reference style image;
[0031] S2. Stylized Text to Image Generation: The text-to-image generation process is guided by extracted representative style features and a normalized hybrid self-attention mechanism. At this stage, after analyzing the strengths and weaknesses of various attention fusion modules, the present invention proposes a norm mixture of self-attention. This mechanism enables the model to retrieve the most relevant style patterns from these style statistics and use them to adjust the content features of the intermediate output. This mechanism also ensures that the generated stylized results are highly consistent with the distribution of the reference style image.
[0032] Optionally, the method utilizes the self-consistency of the latent consistency model to extract representative style features from the input reference style image, including: mapping the reference style image to a low-dimensional latent space through the encoder of a pre-trained variational autoencoder to obtain an initial latent code; adding Gaussian noise to the initial latent code using a diffusion process noise injection formula to obtain a noisy latent code; inputting the noisy latent code into the noise prediction network of the latent consistency model, extracting the style statistical features of each layer of the Transformer as the representative style features through a single-step denoising operation, and the noise prediction network achieves stable mapping of features of adjacent time steps by minimizing the consistency loss function.
[0033] Optionally, the diffusion process noise injection formula is used to add Gaussian noise to the initial latent code to obtain the latent code after adding noise, which is expressed as:
[0034]
[0035] in, is the potential code after noise addition, t is the diffusion time step, ∈∈N(0,1) is Gaussian noise, z s is the initial latent code.
[0036] For example, the representative style features extracted in step S1 are specifically:
[0037] S11, in Figure 2 In this paper, two noise injection methods are applied to the style image. The forward noise addition process of the diffusion model (Eq. 1) and DDIM inversion techniques, and uses two different models, latent consistency models (LCMs) and stable diffusion (SD), for single-step denoising. The present invention uses CLIP's image encoder to extract features of the original style image and the denoised image, and calculates their cosine similarity. Figure 2 The results show that the performance of all three combinations gradually decreases with increasing time steps. However, the performance gap between the combination of Eq. 1 + LCMs and the combination of DDIM Inversion + SD remains small. This demonstrates that LCMs can effectively extract representative style statistics from noisy style images, thus avoiding the time-consuming DDIM inversion operation. The fundamental reason for this is that the optimization objective of LCMs is to minimize the difference in the consistency function output between adjacent samples. This mechanism enables LCMs to maintain representative style statistics, or representative style features, during single-step prediction.
[0038] S12: Based on the findings in step S1, the present invention uses latent consistency models (LCMs) to extract representative style statistics from the reference style image. Input reference style image I s , by pre-training the encoder E of the variational autoencoder, the reference style image I s Mapped to a low-dimensional latent space, the initial latent code z is obtained s =E(I s ); using the diffusion process noise injection formula Add Gaussian noise to the latent code, where t is the diffusion time step and ∈∈N(0,1) is Gaussian noise;
[0039] S13, the latent coding after adding noise in step S12 Noise prediction network F with input latent consistency model θ , extract the style statistical features of each layer of Transformer through a single-step denoising operation Here p represents the text input. The noise prediction network achieves a stable mapping of features of adjacent time steps by minimizing the consistency loss function.
[0040] Optionally, the process of guiding text to generate images through the extracted representative style features and the normalized hybrid self-attention mechanism includes: in the Transformer feature layer of the diffusion model, fusing the content features with the extracted representative style features to obtain the content features of the intermediate output, and adjusting the content features of the intermediate output through the normalized hybrid self-attention mechanism to ensure that the generated stylized result is highly consistent with the distribution of the reference style image.
[0041] Optionally, the normalized hybrid self-attention mechanism is specifically as follows: in the self-attention module of the Transformer feature layer, the content features of the intermediate output are mapped to query features, key features and value features; the key features and value features are replaced with the key features and value features of the representative style features to obtain the replaced key features and value features, and the attention score matrix is calculated using the query features and the replaced key features and value features; a unified Softmax operation is performed on the attention score matrix to obtain a fused attention weight matrix; according to the fused attention weight matrix, the value features are weighted summed to obtain the stylized content features as the stylization result.
[0042] Optionally, the normalized hybrid self-attention mechanism further includes: before performing a unified Softmax operation on the attention score matrix, performing style distribution normalization processing on the content features, which is expressed as:
[0043]
[0044] in, and Represents content characteristics and representative style features The mean and standard deviation of .
[0045] For example, the normalized mixture of self-attention mechanism in step S2 is specifically:
[0046] S21. Define the content features of the diffusion model Transformer special layer as The goal of this invention is to convert the representative style statistics into Seamless integration into content features In order to obtain stylized content features In the Transformer layer of the backbone network, the features After the self-attention module, they are mapped into queries key Sum feature.
[0047] S22. The present invention first directly replaces the key value with the key value of the representative style statistics, that is, the representative style features, because intuitively, the query (Q) can be used to represent semantic information, such as image layout, while the key (K) and value (V) features are used to represent style statistics, such as color, texture and lighting. Stylized content features It can be expressed as:
[0048]
[0049] Where A represents the attention calculation and σ represents the softmax activation function. The present invention observes that the description in this step often prioritizes capturing style statistics at the expense of semantic information obtained from the prompt. For example, Figure 3 As shown in the first row, the method fails to generate the expected content corresponding to the prompt words "girl" and "house". To further analyze this phenomenon, the present invention performs principal component analysis (PCA) on the features or attention maps from different backbone network layers, including ResNet blocks and query (Q) and key (K) layers in self-attention. Figure 3 The second row visualizes the first three principal components, and the results further reveal that the semantic information is closer to the style reference image rather than the given conditional cue words.
[0050] S23. In order to solve the above problems, the present invention further enhances the semantic information by reintroducing the semantic information. The semantic representation in is as follows,
[0051]
[0052] Where λ∈[0,1] is a weight hyperparameter. The present invention found that the robustness of this method is poor and highly dependent on the choice of λ. For example, Figure 4 As shown, under the same λ setting, the stylized images generated in the second row outperform the first row in semantic presentation.
[0053] The present invention believes that the reason for the unstable performance is that the two attention calculations in the formula in this step of the equation are performed separately. Looking back at the calculation process of the standard attention mechanism, negative numbers will be mapped to extremely small values after the exponential operation, thereby effectively reducing their contribution to the attention weight matrix. However, in the stylized text to image (T2I) task, if there is a large difference between the style statistics and the semantic information, the content features will be and representative style statistics The attention score matrix (before Softmax) may contain negative values in some rows, which makes the exponential function ineffective. Therefore, it is necessary to introduce an additional weight coefficient λ to balance the two components on the right side of the equation.
[0054] S24. To avoid the above problems, it is necessary to ensure that the calculation of the attention score matrix considers both the intra-class semantic differences and the inter-class (semantic and style) information differences, rather than calculating these components separately. The present invention first rewrites the formula in step S23 into a matrix form, resulting in the following expression:
[0055]
[0056] To achieve this unified attention score matrix calculation, this method forms a new matrix By merging and Then the present invention rewrites M in the above formula as
[0057]
[0058] The above formula performs a Softmax operation on the inter-class information and the intra-class information as a whole, avoiding the aforementioned failure problem that may occur when the exponential function maps the attention score matrix to the attention weight matrix. It can be obtained in the following ways;
[0059]
[0060] The above formulation enables the model to aggregate semantic and style information more adaptively while reducing the strong dependence on the choice of λ.
[0061] The above formula uses content features Get each input area to query and get representative style statistics The most relevant information. However, this approach sometimes leads to and The style distribution between does not match because it does not consider the global style distribution. Figure 5 As shown in the first row, this mismatch manifests as global color inconsistency between the generated stylized image and the reference style image.
[0062] Therefore, the present invention introduces a style distribution normalization process before applying the above formula, which is inspired by this method. The process can be expressed in the following mathematical form:
[0063]
[0064] in, and Represents content characteristics and representative style statistics The mean and standard deviation of . Figure 5 As shown in Figure 2, after applying style distribution normalization, the global hue is more consistent with the style image than without normalization (top row). Finally, the present invention collectively refers to the above two processes as norm mixture of self-attention.
[0065] In summary, the present invention utilizes the self-consistency properties of latent consistency models (LCMs) to extract representative style statistics from reference style images to guide the stylization process. In addition, the present invention introduces a normalized mixture of self-attention mechanism, which enables the model to query the most relevant style patterns from these style statistics and use them to adjust the content features of the intermediate output. This mechanism also ensures that the generated stylized results are highly consistent with the distribution of the reference style image. The present invention can generate stylized images from text using only one style reference image, and has faster inference speed and higher efficiency than other methods.
[0066] The present invention has conducted a large number of experiments to prove the feasibility of the method of the present invention. See Table 1 for a comparison of the effects of different text-to-image generation methods:
[0067] Table 1 Comparison of different text-to-image generation methods
[0068]
[0069] The embodiment of the present invention further provides a system for rapidly generating stylized text into an image based on a diffusion model, comprising:
[0070] The style feature extraction module takes as input a reference style image with a resolution of 512×512×3 and extracts representative style features from the reference style image using the self-consistency of the latent consistency model.
[0071] Stylized text-to-image module, which guides the text-to-image generation process through extracted representative style features and normalized hybrid self-attention mechanism.
[0072] It should be understood that the diffusion model-based rapid stylized text-to-image generation system of this embodiment is used to implement the corresponding methods in the aforementioned multiple method embodiments and has the beneficial effects of the corresponding method embodiments.
[0073] As another example, the present invention also provides an electronic device, which will now be described as an electronic device that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0074] The electronic device may include: a processor, a communication interface, a memory, and a communication bus.
[0075] The processor, communication interface and memory communicate with each other through a communication bus. The communication interface is used to communicate with other electronic devices or servers.
[0076] The processor is used to execute programs, and specifically can execute the relevant steps in the above method embodiments.
[0077] Specifically, the program may include program codes including computer operation instructions.
[0078] The processor may be a CPU, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or processors of different types, such as one or more CPUs and one or more ASICs.
[0079] The memory is used to store programs and may include high-speed RAM memory or non-volatile memory, such as at least one disk storage.
[0080] When the program is executed by a processor, the program enables the electronic device to execute a method for rapidly generating stylized text to image based on a diffusion model of the present invention.
[0081] In addition, the specific implementation of each step in the program can refer to the corresponding description of the corresponding steps and units in the above method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the above method embodiments, and will not be repeated here.
[0082] An exemplary embodiment of the present invention further provides a computer storage medium storing a computer program, wherein when the computer program is executed by a processor, the methods of the various embodiments of the present invention are implemented. The corresponding process descriptions in the aforementioned method embodiments can be referred to and will not be repeated here.
[0083] The method according to the embodiment of the present invention described above can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded via a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown here.
[0084] Thus far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous.
[0085] It should be understood that although this specification is described according to various embodiments, not every embodiment contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
[0086] Finally, it should be noted that the above implementation methods are only used to illustrate the embodiments of the present invention, and are not limitations on the embodiments of the present invention. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the scope of patent protection of the embodiments of the present invention should be defined by the claims.
Claims
1. A method for rapidly generating stylized text to image based on a diffusion model, characterized in that: include: S1. Style feature extraction: Input a reference style image with a resolution of 512×512×3, and use the self-consistency of the latent consistency model to extract representative style features from the reference style image; S2. Stylized Text to Image Generation: The process of text-to-image generation is guided by the extracted representative style features and the normalized hybrid self-attention mechanism.
2. The method according to claim 1, characterized in that The method of extracting representative style features from an input reference style image using the self-consistency of the latent consistency model includes: Map the reference style image to a low-dimensional latent space through the encoder of the pre-trained variational autoencoder to obtain the initial latent code; The diffusion process noise injection formula is used to add Gaussian noise to the initial latent code to obtain the noisy latent code; The noisy latent code is input into the noise prediction network of the latent consistency model, and the style statistical features of each layer of the Transformer are extracted as the representative style features through a single-step denoising operation. The noise prediction network achieves stable mapping of features of adjacent time steps by minimizing the consistency loss function.
3. The method according to claim 2, characterized in that The diffusion process noise injection formula is used to add Gaussian noise to the initial latent code to obtain the latent code after adding noise, which is expressed as: in, is the potential code after noise addition, t is the diffusion time step, ∈∈N(0,1) is Gaussian noise, z s is the initial latent code.
4. The method according to claim 3, characterized in that The process of generating images from text guided by the extracted representative style features and the normalized hybrid self-attention mechanism includes: In the Transformer feature layer of the diffusion model, the content features are fused with the extracted representative style features to obtain the content features of the intermediate output. The content features of the intermediate output are then adjusted through the normalized hybrid self-attention mechanism to ensure that the generated stylized results are highly consistent with the distribution of the reference style image.
5. The method according to claim 4, characterized in that The normalized hybrid self-attention mechanism is specifically: In the self-attention module of the Transformer feature layer, the content features of the intermediate output are mapped into query features, key features, and value features; The key features and value features are replaced with the key features and value features of the representative style features to obtain the replaced key features and value features, and the attention score matrix is calculated using the query features and the replaced key features and value features; Perform a unified Softmax operation on the attention score matrix to obtain the fused attention weight matrix; According to the fused attention weight matrix, the value features are weighted and summed to obtain stylized content features as the stylization result.
6. The method according to claim 5, characterized in that The normalized hybrid self-attention mechanism also includes: Before performing a unified Softmax operation on the attention score matrix, the content features are normalized by style distribution, which can be expressed as: in, and Represents content characteristics and representative style features The mean and standard deviation of .
7. A rapid stylized text-to-image generation system based on a diffusion model, characterized by: include: The style feature extraction module takes as input a reference style image with a resolution of 512×512×3 and extracts representative style features from the reference style image using the self-consistency of the latent consistency model. Stylized text-to-image module, which guides the text-to-image generation process through extracted representative style features and normalized hybrid self-attention mechanism.
8. An electronic device, characterized in that: include: processor; Memory for storing programs; The program includes instructions, which, when executed by the processor, cause the processor to perform the steps of the method according to any one of claims 1 to 6.
9. A computer storage medium, characterized in that A computer program is stored thereon, and when the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Style migration image processing method and device, electronic equipment and storage medium
CN118447262A
Method for generating style alignment image set based on Diffusion Transform
CN119515669A
Embedded reconstructed text-image alignment style migration method
CN119941492A
Multi-view 3D diffusion
US20250078392A1
Cited By
Constructive information hiding method based on diffusion model and style migration
CN121563756A
A constructional information hiding method based on diffusion model and style transfer
CN121563756B