Embedded reconstructed text-image alignment style migration method
By introducing a style extraction module based on perceptron attention and a text-image alignment enhancement module based on cross attention mechanism in the diffusion model, combined with an explicit modulation module, the problems of inaccurate art image style feature extraction and insufficient text guidance in the prior art are solved, and high-fidelity, diversified art style transfer and text semantic-driven image generation are achieved.
Patent Information
- Application Number
- CN202510005224.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-02
AI Technical Summary
The existing technology is inaccurate and incomplete in the extraction of artistic image style features, resulting in lack of artistic appeal and low style restoration of styles. Text guidance is easily weakened by image data during the generation process, and cannot accurately drive the style and content to be shaped according to text semantics.
The diffusion model is adopted as the infrastructure, combining the style extraction module based on perceptron attention and the text-image alignment enhancement module of the cross-attention mechanism, capture image style details and text semantics through a multi-layer architecture and attention mechanism, optimize the interaction between text and image information, and realize the flexible fusion of multimodal embedding and original embedding through an explicit modulation module.
It significantly improves the ability to capture the details and abstract characteristics of the complex style of artistic images, realizes high-fidelity and diversified presentation of artistic styles, strengthens the guiding role of text in the entire process of image generation, ensures that the generated image is closely in line with the text semantic guidance, and improves the controllability and accuracy of style transfer.
Smart Images

Figure CN119941492A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence computer vision processing technology, and in particular to a method for embedding and reconstructing text-image alignment style transfer. Background Art
[0002] In the field of text-guided image style transfer, some existing methods use traditional pre-trained diffusion models combined with simple adapter modules to handle style transfer tasks. However, there are significant drawbacks: on the one hand, the model architecture and training data are adapted to the characteristics of natural images. When faced with artistic images, they are unable to capture unique style details and abstract style concepts such as oil painting brushstrokes and Chinese painting artistic conception, resulting in a lack of artistic appeal and low style restoration in the style transfer results. On the other hand, when processing image and text embedding, the information imbalance between the two is not effectively compensated. Direct splicing causes the text guidance to be easily weakened by image data during the generation process, and it is impossible to accurately drive the style and content to be shaped according to the text semantics. The final output image style is single and poorly consistent with the text description, which greatly limits the in-depth expansion and efficient application of style transfer technology in multiple scenarios such as artistic creation and design creativity.
[0003] It should be noted that the information disclosed in the above background technology section is only used for understanding the background of the present application, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the invention
[0004] The main purpose of the present invention is to overcome the defects existing in the above-mentioned background technology and provide a method for embedding and reconstructing text-image alignment style transfer.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for embedding and reconstructing text-image alignment style transfer, comprising the following steps:
[0007] S1. Basic model construction: The diffusion model is used as the basic architecture. The model gradually adds noise to the image data through the forward process, simulating the gradual transformation of the image from a clear state to a completely noisy state. Then, the original image is restored from the noisy image through the denoising process to achieve the learning and encoding of image style features.
[0008] S2. Attention-based style extraction: For reference images in the artwork dataset, the input image embedding is extracted through a feature extraction network to capture the visual features of the image; the image embedding is then further processed using the perceptron attention mechanism and the feed-forward network FFN to capture complex style details and generate a style embedding;
[0009] S3. Text embedding generation: The text description is converted into text embedding through a text encoder to guide the style transfer process;
[0010] S4. Text-image alignment enhancement: The style embedding is fused with the text embedding through a cross-attention mechanism to generate a multimodal embedding to optimize the interaction between text and image information;
[0011] S5. Explicit modulation and diffusion model integration: The multimodal embedding is fused with the image embedding through linear interpolation to form a complete prompt embedding, which is then integrated into the diffusion model to generate diverse style images that match the text description.
[0012] Furthermore, the diffusion model uses U-Net to implement the denoising process and is trained through mean square error loss.
[0013] Furthermore, step S2 specifically includes:
[0014] Generate initial latent variables for the reference image, which are normalized to stabilize the training process;
[0015] The latent variables are expanded to match the batch size of the input images to maintain consistency when processing multiple images;
[0016] Updating the latent variables through the perceptron attention mechanism enables the model to selectively focus on different parts of the input image based on the learnable latent variables;
[0017] The updated latent variables are transformed nonlinearly through the feedforward network FFN to enhance the style encoding capability;
[0018] The output of the feed-forward network and the updated latent variables are combined to generate an embedding representing the style of the input image for subsequent text-image alignment enhancement.
[0019] Furthermore, in step S2, the visual part of the CLIP model is used as a feature extraction network to obtain an embedded representation of the input image as a reference basis for style extraction; the output of the perceptron attention mechanism is further processed by a position-by-position feed-forward network FFN to capture more fine-grained style features; the output of the FFN is combined with the updated latent variables to obtain the final style embedding.
[0020] Furthermore, the feedforward network FFN includes at least two linear transformation steps for weighting the latent variables, and a nonlinear activation function is introduced between the two linear transformations to enhance the model's ability to learn complex features.
[0021] Furthermore, step S4 specifically includes:
[0022] The style embedding and text embedding are converted into query, key, and value matrices respectively through a linear layer to interact in a shared feature space;
[0023] The attention weights are computed using the dot product of the query matrix and the key matrix and scaled to prevent gradient vanishing, thereby dynamically prioritizing different aspects of the text prompt.
[0024] Apply the softmax function to the raw attention scores to obtain normalized attention weights, ensuring that the sum of the weights is 1 for effective information integration;
[0025] The value matrix is weighted summed using normalized attention weights to generate a multimodal embedding that can more effectively capture the multimodal context of text and images.
[0026] Through the above process, the model is able to more effectively integrate image style and text embedding, allowing the generation of images that are more consistent with the semantic content of text cues, which is particularly suitable for scenarios that require tight integration of text and visual information.
[0027] Furthermore, step S5 specifically includes:
[0028] The style embedding is fused with the multimodal embedding obtained by text-image alignment enhancement through linear interpolation to achieve a smooth transition and effective combination between the two embeddings.
[0029] Use predefined constants to control the fusion ratio between style embedding and multimodal embedding, ensuring flexibility in adjusting the fusion according to different application scenarios and requirements;
[0030] The fused image embedding is concatenated with the text prompt embedding to form a complete prompt embedding for image generation, which integrates the multimodal information of text and image.
[0031] Integrating the generated full prompt embedding into the diffusion model enables the model to simultaneously consider the multimodal conditions of text and image during the generation process, improving the performance and quality of generated images.
[0032] Through the above fusion and concatenation operations, the model is able to capture multimodal conditions robustly and controllably, thereby achieving effective alignment of text description and image style during image generation.
[0033] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the method for embedding and reconstructing text-image alignment style transfer.
[0034] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the method for embedding and reconstructing text-image alignment style transfer is implemented.
[0035] The present invention has the following beneficial effects:
[0036] The present invention provides a method for embedding and reconstructing text-image alignment style transfer, which innovatively solves the inaccuracy and incompleteness of the existing technology in extracting artistic image style features. By constructing a style extraction module based on perceptron attention and multi-layer architecture, the ability to capture complex style details and abstract characteristics of artistic images is significantly improved, thereby achieving high-fidelity and diversified presentation of artistic styles in style transfer. The present invention optimizes the fusion process of text and image embedding by designing a sophisticated text-image alignment enhancement module and utilizing a cross-attention mechanism, strengthens the guiding role of text in the entire image generation process, ensures that the generated image can closely fit the text semantic guidance, and significantly improves the controllability and accuracy of style transfer. The introduction of the explicit modulation module, through the linear interpolation fusion rule and the fusion ratio control mechanism, the embodiment of the present invention realizes the flexible fusion of multimodal embedding and original embedding, balances content and style, and enables the generated image to show rich style variants while maintaining the core content. The combined effects of these technical advantages not only enrich the style diversity of creative materials, but also improve the accuracy of the fit between creative content and creative ideas, and expand the depth and efficient application of style transfer technology in multiple scenarios such as artistic creation and design creativity, thereby enhancing the practical value and innovation potential of style transfer technology from multiple dimensions.
[0037] The significant advantages of the present invention include the following aspects:
[0038] (1) Excellent style representation
[0039] The style extraction module based on a unique architecture far surpasses traditional technology in artistic style encoding, and can accurately capture the complex style details and abstract characteristics of artistic images, making the generated images have a distinct artistic style and the style transfer effect delicate and realistic, greatly enriching the style diversity of creative materials.
[0040] (2) Text guidance is accurate and efficient
[0041] The text-image alignment enhancement module successfully reverses the disadvantages of text guidance, strengthens the control of text over each stage of image generation, ensures that the generated image faithfully responds to the text description, improves the accuracy of the fit between the creative content and the creative conception, and expands the depth of application of style transfer in the field of creative design.
[0042] (3) Rich and flexible output
[0043] The dynamic fusion mechanism of the explicit modulation module unlocks the presentation of diverse styles. While maintaining the core content, it gives the generated images rich style variations, provides users with massive stylized visual results, strongly supports the style innovation needs of diverse creative scenarios, and enhances the practical value and innovation potential of style transfer technology from multiple dimensions.
[0044] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 The figure is an overall flow chart of the text-to-image style transfer method based on the attention mechanism according to an embodiment of the present invention.
[0046] Figure 2 It is a diagram of the algorithm architecture of text-image alignment enhancement and explicit modulation according to an embodiment of the present invention.
[0047] Figure 3 This is a visualization result of an embodiment of the present invention.
[0048] Figure 4 The results are compared between the embodiment of the present invention and those generated based on a universal adapter.
[0049] Figure 5 Qualitative comparison of the embodiments of the present invention with other state-of-the-art text-guided stylization methods.
[0050] Figure 6 Visual differences between samples with global features and attention-based style features generated for embodiments of the present invention.
[0051] Figure 7 This is a visualization diagram of the ablation study of the text-image alignment enhancement module in an embodiment of the present invention.
[0052] Figure 8 This is a heat map of the ablation study of the text-image alignment enhancement module in an embodiment of the present invention.
[0053] Fig. 9 1 is a diagram showing the effects of different degrees of explicit modulation in an embodiment of the present invention.
[0054] Fig.10 The following are visualization results of training the embodiments of the present invention under different data sets.
[0055] Fig.11 It is the guiding result of the combination of the embodiment of the present invention and the additional conditions. DETAILED DESCRIPTION
[0056] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope and application of the present invention.
[0057] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0058] See also Figure 1 , an embodiment of the present invention provides a method for embedding and reconstructing text-image alignment style transfer, comprising the following steps:
[0059] Step S1. Basic model construction: The diffusion model is used as the basic architecture. The model gradually adds noise to the image data through a forward process, simulating the image gradually changing from a clear state to a completely noisy state. The original image is then restored from the noisy image through a denoising process to achieve learning and encoding of image style features.
[0060] In a preferred embodiment, the diffusion model uses U-Net to implement the denoising process and is trained through mean square error loss.
[0061] Step S2. Attention-based style extraction: For the reference images in the artwork dataset, the input image embedding is extracted through the feature extraction network to capture the visual features of the image; then the image embedding is further processed using the perceptron attention mechanism and the feedforward network FFN to capture complex style details and generate a style embedding.
[0062] In a preferred embodiment, step S2 specifically includes: generating an initial latent variable for the reference image, which is normalized to stabilize the training process; expanding the latent variable to match the batch size of the input image so as to maintain consistency when processing multiple images; updating the latent variable through the perceptron attention mechanism so that the model can selectively focus on different parts of the input image according to the learnable latent variable; performing a nonlinear transformation on the updated latent variable through the feedforward network FFN to enhance the style encoding ability; combining the output of the feedforward network and the updated latent variable to generate an embedding representing the style of the input image for subsequent text-image alignment enhancement. In a further preferred embodiment, the visual part of the CLIP model is used as a feature extraction network to obtain an embedded representation of the input image as a reference basis for style extraction; further processing the output of the perceptron attention mechanism through the position-by-position feedforward network FFN to capture more fine-grained style features; combining the output of the FFN with the updated latent variable to obtain the final style embedding. Preferably, the feedforward network FFN includes at least two linear transformation steps for weighting the latent variables, and introducing a nonlinear activation function between the two linear transformations to enhance the model's learning ability for complex features.
[0063] Step S3. Text embedding generation: The text description is converted into text embedding through a text encoder to guide the style transfer process.
[0064] Step S4. Text-image alignment enhancement: The style embedding obtained in step S2 is fused with the text embedding in step S3 through a cross-attention mechanism to generate a multimodal embedding to optimize the interaction between text and image information.
[0065] In a preferred embodiment, step S4 specifically includes: converting the style embedding (image prompt embedding) and the text embedding into query, key and value matrices respectively through a linear layer so as to interact in a shared feature space; calculating the attention weights using the dot product of the query matrix and the key matrix, and preventing the gradient from disappearing by scaling, thereby dynamically prioritizing different aspects of the text prompts; applying a softmax function to the original attention scores to obtain normalized attention weights, ensuring that the sum of the weights is 1 for effective information integration; using the normalized attention weights to perform weighted summation on the value matrix to generate a multimodal embedding that can more effectively capture the multimodal context of text and images. Through the above process, the model can more effectively integrate image and text embeddings, allowing the generation of images that are more consistent with the semantic content of the text prompts, which is particularly suitable for scenarios that require tight integration of text and visual information.
[0066] Step S5. Integration of explicit modulation and diffusion model: The multimodal embedding in step S4 is fused with the style embedding through linear interpolation, and the fused image embedding is concatenated with the text hint embedding to form a complete hint embedding, which is then integrated into the diffusion model to generate a diversified style image that matches the text description.
[0067] In a preferred embodiment, step S5 specifically includes: fusing the style embedding with the multimodal embedding obtained by text-image alignment enhancement by a linear interpolation method to achieve a smooth transition and effective combination between the two embeddings; using a predefined constant to control the fusion ratio between the style embedding and the multimodal embedding to ensure that the flexibility of the fusion can be adjusted according to different application scenarios and requirements; splicing the image embedding obtained after the fusion with the text prompt embedding to form a complete prompt embedding for image generation, which integrates the multimodal information of text and image; integrating the generated complete prompt embedding into the diffusion model, so that the model can consider the multimodal conditions of text and image at the same time during the generation process, thereby improving the performance and quality of the generated image. Through the above fusion and splicing operations, the model can robustly and controllably capture the multimodal conditions, thereby achieving effective alignment of text description and image style during the image generation process.
[0068] The embedded reconstruction text-image alignment style transfer method of the present invention overcomes the problem of inaccurate and incomplete extraction of artistic image style features in the prior art, constructs an effective mechanism that can deeply mine and accurately encode artistic image style elements, and achieves high-fidelity and diversified presentation of artistic style in style transfer. The present invention resolves the dilemma of information imbalance when text and image are embedded and fused, designs a sophisticated text-image alignment strategy, strengthens the core regulatory position of text guidance in the entire image generation process, ensures that the generated image closely fits the text semantic guidance, and greatly improves the controllability and accuracy of style transfer.
[0069] The following further describes an algorithm example and experimental verification of a specific embodiment of the present invention.
[0070] The present invention is built on the basic architecture of the diffusion model and mainly contains three core modules: an attention-based style extraction module, which captures multi-level style elements of the image through perceptron attention and a multi-layer architecture; a text-image alignment enhancement module, which uses a cross-attention mechanism to optimize the embedding fusion and interaction of text and image; and an explicit modulation module, which fuses multimodal embedding with linear interpolation and splicing strategies to improve the quality and diversity of generated images.
[0071] Basic Model
[0072] The diffusion model consists of two processes: a forward process that gradually adds Gaussian noise ∈ to the data x0 through a Markov chain. In addition, a denoising process removes the Gaussian noise xT ~N(0,1) generates samples and uses a learnable denoising model ∈ θ (x t , t, c), the model is parameterized by θ. This denoising model ∈ θ (·) is implemented with U-Net and trained with the mean squared error loss derived from a simplified variant of the variational bound:
[0073]
[0074] where c represents an optional condition. In diffusion models, c is usually represented by the textual cue using the CLIP-encoded text embedding E t It is represented and integrated into the diffusion model through the following design modules.
[0075] Attention-based style extraction
[0076] Style extraction methods enhance style encoding capabilities by integrating fine-grained features in multi-layer architectures. The goal is to capture complex style details from images using perceptron attention and position-wise feed-forward networks (FFNs).
[0077] Given a reference image, the input image embedding is obtained through CLIP, denoted as x. The latent variable z is initialized as a tensor of shape (1, N, D), where N is the number of queries and D is the dimension of the latent space, and normalized by dividing by the square root of D to stabilize the training process:
[0078]
[0079] To match the batch size of the input x, the latent variable z is expanded by repeating it over the batch dimension. This process can be expressed as:
[0080]
[0081] 1 of them B is a tensor of all ones with shape (B, 1, 1), where B is the batch size of x. The result of this operation is a tensor z with shape (B, N, D), where N is the number of queries and D is the dimension of the latent space.
[0082] The perceptron attention mechanism denoted as P-Attn is then applied to update the latent variable by focusing on the input x and the repeated latent variable:
[0083]
[0084] z′=P-Attn(x,z)+z,
[0085] where d kis the dimension of the key tensor, typically equal to D. This operation allows the model to selectively focus on different parts of the input data based on the learnable latent variables.
[0086] FFN consists of two linear transformations with a GELU activation function in between:
[0087] FFN(z′)=W2·GELU(W1·z′+b1)+b2,
[0088] Where W1 and W2 are weight matrices, and b1 and b2 are bias terms.
[0089] Output E I represents the style embedding extracted from the input image, obtained by combining the FFN output and the updated latent variable z′:
[0090] E I =FFN(z′)+z′.
[0091] Text-image alignment enhancement
[0092] The text-image alignment enhancement method aims to dynamically prioritize different aspects of textual cues by leveraging a cross-attention mechanism. This module allows the model to more effectively integrate image and text embeddings, projecting them into a shared feature space so that the interaction between them can be more subtle.
[0093] First, the image cue is embedded into E through a linear layer I and text prompts embedded in E T Transformed into query, key, and value matrices. These transformations are represented as:
[0094]
[0095] and are the weight matrices associated with image query, text key, and text value, respectively.
[0096] Attention weight w I By taking the query matrix Q I and key matrix The dot product is calculated by key dimension Scaling is done by the square root of to prevent the gradient from vanishing:
[0097]
[0098] The softmax function is then applied to these raw attention scores to obtain the normalized attention weights w′ I :
[0099] w′ I =softmax(wI )
[0100] Use the normalized attention weights w I , calculate the value matrix V T to generate a multimodal embedding E′ IT :
[0101] E′ IT =w′ I ·V T·
[0102] E′ IT Capturing multimodal context more effectively allows the model to generate images that are more consistent with the semantic content of the textual cues. This approach is particularly beneficial in scenarios where textual and visual information need to be tightly integrated to produce a coherent output.
[0103] Explicit Modulation
[0104] Explicit modulation addresses the lack of flexibility of traditional fusion methods by seamlessly fusing image-cue embeddings with multimodal embeddings enhanced by text-image alignment using linear interpolation.
[0105] Specifically, the image cue is embedded into E via linear interpolation. I and multimodal embedding E′ IT To perform the fusion:
[0106] E F =αE I +(1-α)E′ IT ,
[0107] where α is a predefined constant that controls the fusion ratio between the original embedding and the enhanced embedding.
[0108] Finally, the fused image is embedded into E F Embed E with text prompt T Concatenate to form the complete prompt embedding for image generation:
[0109]
[0110] in represents the splicing operation, E P Represents enhanced embedding and is integrated into the diffusion model. By balancing the above embeddings, the model obtains a robust and controllable representation that effectively captures multimodal conditions and improves generation performance.
[0111] Other embodiments
[0112] Attention mechanism optimization: In the style extraction module, fine-tune hyperparameters such as the number of perceptron attention heads and key-value dimensions, or introduce variants such as multi-head attention to balance computational cost and style capture accuracy; the alignment enhancement module attempts position encoding improvements or mixed attention modes to enhance the ability to associate long texts with complex images, and moderately increase model complexity in exchange for advanced style transfer effects.
[0113] Improved fusion strategy: The explicit modulation module explores nonlinear interpolation functions or adaptive fusion weight calculation methods, dynamically adapts the fusion strength according to the image content and text semantics, or guides the selection of embedded fusion areas based on the image semantic segmentation results, so as to improve the quality of generated images with more intelligent fusion strategies. Although it increases the design complexity, it expands the direction of technical optimization.
[0114] Data Augmentation
[0115] Dataset expansion and screening: Expand the art dataset to include rare art styles, specific cultural images, and multimodal derivative data. Optimize data distribution through intelligent screening and weighted sampling to improve the model's adaptability to multiple styles. At the same time, use image enhancement technology to enrich the texture, color, perspective and other changes of training samples to enhance model generalization and robustness, and strengthen the model's style learning ability while increasing the amount of data processing.
[0116] Multimodal data fusion optimization: explore shallow fusion of image and text multimodal features in the data preprocessing stage, or introduce external knowledge graphs, semantic networks and other multimodal knowledge in a specific stage of model training to assist the model in building a better text-image semantic mapping, and innovate the timing and method of multi-source data fusion to enhance the depth and breadth of style transfer knowledge and expand the boundaries of technology.
[0117] Experimental Results
[0118] Table 1
[0119] CLIP-Text↑ CLIP-Image↑ Dinov2↑ LPIPS↑ The present invention 22.57 69.48 40.92 0.4908 StyleShot 22.46 59.47 23.45 0.4561 Style Aligned 18.26 63.82 27.89 0.3892 VSP 22.29 66.31 35.34 0.2653 InstantStyle 19.78 62.59 29.76 0.1478 InstantStyle(Plus) 19.53 64.65 30.62 0.1234 Ip Adapter 15.01 68.90 38.12 0.3451 IP-Adapter(SDXL) 19.14 67.43 32.98 0.2987 CSGO 17.82 55.16 24.53 0.3655 StyleCrafter 19.37 58.68 31.6 0.3194
[0120] Table 1 shows a quantitative comparison of the text and image style transfer performance between the present invention and the prior art, wherein the present invention exhibits superior performance in all indicators.
[0121] In summary, the present invention proposes a method for embedding and reconstructing text-image alignment style transfer. Compared with the traditional technology, the important innovative contributions of the present invention include:
[0122] The design and implementation process of the style extraction module based on perceptron attention and multi-layer architecture, including key links such as latent variable processing, attention calculation, and feedforward network fusion.
[0123] The cross-attention mechanism architecture, matrix transformation process and weight calculation optimization strategy of the text-image alignment enhancement module ensure efficient interactive fusion of text-image embeddings in the shared feature space.
[0124] The linear interpolation fusion rules of the explicit modulation module, the fusion ratio control mechanism and the embedding and splicing method with the diffusion model realize the flexible fusion of multimodal embedding and original embedding to balance content and style.
[0125] The application scenarios of the present invention include but are not limited to:
[0126] Artistic creation assistance: Provides various style templates and intelligent style conversion tools for digital painting and illustration design to inspire artists' creativity and improve creative efficiency and style diversity.
[0127] Film and television special effects production: applied in the stylized processing of film and television scenes and character special effects, accurately transferring the style according to the needs of the plot and character creation, and enhancing the visual narrative and artistic appeal.
[0128] Virtual fashion design: Help fashion designers quickly build virtual clothing series, integrate multiple fashion styles and popular elements, realize the visualization and diversification of design concepts, and innovate the fashion creation process.
[0129] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.
[0130] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.
[0131] An embodiment of the present invention further provides a processor, wherein the processor executes a computer program and at least executes the method described above.
[0132] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a ferromagnetic random access memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0133] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0134] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0135] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0136] A person skilled in the art can understand that: all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, etc. Various media that can store program codes.
[0137] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0138] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0139] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0140] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0141] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art of the present invention, several equivalent substitutions or obvious variations can be made without departing from the concept of the present invention, and the performance or use is the same, which should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for embedding and reconstructing text-image alignment style transfer, characterized in that: The following steps are involved: S1. Basic model construction: The diffusion model is used as the basic architecture. The model gradually adds noise to the image data through the forward process, simulating the gradual transformation of the image from a clear state to a completely noisy state. Then, the original image is restored from the noisy image through the denoising process to achieve the learning and encoding of image style features. S2. Attention-based style extraction: For reference images in the artwork dataset, the input image embedding is extracted through a feature extraction network to capture the visual features of the image; the image embedding is then further processed using the perceptron attention mechanism and the feed-forward network FFN to capture complex style details and generate a style embedding; S3. Text embedding generation: The text description is converted into text embedding through a text encoder to guide the style transfer process; S4. Text-image alignment enhancement: The style embedding is fused with the text embedding through a cross-attention mechanism to generate a multimodal embedding to optimize the interaction between text and image information; S5. Integration of explicit modulation and diffusion model: The multimodal embedding is fused with the style embedding through linear interpolation, and the fused image embedding is concatenated with the text embedding to form a complete prompt embedding, which is then integrated into the diffusion model to generate diversified style images that match the text description.
2. The method for embedding and reconstructing text-image alignment style transfer as claimed in claim 1, characterized in that: The diffusion model uses U-Net to implement the denoising process and is trained through mean square error loss.
3. The method for embedding and reconstructing text-image alignment style transfer as described in claim 1 or 2, characterized in that: Step S2 specifically includes: Generate initial latent variables for the reference image, which are normalized to stabilize the training process; The latent variables are expanded to match the batch size of the input images to maintain consistency when processing multiple images; Updating the latent variables through the perceptron attention mechanism enables the model to selectively focus on different parts of the input image based on the learnable latent variables; The updated latent variables are transformed nonlinearly through the feedforward network FFN to enhance the style encoding capability; The output of the feed-forward network and the updated latent variables are combined to generate an embedding representing the style of the input image for subsequent text-image alignment enhancement.
4. The method for embedding and reconstructing text-image alignment style transfer as claimed in claim 3, characterized in that: In step S2, the visual part of the CLIP model is used as a feature extraction network to obtain an embedded representation of the input image as a reference basis for style extraction; the output of the perceptron attention mechanism is further processed by a position-by-position feed-forward network FFN to capture more fine-grained style features; the output of the FFN is combined with the updated latent variables to obtain the final style embedding.
5. The method for embedding and reconstructing text-image alignment style transfer as claimed in claim 3 or 4, characterized in that: The feedforward network FFN includes at least two linear transformation steps for weighting latent variables, and introduces a nonlinear activation function between the two linear transformations to enhance the model's ability to learn complex features.
6. The method for embedding and reconstructing text-image alignment style transfer according to any one of claims 1 to 5, characterized in that: Step S4 specifically includes: The style embedding and text embedding are converted into query, key, and value matrices respectively through a linear layer to interact in a shared feature space; The attention weights are computed using the dot product of the query matrix and the key matrix and scaled to prevent gradient vanishing, thereby dynamically prioritizing different aspects of the text prompt. Apply the softmax function to the raw attention scores to obtain normalized attention weights, ensuring that the sum of the weights is 1 for effective information integration; The value matrix is weighted summed using normalized attention weights to generate a multimodal embedding that can more effectively capture the multimodal context of text and images.
7. The method for embedding and reconstructing text-image alignment style transfer according to any one of claims 1 to 6, characterized in that: Step S5 specifically includes: The style embedding is fused with the multimodal embedding obtained by text-image alignment enhancement through linear interpolation to achieve a smooth transition and effective combination between the two embeddings. Use predefined constants to control the fusion ratio between style embedding and multimodal embedding, so that the flexibility of fusion can be adjusted according to different application scenarios and requirements; The fused image embedding is concatenated with the text prompt embedding to form a complete prompt embedding for image generation, which integrates the multimodal information of text and image. Integrating the generated full prompt embedding into the diffusion model enables the model to simultaneously consider the multimodal conditions of text and image during the generation process, improving the performance and quality of generated images.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for embedding and reconstructing text-image alignment style transfer as described in any one of claims 1 to 7 is implemented.
9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for embedding and reconstructing text-image alignment style transfer as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Garment style fusion method and system based on diffusion model
CN117315417A
Multi-mode hair style migration generation method and system
CN117593177A
Image style migration method based on lightweight Vision Transform network
CN118279131A
Cited By
Unsupervised anomaly detection method and system based on diffusion model, and readable storage medium
CN120162727A
Method for quickly generating stylized text to image based on diffusion model
CN120543365A
Bidirectional fusion image style migration method and system based on content and style decoupling
CN120765451A
Bidirectional fusion image style transfer method and system based on content and style decoupling
CN120765451B
Style migration method and system based on text inversion and self-attention injection
CN121033214A