Text-driven knitted product image generation and editing method and device
By building a text-driven knitted product image generation and editing model, integrating image generation and editing tasks, and using visual language models and image diffusion modules to realize multimodal feature fusion and editing instruction execution, the problem of lack of a unified framework in the existing technology is solved, and the generalization ability and user experience of the model are improved.
Patent Information
- Application Number
- CN202510680668.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The existing knitted product image generation and editing methods lack a unified framework, and it is difficult to support image generation and editing at the same time, and cannot meet the market needs of personalized design and rapid iteration.
By building a text-driven knitted product image generation and editing model, the image generation and editing tasks are integrated into one model, the visual language model encoder is used to fusion of multimodal features, and the precise understanding and execution of editing instructions is achieved through image diffusion module, multi-level residual upsampling and downsampling.
It realizes a unified framework for image generation and editing of knitted products, simplifies model deployment, improves the generalization ability and convenience of use of models, can more accurately understand user intentions, generate images that conform to text descriptions, and execute editing instructions without destroying the image structure and details.
Smart Images

Figure CN120198548A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of computer graphics, and in particular to a text-driven knitting product image generation and editing method and device. Background Art
[0002] With the accelerated advancement of digital transformation in the textile and fashion industry, traditional knitted product design can no longer meet the personalized needs of consumers. The traditional knitted product design process usually relies on manual drawing or professional design software, which is not only time-consuming and laborious, but also requires high professional skills of designers, limiting the speed and scope of design innovation. In addition, in practical applications, consumers want to generate new knitted product images based on text descriptions, or make partial modifications to existing images (such as changing colors, adjusting textures, etc.), while existing image generation or editing methods mostly focus on the optimization of a single task and lack a unified framework to support image generation and editing at the same time.
[0003] Therefore, exploring an efficient and accurate unified framework for knitted product image generation and editing to meet the market demand for personalized design and rapid iteration has become a key issue that needs to be urgently addressed in the current textile and fashion industry. Summary of the invention
[0004] In view of the above problems, the present invention proposes a text-driven knitted product image generation and editing method and device, which simplifies the complexity of model deployment and improves the generalization ability and ease of use of the model by integrating the knitted product image generation and editing tasks into one model; uses a visual language model encoder to perform high-dimensional semantic joint modeling of knitted product text, input image and editing mask, and realizes deep fusion of multimodal features; uses an image diffusion module, multi-level residual upsampling and multi-level residual downsampling, which can achieve accurate understanding and execution of editing instructions without destroying the overall structure and details of the image.
[0005] On the one hand, the text-driven knitted product image generation and editing method has the following specific steps:
[0006] S1, obtaining knitted product text, an edit mask and an input image, and splicing the edit mask and the input image to obtain a spliced image;
[0007] S2, constructing and training a text-driven knitting product image generation and editing model including a multimodal fusion feature network, a multi-layer denoising network and a knitting image reconstruction network to obtain a trained knitting product image generation and editing model;
[0008] The multi-modal fusion feature network includes a vision-language model encoder, a variational autoencoder, and a multi-level residual downsampling module; the input knitted product text and spliced image are fused by the vision-language model encoder to obtain initial multi-modal features; the input image is encoded by the variational autoencoder and Gaussian noise is embedded to obtain initial latent features, and the initial latent features are subjected to multi-scale feature fusion through the multi-level residual downsampling module to obtain multi-level initial latent features; the multi-modal features and the multi-level initial latent features are concatenated along the channel dimension to obtain and output multi-modal fusion features.
[0009] The multi-layer denoising network includes several denoising network layers, and each denoising network layer includes a first image diffusion module and a second image diffusion module; the first layer of the multi-layer denoising network takes the multi-modal fusion features as the input, and the remaining layers take the output of the previous layer as the input, and the last layer outputs the final denoised multi-modal fusion features to the knitted image reconstruction network; the first image diffusion module denoises the input to obtain preliminarily denoised multi-modal fusion features, and the preliminarily denoised multi-modal fusion features are segmented along the channel dimension to obtain intermediate multi-modal features and intermediate latent features; the intermediate multi-modal features and the intermediate latent features after time step embedding are concatenated along the channel dimension, and then denoised by the second image diffusion module to output the denoised multi-modal fusion features.
[0010] The knitted image reconstruction network segments the final denoised multi-modal fusion features to obtain denoised latent features; the denoised latent features are input into the multi-level residual upsampling module for knitted image reconstruction, and a knitted product image is output.
[0011] S3. Input the knitted product text, edit mask, and input image into the trained text-driven knitted product image generation and editing model to obtain a knitted product image.
[0012] Preferably, the text-driven knitted product image generation and editing model uses time step embedding, which acts on the initial latent features, multi-level residual downsampling module, first image diffusion module, intermediate latent features, second image diffusion module, denoised latent features, and multi-level residual upsampling module respectively.
[0013] Preferably, the time step embedding is expressed as:
[0014] ;
[0015] where represents the time step embedding, represents the feature after time step embedding; represents the feature to be embedded; represents the time step noise attenuation coefficient.
[0016] Preferably, the time-step noise attenuation coefficient is expressed as:
[0017] ;
[0018] where, represents the current time step.
[0019] Preferably, the Gaussian noise embedding is expressed as:
[0020] ;
[0021] where, represents the feature after Gaussian noise embedding; represents the feature to be embedded; represents the standard Gaussian noise; represents the Gaussian noise attenuation coefficient.
[0022] Preferably, the Gaussian noise attenuation coefficient is expressed as:
[0023] ;
[0024] where T represents the total number of steps and t represents the current time step.
[0025] Preferably, the multi-level residual downsampling module is expressed as:
[0026] ;
[0027] where, represents the enhanced initial latent feature; represents the operation repeated N times; represents the convolution operation with a 3×3 convolution kernel and a stride of 2; represents the ReLU activation function; represents the normalization operation; represents the convolution operation with a 3×3 convolution kernel and a stride of 1; represents the initial latent feature.
[0028] Preferably, the multi-level residual upsampling module is expressed as:
[0029] ;
[0030] where, represents the enhanced denoised latent feature; represents the operation repeated N times; represents the transposed convolution operation with a 3×3 convolution kernel and a stride of 2; represents the ReLU activation function; represents the normalization operation; Represents a convolution operation with a 3×3 convolution kernel and a stride of 1; Represents the denoised latent features.
[0031] Preferably, the first image diffusion module and the second image diffusion module have the same structure, and their respective processing procedures are as follows:
[0032] First, the multi-modal fusion features input into the image diffusion module are input into the multi-head self-attention layer to obtain the enhanced multi-modal fusion features after the multi-head self-attention layer, which is expressed as:
[0033] ;
[0034] Among them, Represents the initial multi-modal fusion features; Represents the enhanced multi-modal fusion features after the multi-head self-attention layer; Represents the concatenation operation; Represents the multi-head attention layer; Represents the ReLU activation function; Represents the normalization operation; Represents the time step embedding;
[0035] Then, the enhanced multi-modal fusion features after the multi-head self-attention layer are input into the forward feedback layer, and the enhanced multi-modal fusion features after the forward feedback layer are output; which is expressed as:
[0036] ;
[0037] Among them, Represents the enhanced multi-modal fusion features after the forward feedback layer, Represents the forward feedback layer.
[0038] On the other hand, a text-driven knitted product image generation and editing device includes the following:
[0039] A knitted product data acquisition module, which is used to acquire knitted product texts, editing masks, and input images, and splice the knitted product editing masks and input images to obtain a spliced image;
[0040] A model construction and training module, which is used to construct a text-driven knitted product image generation and editing model including a multi-modal fusion feature network, a multi-layer denoising network, and a knitted image reconstruction network and perform training to obtain a trained text-driven knitted product image generation and editing model;
[0041] The multi-modal fusion feature network includes a vision-language model encoder, a variational autoencoder, and a multi-level residual downsampling module; the input knitted product text and spliced image are fused by the vision-language model encoder to obtain initial multi-modal features; the input image is encoded by the variational autoencoder and Gaussian noise is embedded to obtain initial latent features, and the initial latent features are subjected to multi-scale feature fusion through the multi-level residual downsampling module to obtain multi-level initial latent features; the multi-modal features and the multi-level initial latent features are spliced along the channel dimension to obtain and output multi-modal fusion features;
[0042] The multi-layer denoising network includes several denoising network layers, and each denoising network layer includes a first image diffusion module and a second image diffusion module; the first layer of the multi-layer denoising network takes the multi-modal fusion features as input, and the remaining layers take the output of the previous layer as input, and the last layer outputs the final denoised multi-modal fusion features to the knitted image reconstruction network; the first image diffusion module denoises the input to obtain preliminarily denoised multi-modal fusion features, and divides the preliminarily denoised multi-modal fusion features along the channel dimension to obtain intermediate multi-modal features and intermediate latent features; the intermediate multi-modal features and the intermediate latent features after time step embedding are spliced along the channel dimension, and then denoised by the second image diffusion module to output denoised multi-modal fusion features;
[0043] The knitted image reconstruction network divides the final denoised multi-modal fusion features to obtain denoised latent features; the denoised latent features are input into the multi-level residual upsampling module for knitted image reconstruction, and a knitted product image is output;
[0044] The text-driven knitted product image generation and editing model uses time step embedding, which acts on the initial latent features, the multi-level residual downsampling module, the first image diffusion module, the intermediate latent features, the second image diffusion module, the denoised latent features, and the multi-level residual upsampling module respectively;
[0045] The knitted product image generation module inputs the knitted product text, the editing mask, and the input image into the trained text-driven knitted product image generation and editing model to obtain a knitted product image.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] (1) The present invention constructs a unified knitted product image generation and editing framework, integrating the image generation and editing tasks into one model, which not only simplifies the complexity of model deployment, but also significantly improves the generalization ability and usability of the model, greatly improving work efficiency and user experience;
[0048] (2) The present invention uses a visual language model encoder to perform high-dimensional semantic joint modeling on the knitted product text, input image, and editing mask, achieving deep fusion of multi-modal features; this fusion method enables the model to more accurately understand the user's intention and generate a knitted product image that better conforms to the text description.
[0049] (3) Through the design of an image diffusion module and a multi-level residual network including multi-level residual upsampling and multi-level residual downsampling, the present invention can accurately understand and execute editing instructions without damaging the overall structure and details of the image, such as changing the color of the knitted product image, adjusting the texture, etc., which is of great significance for meeting the personalized needs of consumers. Brief Description of the Drawings
[0050] The following further describes the present invention in detail with reference to the drawings.
[0051] Figure 1 is a flowchart of the text-driven knitted product image generation and editing method according to an embodiment of the present invention.
[0052] Figure 2 is a schematic structural diagram of the knitted product image generation and editing model of the text-driven knitted product image generation and editing method according to an embodiment of the present invention.
[0053] Figure 3 is a schematic structural diagram of the image diffusion module of the text-driven knitted product image generation and editing method according to an embodiment of the present invention.
[0054] Figure 4 is a schematic diagram of the knitted product image generation and editing functions of the text-driven knitted product image generation and editing method according to an embodiment of the present invention.
[0055] Figure 5 is a structural block diagram of the text-driven knitted product image generation and editing device according to an embodiment of the present invention. Detailed Embodiments
[0056] The following further describes the present invention through specific embodiments.
[0057] As Figure 1 shown, the text-driven knitted product image generation and editing method includes the following specific steps:
[0058] S1, obtain the knitted product text, editing mask, and input image, and splice the knitted product editing mask and the input image to obtain a spliced image.
[0059] The knitted product text refers to the natural language description used to describe information such as the image features, styles, colors, textures, and patterns of knitted products; the image is a reference image or the source image to be edited; the mask is a matrix with the same size as the input image, serving the editing task and used to define the editing area, and the element values in the matrix are 0 or 1, where 1 represents the editing area.
[0060] S2. Construct a text-driven knitted product image generation and editing model including a multi-modal fusion feature network, a multi-layer denoising network, and a knitted image reconstruction network, and train it to obtain a trained knitted product image generation and editing model.
[0061] See Figure 2 As shown, the knitted product image generation and editing model includes a multi-modal feature fusion network, a multi-layer denoising network, and a knitted image reconstruction network, which are respectively used for high-dimensional semantic joint modeling, layer-by-layer denoising processing, and knitted image reconstruction; the multi-modal feature fusion network includes a vision-language model encoder (Vision-Language Model, VLM, or VLM encoder for short), initial multi-modal features, a variational auto-encoder (Vibrational Auto-Encoder, VAE, or VAE encoder for short), initial latent features, and a multi-level residual downsampling module; the multi-layer denoising network includes an image diffusion module (DiffusionImage Transformer, DIT, or DIT module for short), intermediate multi-modal features, and intermediate latent features, and the multi-layer denoising network includes several layers, and each layer includes two diffusion modules; the knitted image reconstruction network includes denoised multi-modal features, denoised latent features, a multi-level residual upsampling module, and a knitted product image output.
[0062] The multi-modal feature fusion network receives text input, image input, and mask input (i.e., the knitted product text, the editing mask, and the input image), and performs high-dimensional semantic joint modeling on these input information through the vision-language model encoder to obtain initial multi-modal features; encodes the image using the variational auto-encoder and embeds Gaussian noise to obtain initial latent features; performs multi-scale feature fusion on the initial latent features through the multi-level residual downsampling module to obtain multi-level initial latent features; concatenates the initial multi-modal features and the multi-level initial latent features along the channel dimension to form multi-modal fusion features. The multi-layer denoising network receives the multi-modal fusion features and uses the image diffusion module to perform layer-by-layer denoising processing on the multi-modal fusion features; the image diffusion module consists of a multi-head self-attention layer and a forward feedback layer, and realizes fine denoising of the multi-modal fusion features through time-step embedding; after multiple denoising processes, denoised multi-modal features and denoised latent features are obtained. The knitted image reconstruction network receives the denoised latent features and gradually restores the image details through the multi-level residual upsampling module, and finally outputs a high-resolution knitted product image.
[0063] In a specific embodiment, the vision-language model encoder selects the existing Qwen-VL 7B model, which supports multi-resolution image input (256×256 to 1024×1024) and can better perform fine-grained semantic parsing on knitted products (such as "changing the round neck to a V-neck"); the variational autoencoder adopts the existing FLUX-schnell architecture, and the number of channels in the latent space is expanded to 128 to retain the knitted texture details (such as ribbing, cable knitting, etc.); the image diffusion module has 12 layers and supports the input of long text instructions, and can converge within a short inference step length.
[0064] In a specific embodiment, during the encoding stage of the variational autoencoder, temporally correlated Gaussian noise is injected into the latent representation of the input image I. The formula for embedding Gaussian noise is as follows:
[0065] ;
[0066] where represents the feature after Gaussian noise embedding of X; represents the feature to be embedded; represents the standard Gaussian noise, , with the dimension consistent with the latent space (default 256×256); represents the Gaussian noise attenuation coefficient, following a cosine schedule, , T represents the total number of steps, and t represents the current time step.
[0067] In a specific embodiment, the knitted product image has overall structural stability and coherence, as well as complex local textures, which determine the task characteristics of its generation process from the whole to the part. Therefore, according to the task characteristics of knitted product generation, the time step embedding adopts a phased dynamic scheduling strategy; the formula for embedding time step noise is expressed as:
[0068] ;
[0069] where represents the feature after time step embedding of X; represents the time step noise attenuation; X represents the feature to be embedded.
[0070] Specifically, the time step embedding realizes the progressive embedding of noise through a cosine schedule. In the initial stage , low-intensity base noise is retained to ensure the high-frequency textures of the input image (such as ribbing, cable knitting, etc.), and it smoothly transitions to the middle stage through a cosine schedule; in the middle stage , a scaling factor of 0.8 is used to control the noise attenuation rate, and the noise intensity gradually decreases according to a cosine schedule, and it is ensured that at Decays to zero at time; in the later stage ,Completely stop noise injection, and the residual upsampling module restores high-resolution details. The specific parameter configuration of the time-step noise decay is as follows:
[0071] ;
[0072] Among them, Represents the time-step noise decay coefficient, Represents the time step.
[0073] During the downsampling process, the time-step embedding is used to adjust the intensity of the noise in order to extract image features at different scales. The image diffusion module uses the time-step embedding through the multi-head self-attention layer and the forward feedback layer to achieve layer-by-layer denoising of the multi-modal fusion features. The time-step embedding is used in the diffusion module to control the denoising process, gradually reducing the noise and restoring the clarity and details of the image. In the upsampling stage, the time-step embedding is used to control the process of restoring image details. The time-step embedding ensures that noise injection completely stops in the later stage, and the residual upsampling module can effectively restore the high-resolution details of the image.
[0074] Visual language model encoder. In a specific embodiment, the knitted product text, edit mask, and input image are input into the trained knitted product image generation and editing model. The visual language model encoder processes the input information to obtain the initial multi-modal features of the input information, specifically as follows:
[0075] Using the visual language model encoder to perform high-dimensional semantic joint modeling on the input text, image, and mask to obtain the initial multi-modal features of the input information, expressed as:
[0076] ;
[0077] Among them, Represents the initial multi-modal features of the input information, Represents the visual language model encoder, Represents the input text information, Represents the input image information, Represents the input mask information, Represents the concatenation operation.
[0078] Variational autoencoder. In a specific embodiment, the image is input into the variational autoencoder and Gaussian noise is embedded to obtain the initial latent features; the multi-level residual downsampling module performs multi-scale feature fusion on the initial latent features to obtain the multi-level initial latent features, specifically as follows:
[0079] Using the variational autoencoder to encode the image and embed Gaussian noise to obtain the initial latent features of the input image, expressed as:
[0080] ;
[0081] Among them, represents the initial latent feature of the input image, represents the Gaussian noise embedding, represents the variational autoencoder.
[0082] Multilevel residual downsampling module. The initial latent feature of the input image is input into the multilevel residual downsampling module for multi-scale feature fusion to obtain the multilevel initial latent feature, which is expressed as:
[0083] ;
[0084] Among them, represents the multilevel initial latent feature, represents the multilevel residual downsampling module, represents the noise embedding at the current time step.
[0085] In a specific embodiment, a multilevel residual downsampling module is composed of N repeated residual convolutional layers and downsampling layers. The initial latent feature of the input image is input into the multilevel residual downsampling module to obtain the enhanced initial latent feature of the input image, which is expressed as:
[0086] ;
[0087] Among them, represents the enhanced initial latent feature of the input image, represents the operation repeated N times, represents the convolutional operation with a 3×3 convolutional kernel and a stride of 2, represents the ReLU activation function, represents the normalization operation, represents the convolutional operation with a 3×3 convolutional kernel and a stride of 1.
[0088] In a specific embodiment, the multi-modal features of the input information and the multilevel initial latent features are concatenated along the channel dimension to obtain the multi-modal fusion feature, which is expressed as:
[0089] ;
[0090] Among them, represents the multi-modal fusion feature, represents the concatenation operation.
[0091] In a specific embodiment, the multimodal fusion features are input into the image diffusion module to obtain the multimodal fusion features for preliminary denoising, segmented along the channel dimension to obtain the intermediate multimodal features and the intermediate potential features, the intermediate multimodal features and the intermediate potential features after time step embedding are spliced along the channel dimension, and then segmented after passing through the image diffusion module to obtain the denoised multimodal features and the denoised potential features, which specifically include:
[0092] The multimodal fusion features are input into the image diffusion module to obtain the multimodal fusion features for preliminary denoising, as shown below:
[0093] ;
[0094] in, represents the multimodal fusion features of preliminary denoising, Represents the image diffusion module.
[0095] The multimodal fusion features of the preliminary denoising are split along the channel dimension to obtain intermediate multimodal features and intermediate latent features, as shown below:
[0096] ;
[0097] in, represents the intermediate multimodal features, represents the intermediate latent feature, Represents a matrix split operation.
[0098] The intermediate multimodal features and the intermediate latent features after time step embedding are concatenated along the channel dimension, and then segmented along the channel dimension after passing through the image diffusion module to obtain the denoised multimodal features and denoised latent features, as shown below:
[0099] ;
[0100] in, represents the denoising multimodal feature, represents the denoised latent features.
[0101] In a specific embodiment, reference Figure 3 ,The image diffusion module consists of a multi-head self-attention layer and a forward feedback layer, and uses time step embedding to achieve layer-by-layer denoising of multimodal fusion features, as follows:
[0102] First, the multimodal fusion features are input into the multi-head self-attention layer. The specific calculation is as follows:
[0103] ;
[0104] in, represents the initial multimodal fusion features, Represents the enhanced multimodal fusion feature passing through the multi-head self-attention layer, represents the multi-head attention layer; represents the ReLU activation function.
[0105] Then, the enhanced multimodal fusion feature is input into the forward feedback layer, and the specific calculation is as follows:
[0106] ;
[0107] Among them, represents the enhanced multimodal fusion feature passing through the forward feedback layer, represents the forward feedback layer; represents the ReLU activation function.
[0108] In a specific embodiment, the denoised latent feature is input into the multi-level residual upsampling module for knitting image reconstruction, and finally a high-resolution knitting image is obtained, denoted as:
[0109] ;
[0110] Among them, represents the finally obtained high-resolution knitting image, represents the multi-level residual upsampling module.
[0111] In a specific embodiment, a multi-level residual upsampling module is composed of N repeated residual convolutional layers and upsampling layers. The denoised latent feature is input into the multi-level residual upsampling module to obtain an enhanced denoised latent feature, as shown below:
[0112] ;
[0113] Among them, represents the enhanced denoised latent representation, represents the transposed convolution operation with a convolution kernel of 3×3 and a stride of 2.
[0114] S3. Input the knitting product text, knitting product editing mask, and input image into the trained text-driven knitting product image generation and editing model to obtain a knitting product image.
[0115] As Figure 4 shown, the knitting product image generation and editing functions are implemented as follows:
[0116] The editing function is realized by combining the knitting product text, knitting product editing mask, and input image. Specifically: The mask is used to specify the area in the image that needs to be edited, clarifying the scope of the editing operation. The text instruction provides specific information on how to edit this area, such as color change, texture adjustment, etc. requirements.
[0117] If only the text description is input without providing the original image, the model recognizes it as an image generation task and generates a completely new knitted product image from scratch according to the text description. If both the text description and the original knitted product image are input, when no mask is input, the model tends to perform the image generation operation. When a mask is provided, the model will perform the image editing operation by combining the mask information, because the presence of the mask usually means that specific regions of the image need to be modified.
[0118] As Figure 5 shown, the present invention also discloses a text-driven knitted product image generation and editing device, including:
[0119] A knitted product data acquisition module 501, configured to acquire knitted product text, an editing mask, and an input image, and splice the knitted product editing mask and the input image to obtain a spliced image;
[0120] A model construction and training module 502, configured to construct and train a text-driven knitted product image generation and editing model including a multi-modal fusion feature network, a multi-layer denoising network, and a knitted image reconstruction network to obtain a trained knitted product image generation and editing model;
[0121] The multi-modal fusion feature network includes a vision-language model encoder, a variational autoencoder, and a multi-level residual downsampling module; the input knitted product text and the spliced image are fused by the vision-language model encoder to obtain an initial multi-modal feature; the input image is encoded by the variational autoencoder and Gaussian noise is embedded to obtain an initial latent feature, and the initial latent feature undergoes multi-scale feature fusion through the multi-level residual downsampling module to obtain a multi-level initial latent feature; the multi-modal feature and the multi-level initial latent feature are spliced along the channel dimension to obtain a multi-modal fusion feature and output it;
[0122] The multi-layer denoising network includes several layers of denoising network layers, and each layer of denoising network layer includes a first image diffusion module and a second image diffusion module; the first layer of the multi-layer denoising network takes the multi-modal fusion feature as the input, and the remaining layers take the output of the previous layer as the input, and the last layer outputs the final denoised multi-modal fusion feature to the knitted image reconstruction network; the first image diffusion module denoises the input to obtain a preliminarily denoised multi-modal fusion feature, divides the preliminarily denoised multi-modal fusion feature along the channel dimension to obtain an intermediate multi-modal feature and an intermediate latent feature; the intermediate multi-modal feature and the intermediate latent feature after time step embedding are spliced along the channel dimension, and then denoised by the second image diffusion module to output the denoised multi-modal fusion feature;
[0123] The knitting image reconstruction network segments the final denoised multi-modal fusion features to obtain denoised latent features; the denoised latent features are input into a multi-level residual upsampling module for knitting image reconstruction, and a knitted product image is output;
[0124] The knitted product image generation module 503 inputs the knitted product text, the editing mask, and the input image into the trained text-driven knitted product image generation and editing model to obtain a knitted product image.
[0125] The specific implementation of the text-driven knitted product image generation and editing device is the same as that of the text-driven knitted product image generation and editing method, and will not be repeated in this embodiment.
[0126] The above is only the specific implementation manner of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantive modification made to the present invention using this concept shall fall within the scope of infringement of the protection scope of the present invention.
Claims
1. A text-driven knitting product image generation and editing method, characterized in that It includes the following steps: S1. Obtain the knitted product text, the editing mask, and the input image, and splice the editing mask and the input image to obtain a spliced image; S2. Construct a text-driven knitted product image generation and editing model including a multi-modal fusion feature network, a multi-layer denoising network, and a knitted image reconstruction network, and train it to obtain a trained text-driven knitted product image generation and editing model; The multi-modal fusion feature network includes a vision-language model encoder, a variational auto-encoder, and a multi-level residual downsampling module; the input knitted product text and the spliced image are fused by the vision-language model encoder to obtain an initial multi-modal feature; The input image is encoded by the variational auto-encoder and embedded with Gaussian noise to obtain an initial latent feature, and the initial latent feature undergoes multi-scale feature fusion through the multi-level residual downsampling module to obtain a multi-level initial latent feature; The multi-modal feature and the multi-level initial latent feature are spliced along the channel dimension to obtain a multi-modal fusion feature and output it; The multi-layer denoising network includes several denoising network layers, and each denoising network layer includes a first image diffusion module and a second image diffusion module; the first layer of the multi-layer denoising network takes the multi-modal fusion feature as the input, and the remaining layers take the output of the previous layer as the input, and the last layer outputs the final denoised multi-modal fusion feature to the knitted image reconstruction network; The first image diffusion module denoises the input to obtain a preliminarily denoised multi-modal fusion feature, and divides the preliminarily denoised multi-modal fusion feature along the channel dimension to obtain an intermediate multi-modal feature and an intermediate latent feature; The intermediate multi-modal feature and the intermediate latent feature after time step embedding are spliced along the channel dimension, and then denoised by the second image diffusion module to output a denoised multi-modal fusion feature; The knitted image reconstruction network divides the final denoised multi-modal fusion feature to obtain a denoised latent feature; inputs the denoised latent feature into a multi-level residual upsampling module for knitted image reconstruction, and outputs a knitted product image; S3. Input the knitted product text, the editing mask, and the input image into the trained text-driven knitted product image generation and editing model to obtain a knitted product image.
2. The text-driven knitting product image generation and editing method according to claim 1, characterized in that The text-driven knitted product image generation and editing model uses time step embedding, which acts on the initial latent feature, the multi-level residual downsampling module, the first image diffusion module, the intermediate latent feature, the second image diffusion module, the denoised latent feature, and the multi-level residual upsampling module respectively.
3. The text-driven knitting product image generation and editing method according to claim 2, characterized in that The time step embedding is expressed as: ; Among them, represents the time step embedding, represents the feature after time step embedding; represents the feature to be embedded; represents the time step noise attenuation coefficient.
4. The method for generating and editing a text-driven knitted product image according to claim 3, characterized in that, The time step noise attenuation coefficient is expressed as: ; Among them, represents the current time step.
5. The method for generating and editing a text-driven knitted product image according to claim 2, wherein The Gaussian noise embedding is expressed as: ; Among them, represents the feature after Gaussian noise embedding; represents the feature to be embedded; represents the standard Gaussian noise; represents the Gaussian noise attenuation coefficient.
6. The text-driven knitting product image generation and editing method according to claim 5, characterized in that, The Gaussian noise attenuation coefficient is expressed as: ; Where T represents the total number of steps, and t represents the current time step.
7. The text-driven knitting product image generation and editing method according to claim 1, characterized in that The multi-level residual downsampling module is expressed as: ; Among them, represents the enhanced initial latent feature; represents the operation repeated N times; represents a convolution operation with a 3×3 convolution kernel and a stride of 2; represents the ReLU activation function; represents the normalization operation; represents a convolution operation with a 3×3 convolution kernel and a stride of 1; represents the initial latent feature.
8. The method for generating and editing a text-driven knitted product image according to claim 1, wherein The multi-level residual upsampling module is expressed as: ; Among them, represents the enhanced denoised latent feature; represents the operation repeated N times; represents the transposed convolution operation with a convolution kernel of 3×3 and a stride of 2; represents the ReLU activation function; represents the normalization operation; represents the convolution operation with a convolution kernel of 3×3 and a stride of 1; represents the denoised latent feature.
9. The text-driven knitting product image generation and editing method according to claim 2, characterized in that The structures of the first image diffusion module and the second image diffusion module are the same, and their specific processing processes are as follows: First, input the multi-modal fusion feature of the input image diffusion module into the multi-head self-attention layer to obtain an enhanced multi-modal fusion feature after the multi-head self-attention layer, expressed as: ; Among them, represents the initial multi-modal fusion feature; represents the enhanced multi-modal fusion feature after passing through the multi-head self-attention layer; represents the concatenation operation; represents the multi-head attention layer; represents the ReLU activation function; represents the normalization operation; represents the time step embedding; Then, the enhanced multi-modal fusion features that have passed through the multi-head self-attention layer are input into the forward feedback layer, and the enhanced multi-modal fusion features that have passed through the forward feedback layer are output; expressed as: ; Among them, represents the enhanced multi-modal fusion feature after passing through the forward feedback layer, represents the forward feedback layer.
10. A text-driven knitting product image generation and editing device, including the following: A knitting product data acquisition module, configured to acquire knitting product text, an editing mask, and an input image, and splice the knitting product editing mask and the input image to obtain a spliced image; A model construction and training module, configured to construct and train a text-driven knitting product image generation and editing model including a multi-modal fusion feature network, a multi-layer denoising network, and a knitting image reconstruction network to obtain a trained knitting product image generation and editing model; The multi-modal fusion feature network includes a vision-language model encoder, a variational autoencoder, and a multi-level residual downsampling module; the input knitted product text and spliced image are fused by the vision-language model encoder to obtain initial multi-modal features; The input image is encoded by a variational autoencoder and Gaussian noise is embedded to obtain initial latent features, and the initial latent features are subjected to multi-scale feature fusion through a multi-level residual downsampling module to obtain multi-level initial latent features; The multi-modal features and the multi-level initial latent features are spliced along the channel dimension to obtain and output multi-modal fusion features; The multi-layer denoising network includes several denoising network layers, and each denoising network layer includes a first image diffusion module and a second image diffusion module; the first layer of the multi-layer denoising network takes the multi-modal fusion features as input, and the remaining layers take the output of the previous layer as input, and the last layer outputs the final denoised multi-modal fusion features to the knitting image reconstruction network; The first image diffusion module denoises the input to obtain preliminarily denoised multi-modal fusion features, and divides the preliminarily denoised multi-modal fusion features along the channel dimension to obtain intermediate multi-modal features and intermediate latent features; The intermediate multi-modal features and the intermediate latent features after time step embedding are spliced along the channel dimension, and then denoised by the second image diffusion module to output denoised multi-modal fusion features; The knitting image reconstruction network divides the final denoised multi-modal fusion features to obtain denoised latent features; the denoised latent features are input into a multi-level residual upsampling module for knitting image reconstruction, and a knitting product image is output; A knitting product image generation module, which inputs the knitting product text, the editing mask, and the input image into the trained text-driven knitting product image generation and editing model to obtain a knitting product image.
Citation Information
Patent Citations
Method and device for generating image from text, storage medium and electronic equipment
CN118587303A
Knitwear defect detection method and system based on image processing
CN118781128A
Visual language corresponding AI generated panoramic image quality evaluation method and system
CN119919423A
Tool for controlling a knitted image embedded in amobile telecommunications terminal
KR1020030063080A
Systems and methods for subject-driven image generation
US20240161369A1