Clothing image generation method and system based on sketch and text
By introducing sketch prior embedding modules and multi-scale attention mechanisms of structural perception in clothing image generation, combined with variance-guided network simplification strategy, the problems of inaccurate outlines, poor style consistency and high model complexity are solved, and clothing image generation with high quality, stability and computing efficiency are achieved.
Patent Information
- Application Number
- CN202510292698.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-12
AI Technical Summary
The prior art has problems in the generation of clothing images that are inaccurate outlines, poor style consistency, and high model complexity.
By introducing a sketch prior embedding module and a multi-scale attention mechanism of structure perception, combined with a variance-guided network simplification strategy, the structural and semantic features in sketches and text descriptions are extracted and integrated to achieve accurate restoration of sketch outlines and style consistency of text information.
It effectively improves the quality, stability and computing efficiency of generated images, and ensures the fidelity of the details and style uniformity of generated clothing images.
Smart Images

Figure CN120219573A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a method and system for generating clothing images based on sketches and text. Background Art
[0002] In recent years, with the development of artificial intelligence (AI) technology, especially the breakthroughs in the field of deep learning, computer vision and multi-modal generation technologies have been widely applied. Generative models such as Generative Adversarial Networks (GANs) and Diffusion Models (DMs) have shown great potential in tasks such as text-to-image generation and multi-modal image generation, and have achieved remarkable progress especially in applications such as fashion design and virtual fitting. Text-to-image generation technology, as a cross-modal conversion task, has become a popular research direction in the fields of computer vision and multimedia.
[0003] In traditional fashion design, designers usually rely on hand-drawn sketches, fabric samples and manual adjustments to complete the design. However, with the development of deep generative models, AI technology has gradually changed this traditional process. Designers can automatically generate clothing images through text descriptions or sketches, which greatly improves the design efficiency and accuracy. In this context, the method of text-to-image generation has received particular attention, aiming to automatically generate images that match the content of the text description.
[0004] Existing text-to-image generation technologies, especially in fashion design, mainly face the following challenges: First, how to maintain the contour accuracy and details of the sketch during the image generation process to avoid distorted shapes or inconsistent styles in the generated images; Second, in the case of multi-modal input (such as sketches and text), how to effectively fuse text information with the image structure to ensure that the image content is consistent with the semantics of the input text and that the styles of all parts of the image are consistent. In addition, existing generative models usually have the problem of high complexity. How to reduce the complexity and computational overhead of the model while ensuring the quality of the generated images is also a technical problem to be solved. Summary of the Invention
[0005] Aiming at the problems of inaccurate generated image contours, poor style consistency and model complexity in the prior art, the present invention proposes a method and system for generating clothing images based on sketches and text. By introducing a sketch prior embedding module and a structure-aware multi-scale attention mechanism, this method can extract and fuse key structural and semantic features from sketches and text descriptions, thereby achieving accurate restoration of sketch contours and style consistency of text information. At the same time, by introducing a variance-guided network simplification strategy into the generation network, the present invention effectively improves the quality, stability and computational efficiency of the generated images, ensuring the detail fidelity and style unity of the generated clothing images.
[0006] The technical solution of the present invention is a method for generating clothing images of sketches and texts and controlling the consistency of their contours and styles, including the following steps:
[0007] Step 1, input the text information describing the clothing, and encode the input text into a feature matrix;
[0008] Step 2, input the sketch image of the clothing, and extract features of the input sketch through a sketch prior embedding module;
[0009] Step 3, fuse the input text feature matrix and the features of the sketch image, and input them into a UNet encoder to generate preliminary clothing image features;
[0010] Step 4, utilize a structure-aware multi-scale attention mechanism to further process the sketch and text features by calculating the correlation between different parts of the feature map, thereby enhancing the consistency of the contour and style of the clothing image;
[0011] Step 5, process the result obtained in Step 4 through zero convolution to extract more precise features, where zero convolution is reduced by variance-guided network simplification;
[0012] Step 6, randomly sample a Gaussian noise from a Gaussian distribution and input it as an initial noise image into another UNet architecture identical to the encoder in Step 3. At the same time, input the text features extracted in Step 1 as conditions into the UNet to guide the generation process. During the denoising process, by fusing with the result extracted in Step 5, the noise is gradually reduced and intermediate images are generated. Through multiple iterations of denoising, a clear clothing image is finally obtained, which conforms to both the contour of the input sketch and the style described in the text;
[0013] Step 7, during the denoising process, calculate the difference between the image after denoising at each stage and the target image through the mean squared error loss (MSE Loss). Specifically, the mean squared error loss measures the error between the image after denoising and the target image at the pixel level. The error gradient is transmitted back to the model through the backpropagation algorithm, thereby obtaining gradually optimized gradient information. These gradient information are used to update the weight parameters of the UNet model, enabling the model to better reduce noise and generate intermediate results closer to the target image in each step of denoising. Through multiple iterations of optimization, the UNet model gradually learns how to recover a clear clothing image that conforms to the sketch and text description from the noise image.
[0014] Further, the specific process of Step 2 is: based on the input sketch S input, first, the basic edge features are captured by the preliminary feature extraction module. This module includes a first convolutional layer to extract edge features, which are processed through batch normalization and the ReLU activation function to obtain the preliminary feature map S1. Subsequently, a second convolutional layer further refines the features to obtain the refined feature map S2. The feature map S2 captures the key features in the sketch, which is crucial for subsequent processing. This part of the structure enables the model to better capture and process the important features in the input sketch. To remove redundant information, then the global feature weighting module performs global pooling on S2 to obtain the global feature vector S gap :
[0015]
[0016] where S gap represents the output of the global average pooling layer, and S 2c represents the c-th channel of S2, and c ∈ [1, C].
[0017] And the channel weights are calculated through a fully connected layer and the Sigmoid activation function to obtain the weighted feature map S g . These weight factors are used to adjust the importance of each channel, thereby highlighting important features and suppressing unnecessary parts:
[0018] S g = σ(U2δ(U1S gap ))
[0019] where δ represents the Silu activation function, σ represents the Sigmoid function, U1 is the weight matrix used to reduce the channel dimension, U2 is the weight matrix used to restore the channel dimension, and r is the scaling parameter, aiming to reduce the number of channels and thus reduce the computational amount.
[0020] Finally, to prevent the feature information from gradually disappearing or being over-transformed in the deep network, combined with the residual connection mechanism, the weighted feature map is fused with the preliminary feature map to obtain the stable feature map S res . Through downsampling operations and convolutional processing, the final output S out provides stable structural features for subsequent image generation, and the process is as follows:
[0021] S res = S2 ⊙ S g ,
[0022]
[0023] where ⊙ represents the element-wise Hadamard product, represents the downsampling operation, β represents the normalization, represents the convolutional operation.
[0024] Further, the specific process of step 4 is as follows:
[0025] According to the features output in step 3, the features are further processed through the self-attention mechanism and the cross-attention mechanism to obtain the feature map h. The feature map h is input into the structure-aware multi-scale attention mechanism. First, the feature map is split into multiple sub-feature maps h = [h0, h1, …, h G-1 along the channel dimension, and each sub-feature map focuses on different spatial regions. Then, a global pooling operation is performed to extract global context information, obtaining the global feature dependencies h x and h y , and the importance of each channel is adjusted through 1×1 convolution and the Sigmoid activation function to obtain the globally weighted feature map h spatial :
[0026]
[0027] where d represents the global feature of the d-th channel of the feature map, d ∈ [1, K / G], h d (i, j) represents the value at position (i, j), h x and h y represent the average pooling results in the horizontal and vertical directions respectively, σ is the Sigmoid activation function, W represents the weight of the 1×1 convolution kernel, represents the concatenation operation in the channel dimension.
[0028] Meanwhile, 3×3 convolution is used to extract the local feature h local . Then, the spatial attention weights α1 and α2 are calculated through a cross-space mechanism. The cross-space mechanism enables the model to focus on key spatial regions, thereby improving the long-range dependence capture ability and enhancing feature consistency. These spatial attention weights are used to optimize the spatial features, making the generated image more coherent. This step strengthens the model's understanding of the relationships between different regions in the feature map, thus contributing to improving the overall consistency:
[0029]
[0030] α1 = Softmax(AvgPool(h local )), α2 = Softmax(AvgPool(h spatial ))
[0031] where h local represents the locally extracted feature, and W 3×3 represents the 3×3 convolution operation.
[0032] Combine the weight α1 with the globally weighted feature map h spatial, the feature map is weighted by element-wise multiplication to obtain h out1 , similarly, the weight α2 combines with the local feature map h local for element-wise multiplication to obtain h out2 , then the feature h out1 and h out2 are fused and added to obtain the feature map h out , to enhance the overall consistency and detail retention of the generated image, as follows:
[0033] h out1 = α1·h spatial , h out2 = α2·h local , h out = σ(h out2 + h out1 )
[0034] where h out is the finally output feature map, σ is the Sigmoid activation function, h out1 and h out2 respectively represent the output results of the feature maps weighted by the spatial attention weights α1 and α2.
[0035] Further, the specific process of step 5:
[0036] For the feature h out output by each encoding block in step 4, these features are further processed by zero convolution to extract features. To measure the dispersion of the output features, we calculate the variance of the features h out directly output by each layer of the encoding block, and call it the left variance. In addition, we also calculate the variance of the feature output after zero convolution processing, and define it as the right variance. We use the following formula to determine the impact of zero convolution on each encoding block:
[0037]
[0038] where σ L represents the left variance, σ R represents the right variance, η represents the variance reduction amount, which enables us to measure the variance change caused by zero convolution for each layer. For each obtained variance reduction amount η, it is sorted in ascending order, and starting from the layer with the smallest variance reduction amount, the six zero convolution layers with the smallest variance reduction amounts are retained.
[0039] The present invention also provides a clothing image generation system based on sketches and texts, including a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the clothing image generation method based on sketches and texts as described in the above technical solution.
[0040] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:
[0041] 1. Accurately retain the sketch contour information: The sketch prior embedding module designed in the present invention can better capture the key contours and structural features in the sketch, and the generated clothing image can accurately reflect the contours and details in the sketch.
[0042] 2. Style consistency: By introducing a structure-aware multi-scale attention mechanism, the present invention can maintain the style consistency of different parts of the clothing during the generation process, effectively avoiding the problem of inconsistent styles in traditional methods.
[0043] 3. Reduce the model complexity: Through the variance-guided network simplification strategy, the present invention effectively reduces the complexity of the network while ensuring the quality and result stability of the generated image, and improves the efficiency of the model. Brief Description of the Drawings
[0044] Figure 1 is the overall flowchart of the present invention for generating clothing images from sketches and texts.
[0045] Figure 2 is the schematic flowchart of the sketch prior embedding module of the present invention.
[0046] Figure 3 is the schematic diagram of the structure-aware multi-scale attention mechanism of the present invention.
[0047] Figure 4 is the schematic diagram of the process of the variance-guided network simplification strategy of the present invention.
[0048] Figure 5 is the schematic diagram for comparing the generation results of the method of the present invention with other multi-modal methods.
[0049] Figure 6 is the schematic diagram for comparing the generation results of the present invention with and without the sketch prior embedding module.
[0050] Figure 7 is the schematic diagram for comparing the generation results of the present invention with and without the structure-aware multi-scale attention mechanism.
[0051] Figure 8 is the schematic diagram for comparing the present invention with and without the variance-guided network simplification strategy. Detailed Embodiments
[0052] The present invention will be further described below through specific embodiments and the accompanying drawings. The embodiments of the present invention are for better enabling those skilled in the art to understand the present invention and do not impose any limitations on the present invention.
[0053] Such as Figure 1As shown in the figure, the method for generating clothing images based on sketches and text provided by the embodiments of the present invention includes the following steps:
[0054] Step 1: Input the text information describing the clothing, and use the clip text encoder to encode the input text into a feature matrix;
[0055] Step 2: Input the sketch image of the clothing, and extract features from the input sketch through the sketch prior embedding module;
[0056] Step 3: Fuse the input text feature matrix with the features of the sketch image and input them into the UNet encoder (SD Encoder) to generate preliminary clothing image features;
[0057] Step 4: Use the structure-aware multi-scale attention mechanism to further process the sketch and text features by calculating the correlation between different parts of the feature map, thereby improving the contour and style consistency of the clothing image;
[0058] Step 5: Process the result obtained in Step 4 through zero convolution to extract more accurate features, where zero convolution is reduced by variance-guided network simplification;
[0059] Step 6: Randomly sample a Gaussian noise from a Gaussian distribution and input it as an initial noise image into another UNet architecture (including SD Encoder and SD Decoder) that is the same as the encoder in Step 3. During the denoising process, by fusing with the result extracted in Step 5, the noise is gradually reduced and an intermediate image is generated. Through multiple iterations of denoising, a clear clothing image is finally obtained, which conforms to both the contour of the input sketch and the style described in the text;
[0060] Step 7: During the denoising process, calculate the difference between the image after denoising at each stage and the target image through the mean squared error loss (MSE Loss). Specifically, the mean squared error loss measures the error between the denoised image and the target image at the pixel level. The error gradient is transmitted back to the model through the backpropagation algorithm, thereby obtaining gradually optimized gradient information. These gradient information are used to update the weight parameters of the UNet model, enabling the model to better reduce noise and generate intermediate results closer to the target image in each step of denoising. Through multiple iterations of optimization, the UNet model gradually learns how to recover a clear clothing image that conforms to the sketch and text description from the noise image.
[0061] Based on the above method, the network designed by the present invention, as Figure 1 shown, includes a sketch prior embedding module and a structure-aware multi-scale attention mechanism, where the process of the sketch prior embedding module includes the following steps:
[0062] Step 11: Input the sketch image S and process it through the initial feature extraction module. First, the first layer of convolution extracts edge features, and then batch normalization and ReLU activation are used to obtain the feature map S1. Then, the second layer of convolution further refines the features, and after BN2, the refined feature map S2 is obtained to capture the key information in the sketch.
[0063] Step 12: Based on the feature map S2 obtained in the previous step, we designed a global feature weighting sub-module. This sub-module first performs global pooling on S2 to compress the spatial dimension into the global feature vector S gap , and the specific implementation is as follows:
[0064]
[0065] where S 2c represents the c-th channel of S2, c ∈ [1, C], and H, W, and C represent the height, width, and number of channels of the feature map respectively.
[0066] Step 13: To model complex feature relationships and capture the key information in the sketch, the global feature vector S gap is processed through a fully connected layer and the Silu activation function. Then, the weight factor for each channel is calculated through a fully connected layer and the Sigmoid activation function to obtain the feature map S g . These weight factors are used to adjust the importance of each channel, highlight the key features, and suppress the redundant parts. The specific process is as follows:
[0067] S g = σ(U2δ(U1S gap ))
[0068] where δ represents the Silu activation function, σ represents the Sigmoid function, U1 is the weight matrix used to reduce the channel dimension, U2 is the weight matrix used to restore the channel dimension, and r is the scaling parameter used to reduce the number of channels and reduce the computational amount.
[0069] Step 14: To prevent the feature information from gradually disappearing in the deep network, we introduce the residual connection mechanism of ResNet to fuse the weighted feature map with the preliminary feature map to obtain the stable feature map S res , and through downsampling operations and convolution processing, the final output feature S out is obtained. The specific process is as follows:
[0070] S res = S2 ⊙ S g ,
[0071]
[0072] where, ⊙ represents the element-wise Hadamard product, represents the downsampling operation, β represents normalization, represents the convolution operation.
[0073] The process of the structure-aware multi-scale attention mechanism includes the following steps:
[0074] Step 15, through the feature S output by Step 14 out , the input text feature matrix is fused with the feature S of the sketch image out and input into the UNet encoder to generate the preliminary clothing image feature, obtaining the feature h. At this time, the feature map h is divided into G sub-features h = [h0, h1,..., h G-1 , and the size of each sub-feature h i is M×N×K / G, where M, N, and K respectively represent the height, width, and number of channels of the sub-feature h i . Each sub-feature focuses on different spatial regions, helping the model understand the multi-dimensional features of the input, thereby improving the overall consistency and detail performance of the generated clothing image.
[0075] Step 16, we construct a global pooling sub-module to extract global context information from the feature map. For each sub-feature h i in Step 15, the global context is extracted by performing average pooling operations in the spatial dimensions (horizontal and vertical directions). This operation compresses the spatial information into a global feature for each channel, helping the model capture long-range dependencies. The specific implementation method is:
[0076]
[0077] where, d represents the global feature of the d-th channel of the feature map, d ∈ [1, K / G], h d (i, j) represents the value at position (i, j), h x and h y represent the average pooling results in the horizontal and vertical directions respectively, and σ is the Sigmoid activation function.
[0078] Step 17, in order to enhance the model's ability to focus on more relevant features, after extracting the global features h x and h y , they are concatenated and processed through a 1×1 convolution, and then the weights for each channel are calculated through the Sigmoid activation function. These weights are used to adjust the importance of each channel in the feature map according to the global context. The specific process is:
[0079]
[0080] where, hspatial denotes the weight matrix for each channel, σ is the Sigmoid activation function, and W is the weight of the 1×1 convolutional kernel. denotes the concatenation operation along the channel dimension.
[0081] Step 18: Extract local features from the original feature map h through a 3×3 convolution to capture fine-grained spatial details and supplement global context. The specific process is as follows:
[0082]
[0083] where h local denotes the locally extracted features, and W 3×3 denotes the 3×3 convolution operation.
[0084] Step 19: Calculate the spatial attention weights α1 and α2 through a cross-space mechanism. The cross-space mechanism enables the model to focus on key spatial regions, thereby improving the long-range dependence capture ability and enhancing feature consistency. These spatial attention weights are used to optimize the spatial features to make the generated image more coherent. This step strengthens the model's understanding of the relationships between different regions in the feature map, thus contributing to improving the overall consistency. Specifically as follows:
[0085] α1 = Softmax(AvgPool(h local )), α2 = Softmax(AvgPool(h spatial ))
[0086] where AvgPool represents average pooling, and Softmax represents the normalization activation function.
[0087] Step 110: Finally, by optimizing the weighted combination of the feature maps, the detail performance and overall consistency of the generated image are further improved. Combine the weights α1 and α2 with the weighted feature map h spatial and the local feature map h local to obtain the optimized feature map h out to enhance the overall consistency and detail retention of the generated image. Specifically as follows:
[0088] h out1 = α1·h spatial , h out2 = α2·h local , h out = σ(h out2 + h out1 )
[0089] where h out is the finally output feature map, and σ is the Sigmoid activation function.
[0090] The variance-guided network simplification strategy is as follows:
[0091] Step 111, for the feature h output by each encoding block in Step 110 out , these features are further processed by zero convolution to extract features. To measure the dispersion of the output features, we calculate the variance of the features h out directly output by each layer of encoding blocks, and call it the left variance. In addition, we also calculate the variance of the feature output after zero convolution processing, and define it as the right variance. We use the following formula to determine the impact of zero convolution on each encoding block, specifically as follows:
[0092]
[0093] where, σ L represents the left variance, σ R represents the right variance, and η represents the variance reduction amount, which enables us to measure the variance change caused by zero convolution for each layer.
[0094] Step 112, for each obtained variance reduction amount η, we first sort it in ascending order. Then, we compare the effects of the layer with the smallest variance reduction amount and the layer with the largest variance reduction amount. The experimental results show that the layer with the smallest variance reduction amount has a better effect when retaining zero convolution. Based on this finding, starting from the layer with the smallest variance reduction amount, we sequentially add other layers until the overall effect is close to the performance of the original model. Finally, we retain six layers with the smallest variance reduction amount, which show the smallest impact of zero convolution and can maintain the performance of the model.
[0095] On the other hand, the embodiment of the present invention also provides a clothing image generation system based on sketches and texts, including a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the clothing image generation method based on sketches and texts as described in the above technical solution.
[0096] The effects of the present invention are illustrated below through specific embodiments:
[0097] 1. Implementation details
[0098] We adopted Stable Diffusion v1.5 as our basic text-to-image model, which contains 16 cross-attention layers. During the training process, we used the AdamW optimizer based on PyTorch Lightning. The learning rate was set to 1×10^-5. Our model was trained on an NVIDIA RTX 4090-24G GPU with a batch size of 4. The sizes of the input and output images were both 512×512.
[0099] 2. Data Preparation
[0100] We collected a new dataset. First, we collected approximately 30,000 fashion images from multiple websites and attached corresponding text descriptions. To ensure the consistency of image size and quality, we cropped and adjusted the resolution of all the collected images. Then, we used the Photosketch network to convert all fashion images into sketches. To further improve the clarity and usability of the sketches, we performed binarization to remove background noise and irrelevant details, thus simplifying the sketches. Each pair of data includes an original image, a corresponding sketch, and a detailed text description, all with dimensions adjusted to 512×512.
[0101] 3. Evaluation Metrics
[0102] ● FID (Frechet Inception Distance): FID is a metric for measuring the difference between generated images and real images by calculating the distribution difference between generated images and real images in the deep feature space. The smaller the FID value, the closer the feature distribution of the generated image is to that of the real image, and the better the generation effect. It is widely used to evaluate the performance of generative models such as generative adversarial networks.
[0103] ● SSIM (Structural Similarity Index): SSIM is an image quality evaluation criterion that assesses the similarity in luminance, contrast, and structure between generated images and real images. Its value ranges from 0 to 1, and the higher the value, the more similar the images are and the better the quality. SSIM can better reflect the visual perception effect of images in the human eye and is commonly used in the field of image generation.
[0104] ● LPIPS (Learned Perceptual Image Patch Similarity): LPIPS is a perceptual image similarity evaluation method that extracts high-level features of images through a deep neural network and measures the perceptual difference between generated images and real images. The smaller the LPIPS value, the better the effect.
[0105] 4. Comparison with Other Models
[0106] Quantitative Evaluation: Four advanced models that currently generate images under sketch and text conditions were selected for the experiment: ControlNet, Gligen, UniControl, and FineControlNet. To make a fair comparison with our method, we retrained all the selected models on our custom dataset. By training these models on the same comparison dataset, we eliminated potential biases that might be caused by data differences, ensuring that the evaluation results accurately reflect the performance of the models under the same conditions. Then, we quantified the comparison results and calculated the scores of LPIPS, FID, and SSIM. The metric results are shown in Table 1. As can be seen from the results in Table 1, our method is significantly superior to the comparative methods in terms of perceptual quality and feature alignment. The significant reduction in LPIPS and FID, as well as the improvement in SSIM, indicate that the fashion images generated by our method are not only perceptually closer to real images but also maintain a more accurate feature distribution and structural integrity. This is crucial for both the visual fidelity of the generated images and their consistency with the text description. The generated results are as Figure 5 shown. Compared with several other methods, the method of the present invention can not only excellently reproduce the input sketch and text description in terms of details and accuracy, but also the generated fashion images are more visually consistent and diverse, showing obvious advantages over the competing methods. Thanks to the collaborative integration of the sketch prior embedding module and the structure-aware multi-scale attention mechanism, the method of the present invention has shown significant advantages in detail retention and semantic consistency.
[0107] Table 1 Quantitative Comparison of Multimodal Text-Condition Image Generation Models
[0108] Method FID SSIM LPIPS ControlNet 22.47 0.7007 0.2612 Gligen 16.29 0.7488 0.2119 UniControl 18.86 0.7429 0.2143 FineControlNet 19.39 0.7227 0.2255 The present invention 15.31 0.7644 0.1966
[0109] 5. Ablation Experiments
[0110] In this section, the impacts of the proposed sketch prior embedding module, structure-aware multi-scale attention mechanism, and variance-guided network simplification on model efficiency and performance are mainly discussed. To verify the contribution of each proposed module, we successively removed these modules from the entire model. The results are shown in Table 2, where w / o(SAMA) means the structure-aware multi-scale attention (SAMA) module is removed, and w / o(SPEM) means the sketch prior embedding module (SPEM) is removed.
[0111] Table 2 Ablation Experiments
[0112] Method FID SSIM LPIPS ControlNet 22.47 0.7007 0.2612 w / o SAMA 16.87 0.7601 0.2024 w / o SPEM 18.48 0.7378 0.2015 The present invention (w / o VNS) 15.07 0.7685 0.1910
[0113] Impact of the Sketch Prior Embedding Module on Image Contours: As shown in the second row of Table 2, the sketch prior embedding module is crucial for maintaining fine details of the image contours. After removing this module, the generated images cannot be closely aligned with the input sketches, especially showing distortions in key structural areas such as the hems and cuffs of the clothing. As Figure 6 shown, the lack of SPEM leads to shape inconsistencies and feature alignment problems, resulting in a significant decrease in the fidelity of the overall contour. This fact proves that SPEM plays an important role in maintaining contour accuracy, ensuring that the generated images match the input sketches.
[0114] Impact of the Structure-Aware Multi-Scale Attention on Image Style Consistency: As shown in the third row of Table 2, the structure-aware multi-scale attention enhances the model's ability to maintain style consistency across different image regions while effectively retaining text information. From Figure 7 the examples provided, it can be seen that after removing SAMA, there are obvious inconsistencies in key style elements such as texture and color distribution. For example, in Figure 7 the first set of images, although the text description is "pure black", removing SAMA results in the loss of text information, and the generated clothing images do not match the text description. In addition, in the second set of examples, when SAMA is removed, the style of the clothing images lacks consistency. On the contrary, after adding SAMA, the clothing images we generated ensure smooth style transfer and alignment with the text description of the generated images. In summary, the lack of the multi-scale attention mechanism reduces the model's ability to align text information and coordinate global and local features, resulting in a decrease in the coherence and realism of clothing style performance.
[0115] Impact of the Variance-Guided Network Simplification on Model Efficiency and Performance: The inference time comparison shown in Table 3 highlights the significant impact of the variance-guided network simplification (VNS) module on improving computational efficiency. By introducing variance-guided network simplification, this module effectively reduces unnecessary zero convolutions while retaining key features. This optimization reduces the inference time by approximately 10%, from 6.73 seconds (without VNS) to 6.05 seconds (with VNS). Notably, as Figure 8 shown, there are almost no visual differences between the images generated with and without VNS, and the FID scores are also very close, reaching 15.31. This further proves that VNS successfully balances computational efficiency while maintaining high-quality generation results. Specifically, the saved time demonstrates the effectiveness of selecting layers with the least variance reduction for simplification, thus achieving a balance between performance and efficiency. This efficiency improvement is achieved without sacrificing the quality of the generated output. These results verify the role of VNS in optimizing computational overhead while maintaining high model performance and provide a practical solution for scenarios that require both speed and accuracy.
[0116] Table 3 Comparison of Inference Time with and without Variance-guided Network Simplification (VNS) Module
[0117] Method w / o VNS The present invention Time (s) 6.73 6.05
[0118] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar means for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A method for generating clothing images based on sketches and texts, characterized in that: The steps include: Step 1: Input text information describing clothing and encode the input text into a feature matrix; Step 2, inputting a sketch image of clothing and performing feature extraction on the input sketch; Step 3, the input text feature matrix is fused with the features of the sketch image and input into the Unet encoder to generate preliminary clothing image features; Step 4: Use the structure-aware multi-scale attention mechanism to further process the sketch and text features by calculating the correlation between different parts of the feature map; Step 5, the result obtained in step 4 is further processed by zero convolution to extract more accurate features, wherein the zero convolution is reduced by variance-guided network simplification; Step 6: Generate a Gaussian noise by randomly sampling from the Gaussian distribution and input it as the initial noise image into another UNet model with the same encoder as step 3. At the same time, the text features extracted in step 1 are also input into the UNet model as conditions to guide the generation process. In the denoising process, by fusing with the result extracted in step 5, the noise is gradually reduced and an intermediate image is generated. After multiple iterations of denoising, a clear clothing image is finally obtained. Step 7: During the denoising process, the difference between the denoised image at each stage and the target image is calculated, and a clear clothing image that conforms to the sketch and text description is restored from the noisy image through the iteratively optimized Unet model.
2. The method for generating clothing images based on sketches and texts as claimed in claim 1, characterized in that: In step 2, the sketch prior embedding module is used to extract features from the input sketch. The specific process is as follows: based on the input sketch S input First, the basic edge features are captured by the preliminary feature extraction module, which includes: the first convolution layer extracts edge features and processes them through batch normalization and ReLU activation function to obtain a preliminary feature map S1; then, the second convolution layer further refines the features to obtain a refined feature map S2; in order to remove redundant information, S2 is then globally pooled through the global feature weighting module to obtain a global feature vector S gap : The channel weights are calculated through the fully connected layer and the Sigmoid activation function to obtain the weighted feature map S g ; Finally, combined with the residual connection mechanism, the weighted feature map is fused with the preliminary feature map to obtain a stable feature map S res , through downsampling and convolution processing, the final output feature S out .
3. The method for generating clothing images based on sketches and texts as claimed in claim 2, characterized in that: Global eigenvector S gap The calculation method is as follows: Among them, S gap represents the output of the global average pooling layer, S 2c represents the c-th channel of S2, and c∈[1,C], H, W, C represent the height, width and number of channels of the feature map respectively.
4. The method for generating clothing images based on sketches and texts as claimed in claim 2, characterized in that: Weighted feature map S g The calculation method is as follows: S g =σ(U2δ(U1S gap )) Among them, δ represents the Silu activation function, σ represents the Sigmoid function, U1 is the weight matrix used to reduce the channel dimension, U2 is the weight matrix used to restore the channel dimension, and r is the scaling parameter.
5. The method for generating clothing images based on sketches and texts as claimed in claim 2, characterized in that: Output feature S out The calculation method is as follows: S res =S2⊙S g , Among them, ⊙ represents the element-wise Hadamard product, represents the downsampling operation, β represents normalization, Represents a convolution operation.
6. The method for generating clothing images based on sketches and texts as claimed in claim 1, characterized in that: The specific process of step 4 is: According to the features output in step 3, the features are further processed through the self-attention mechanism and the cross-attention mechanism to obtain the feature map h, and the feature map h is input into the structure-aware multi-scale attention mechanism. First, the feature map is split into G sub-feature maps h = [h0,h1,…,h G-1 ], each sub-feature h i The size of is M×N×K / G, where M, N, and K represent the sub-feature h respectively. i The height, width and number of channels of the feature vector are then pooled to extract the global context information and obtain the global feature dependency h. x and h y , and adjust the importance of each channel through 1×1 convolution and Sigmoid activation function to obtain the global weighted feature map h spatial ; At the same time, a 3×3 convolution is used to extract local features h local ,Then, the spatial attention weights α1 and α2 are calculated through a cross-space mechanism; α1=Softmax(AvgPool(h local )),α2=Softmax(AvgPool(h spatial )) Among them, h local represents the local extracted features, W 3×3 Represents a 3×3 convolution operation, and AvgPool represents average pooling; Combine the weight α1 with the global weighted feature map h spatial , weight the feature map by dot multiplication to get h out1 Similarly, the weight α2 is combined with the local feature map h local Dot product to get h out2 , then the feature h out1 and h out2 Fusion and addition get the feature map h out , as follows: h out1 =α1·h spatial ,h out2 =α2·h local ,h out =σ(h out2 +h out1 ) Among them, h out is the feature map of the final output, σ is the Sigmoid activation function, and h out1 and h out2 They represent the feature map output results weighted by spatial attention weights α1 and α2 respectively.
7. The method for generating clothing images based on sketches and texts as claimed in claim 6, characterized in that: Global weighted feature map h spatial The calculation method is as follows: Among them, d represents the global feature of the dth channel of the feature map, d∈[1,K / G], h d (i,j) represents the value of position (i,j), h x and h y Respectively represent the average pooling results in the horizontal and vertical directions, σ is the Sigmoid activation function, W represents the weight of the 1×1 convolution kernel, Indicates the concatenation operation in the channel dimension; h spatial Represents the weight matrix for each channel.
8. The method for generating clothing images based on sketches and texts as claimed in claim 1, characterized in that: The specific process of step 5: For each feature h output by the encoding block in step 4 out , calculate each layer of encoding block to directly output the feature h out The variance of the feature output after zero convolution is calculated and called the left variance. In addition, the variance of the feature output after zero convolution is calculated and defined as the right variance. The following formula is used to determine the impact of zero convolution on each encoding block: Among them, σ L represents the left variance, σ R represents the right variance, and η represents the variance reduction; For each variance reduction η obtained, sort it in ascending order, starting from the layer with the smallest variance reduction, and retain several zero convolution layers with the smallest variance reduction.
9. The method for generating clothing images based on sketches and texts as claimed in claim 8, characterized in that: Keep the six zero-convolutional layers that have the smallest variance reduction.
10. Sketch and text-based clothing image generation system, characterized in that: The invention comprises a processor and a memory, wherein the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the method for generating clothing images based on sketches and texts as claimed in any one of claims 1 to 9.
Citation Information
Patent Citations
Clothing sketch-to-image generation method based on multi-modal information
CN115393456A
Virtual fitting method based on diffusion model
CN118278291A
Two-stage low-illumination image enhancement method based on wavelet transform
CN119494792A
Underwater image enhancement method based on brightness-mask-guided multi-attention mechanism
WO2024208188A1
Cited By
Clothing attribute identification method and system based on multi-modal information hierarchy semantic modeling
CN121708402A
Garment image generation method based on basic element retrieval and replacement
CN122049119A