A Clothing Sketch-to-Image Generation Method Based on Multimodal Information

By combining the clothing image layered encoding model and the Transformer-G/L model, using multimodal information to generate clothing images, the problem of single-modal method generating image single color and uncontrollable attributes is solved, and a more realistic and flexible clothing image generation is achieved.

CN115393456BActive Publication Date: 2025-06-03WUHAN TEXTILE UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202210885260.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-26
Publication Date
2025-06-03
Estimated Expiration
2042-07-26

AI Technical Summary

Technical Problem

The existing clothing image generation methods are mainly single-modal methods. The generated images are single in color and uncontrollable in attributes. The traditional methods require a lot of time and require high data sets, and rely on the label information of the sketch.

Method used

Using a clothing sketch-to-image generation method based on multimodal information, the clothing image layered encoding model and the Transformer-G/L model are used to combine sketch encoding, text encoding and clothing global/local information encoding to generate clothing images with specified attributes.

Benefits of technology

The generated clothing images are more realistic, the attribute control is more flexible, and the diversity and fidelity are significantly improved. Compared with the most advanced single-modal generation method, FID is reduced by 24.19, the IS value is increased by 10.73%, and the LPIPS value is increased by 10.8%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393456B_ABST
    Figure CN115393456B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for generating clothing images from clothing sketches based on multi-modal information. By utilizing the collaborative feature representation of multi-modal information, the present invention proposes a multi-modal generation model for clothing sketches to images. This model combines the advantages of CNN multi-scale feature extraction with the powerful masked attention mechanism of Transformer. For clothing images, a hierarchical encoding model is proposed, and a feature matching loss is introduced to make the texture of the generated clothing images clearer. At the same time, a Duplicate-Transformer is proposed to learn the association between different modal information and collaboratively guide the generation of clothing images with specified attributes. The method of the present invention can generate highly realistic clothing images, and has greater flexibility in attribute control, with obvious improvements in both diversity and fidelity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image generation methods, and particularly to a method for generating an image from a clothing sketch based on multimodal information. Background Art

[0002] With the continuous improvement of people's living standards, the demand for personalized clothing design is increasing. In the field of clothing design, designers will show the new designs conceived in their minds in the form of hand-drawn sketches before designing specific clothing, and then add more detailed information, such as folds, to form a pattern diagram, and finally color the pattern diagram to generate the final clothing effect diagram. However, the original clothing design process is complex, time-consuming and laborious, and it is difficult for the general public to design their favorite clothing styles according to their own wishes. With the development of deep learning, artificial intelligence technology has gradually been used to assist in clothing generation. However, most of the existing auxiliary generation methods are single-modal methods, and the generated clothing images are too single.

[0003] Traditional clothing image generation methods are mainly based on image retrieval, which is an indirect image generation technology. First, relevant pictures are searched in the database using keywords, and then the returned pictures are compared with the sketches one by one to find the pictures with higher matching degrees as the generated images. This method not only takes a lot of time, but also has high requirements for the dataset and strong dependence on the label information of the sketches.

[0004] With the development of deep convolutional neural network (DCNN), generative adversarial network (GAN) has shown great potential in image generation. This type of method regards the conditional image generation task as an image-to-image conversion task. When converting, the sketch information is shared between the input and the output, thus ensuring the similarity between the input image and the output image. However, most of the inputs of these methods are single-modal, and they only focus on the mapping between two image domains, ignoring other semantic attribute information, which makes the generated images have single colors and uncontrollable attributes. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method for generating an image from a clothing sketch based on multimodal information in view of the above problems and requirements.

[0006] To solve the above technical problems, the present invention adopts the following technical solutions:

[0007] A method for generating an image from a clothing sketch based on multimodal information, comprising the following steps:

[0008] Step 1. Input the clothing image into the trained clothing image hierarchical encoding model. Use the clothing image hierarchical encoding model to process the clothing image to obtain the local feature map and the global feature map, and use the corresponding basis vector space to perform vector quantization on the local feature map and the global feature map respectively to obtain the global information encoding of the clothing and the local information encoding of the clothing;

[0009] Input the clothing sketch into the sketch encoding model to obtain the sketch encoding;

[0010] Input the text into the input text encoding model to obtain the text encoding;

[0011] Step 2. Combine the sketch encoding and the text encoding to form conditional information, connect them with the global information encoding of the clothing and input them into Transformer-G together. The masked attention mechanism in Transformer-G automatically generates the global information sequence of the clothing; Use the global information encoding of the clothing image as conditional information, connect it with the local information encoding of the clothing and input them into Transformer-L together to automatically generate the local information sequence of the clothing;

[0012] Step 3. Input the generated global information sequence of the clothing and the local information sequence of the clothing into the corresponding basis vector space respectively to find the corresponding vectors, and convert them into the global two-dimensional feature map and the local two-dimensional feature map respectively. Input the global two-dimensional feature map and the local two-dimensional feature map into the trained decoder D to generate the final image.

[0013] Furthermore, the training methods of the clothing image hierarchical encoding model and the decoder D include the following steps:

[0014] Step 1.1. Input the image x into the encoder E of the clothing image hierarchical encoding model. The encoder E downsamples the image x to 1 / 4 of its original size to obtain the local feature map Z containing detailed information local , the local feature map Z local Continue to downsample to 1 / 8 of the image to obtain the global feature map Z containing global information global ;

[0015] Step 1.2. Input the global feature map Z global into the global basis vector space V global , to obtain the global information encoding Z of the clothing global ; Upsample Z global to the same size as the local feature map Z local and perform residual connection with it. Input the fused feature after connection into the local basis vector space V local , to obtain the global information encoding Z of the clothing local ;

[0016] Step 1.3. Input the global information encoding Z of the clothinglocal and the global information encoding Z of the clothing global are input into the decoder D to obtain the reconstructed image: x' = D(Z local , Z global )

[0017] Step 1.4. Use the discriminator to calculate the total loss function L = L encoding + L fm , where L fm is the average value of the loss values of all layers of the discriminator. The formula for calculating the loss value of the t-th layer of the discriminator is: where N is the number of features of the t-th layer of the discriminator, is the i-th feature of the image x extracted by the t-th layer discriminator;

[0018]

[0019] where is to find the L2 norm loss, ||·|| 2 is to find the sum of squares loss, and sg[] represents stopping the gradient propagation and stopping updating the parameters of this part; by minimizing the total loss function, the encoder E, the local basis vector space V global , the global basis vector space V global , the decoder D and the discriminator are trained and the parameters are updated. After training, a trained hierarchical encoding model of clothing images and a trained decoder D are obtained.

[0020] Furthermore, in the step 1.1, the method for obtaining Z global is as follows: for each vector in the global feature map Z global , the quantizer q(·) finds the vector closest to it in V global through the nearest neighbor algorithm and replaces it to obtain Z global ; the method for obtaining Z local is as follows: upsample Z global to the same size as Z local and perform a residual connection with Z local to obtain a new fused feature, and then find the vector closest to each position in the fused feature map in V local through the nearest neighbor algorithm and replace it to obtain Z local .

[0021] Furthermore, in the step 3, Transformer-G and Transformer-L are used to generate clothing sequences by means of autoregressive prediction.

[0022] After the present invention adopts the above technical solutions, compared with the prior art, it has the following advantages:

[0023] The present invention utilizes the collaborative feature representation of multi-modal information to propose a multi-modal generation model for clothing sketches to images. This model combines the advantages of CNN multi-scale feature extraction with the powerful masked attention mechanism of Transformer. For clothing images, a hierarchical encoding model is proposed, and a feature matching loss is introduced to make the texture of the generated clothing images clearer. At the same time, a Duplicate-Transformer is proposed to learn the associations between different modal information and collaboratively guide the generation of clothing images with specified attributes. In the experiments, the method of the present invention can generate highly realistic clothing images and has greater flexibility in attribute control. After comparison, the FID of the images generated by the method of the present invention is reduced by 24.19 compared with the state-of-the-art attention-guided single-modal generation method U-GAT-IT, the IS value is increased by 10.73% compared with the state-of-the-art MUNIT, and the LPIPS value is increased by 10.8%. Both the diversity and fidelity are significantly improved.

[0024] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0025] Figure 1 Schematic diagram of the hierarchical CNN and Transformer fusion model;

[0026] Figure 2 Schematic diagram of the clothing image hierarchical encoding model;

[0027] Figure 3 Schematic diagram of the autoregressive Duplicate-Transformer;

[0028] Figure 4 Schematic diagram of the comparison between the method of the present invention and traditional single-modal methods;

[0029] Figure 5 Multi-modal information generation result diagram of the present invention;

[0030] Figure 6 Schematic diagram of the comparison of the images reconstructed by different clothing encoding methods;

[0031] Figure 7 Mean square error diagram during training;

[0032] Figure 8 Ablation experiment result diagram. Detailed Implementation Manner

[0033] The principles and features of the present invention will be described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0034] I. Clothing Image Generation Framework

[0035] The present invention proposes a clothing sketch to clothing image generation model based on multi-modal information, which combines the advantages of CNN multi-scale feature extraction with the powerful masked attention mechanism of Transformer. Its overall architecture is as Figure 1 shown, mainly divided into two stages: In the first stage, each encoder separately learns the encoding of the corresponding modal data and performs feature fusion. At the same time, in view of the problem of unclear texture of clothing images, the present invention designs a hierarchical CNN encoding method, as detailed in Section 2.1; In the second stage, according to the feature fusion encoding obtained in the first stage, a multi-modal information association learning model: Duplicate-Transformer is proposed to generate clothing images with specific attributes. This model consists of a pair of unidirectional autoregressive Transformers with exactly the same structure, which is used to learn the association between different modal information.

[0036] II. Multi-modal Information Encoding

[0037] In the encoding stage of multi-modal information, clothing sketches, texts, and clothing images are used to learn the corresponding encoders. Each encoder separately extracts the feature encodings of the corresponding modal data and performs fusion.

[0038] 2.1 Clothing Image Encoder

[0039] In order to convert the continuous clothing image content into a discrete sequence, the present invention no longer focuses on pixels, but uses the idea of vector quantization to perform discrete encoding on the image. The basic idea of vector quantization is to find the nearest neighbor of the target vector in the basis vector space through the nearest neighbor algorithm, and then use the number of the basis vector to encode the target vector.

[0040] Since clothing images contain many complex textures and colors, it is necessary to make the encoder pay more attention to these detail information. In the multi-scale feature extraction of CNN, features at different levels contain different information. Generally speaking, high-level features pay more attention to semantic information and less attention to detail information; low-level features contain more detail information. Inspired by this, the present invention combines the advantages of different layers in a simple and effective way: that is, vector quantization of clothing images is performed in different feature layers respectively to retain more useful information. In order to make the generated image clearer, the present invention separately models the local information (such as texture) and global information (such as shape) of the clothing.

[0041] As Figure 2 shown, the encoder E maps the sketch to two different size spaces through convolution. Specifically, for a clothing image x ∈ R H×W×3 , the encoder first downsamples it to 1 / 4 of its original size to obtain a local feature map containing detail information Then continue downsampling to 1 / 8 of the original size to obtain a global feature map containing global information The present invention adopts the idea of vector quantization to learn two basis vector spaces for the model, where the local basis vector space is used to encode texture details and is called V local , while the global basis vector space is used to encode global information and is called V global . For each vector at position (i, j) in the global feature map Z global , the quantizer q(·) finds the closest vector e gl o bal in V k through the nearest neighbor algorithm and replaces it, where k is the index, and the process is as follows:

[0042]

[0043] To fuse the feature information of different layers, upsample Z global to the same size as Z local and perform a residual connection with it to obtain a new fused feature. Then, find the closest vector e local for each position in this fused feature map through the nearest neighbor algorithm k and replace it to obtain:

[0044] Z local =q(Z local +upsample(Z global ))

[0045] Finally, input the two quantized feature maps into the decoder D to obtain the reconstructed image:

[0046] x' = D(Z local , Z global )

[0047] Through this encoder, any input image x ∈ R H×W×3 can be represented by a combination of several vectors in the global basis vector space and the local basis vector space. The indices corresponding to these vectors in their respective basis vector spaces are the global information encoding and local information encoding of the clothing, and the encoding loss function is:

[0048]

[0049] where ||x - x'|| 2 is the sum of squared errors of the difference between the reconstructed image and the real image. Since the quantization process is not differentiable, during the training process, the gradient is directly copied from the decoder to the encoder, and sg[] represents stopping the gradient propagation and stopping updating the parameters of this part.

[0050] It was found in the experiment that since there are many blank areas around the clothing images, the mean squared error between the generated images and the real images is very low at the beginning, only about 0.4. After 30 epochs of iterative optimization, the convergence speed is extremely slow, and the texture of the clothing edge part is not clear. See Section 4.5 for details. In order to enable the model to pay attention to the detailed texture of the clothing images during optimization, a discriminator is designed in the model in the present invention, and the original cross-entropy loss of the discriminator is changed to a feature matching loss. The feature matching loss extracts features from different scales of the real images and the generated images and performs "matching". It can pay attention to the most significant gap between the features of the generated samples and the real samples, such as the texture difference in the feature space. The loss function is as follows:

[0051] where L fm is the average value of the loss values of all layers of the discriminator. The formula for calculating the loss value of the t-th layer of the discriminator is: where N is the number of features of the t-th layer of the discriminator, is the i-th feature of the image x extracted by the t-th layer discriminator;

[0052] The total loss function is:

[0053] L = L encoding + L fm

[0054] 2.2 Text Encoder and Sketch Encoder

[0055] The text and the clothing sketch together constitute the conditional information. For the encoding of the text, the present invention directly uses Word2vec. This method uses a lightweight neural network to train a word-related model on a corpus, which can map each word in the corpus to a vector space of a specified dimension and represent the relationship between different words through the correlation between vectors.

[0056] For the clothing sketches, since they only contain some simple black lines and carry less effective information and do not require much attention to their detailed information, it is very easy to obtain their image representations. In the present invention, the VQ-VAE model is used to directly encode them into a series of discrete sequences.

[0057] III. Multimodal Information Association Learning Model

[0058] The multimodal information association learning model uses the features extracted by each encoder to learn the association between different modal information, so as to generate corresponding clothing images according to the given sketches and texts.

[0059] The present invention uses the method of autoregressive prediction to generate clothing sequences. In the autoregressive model, given a series of vectors x = (x 1 , …, x n),The model calculates the joint distribution by taking each feature as a condition, that is

[0060]

[0061] Then, the feature with the highest probability is found by maximizing the likelihood estimation and used as the predicted value. In the task of the present invention, additional conditional information needs to be added to control the generation process. The conditional information is defined as c, and the model needs to learn the conditional probability distribution p(x, c):

[0062]

[0063] On this basis, the present invention proposes an autoregressive Duplicate-Transformer, which consists of a pair of unidirectional autoregressive Transformers with exactly the same structure. Among them, Transformer-L is used to generate the encoding of clothing detail information, and Transformer-G is used to generate the encoding of clothing global information. As Figure 3 shown, in the training stage, the sketch and its corresponding clothing image and text description information are respectively fed into the corresponding encoders. For Transformer-G, the sketch encoding, text encoding, and clothing global information encoding are fused as the input. Among them, the sketch and text jointly constitute the conditional information, and the clothing global information sequence is automatically generated through the masked attention mechanism in the Transformer. For Transformer-L, the global information encoding of the clothing image is used as the conditional information and connected with the local information encoding as the input to automatically generate the local information sequence of the clothing. Finally, the generated global information sequence and local information sequence respectively find the corresponding vectors in the basis vector space and are converted into two-dimensional feature maps, which are sent to the decoder D to generate images.

[0064] IV. Experimental part

[0065] 4.1 Data preparation

[0066] Two datasets are used in the experiment to verify the model of the present invention. The first is the VITON dataset, which is used to evaluate the clothing unimodal generation task. The clothing images in this dataset only contain some simple colors and textures; the other is the FEIDEGGER dataset, which is used to evaluate the clothing multimodal generation task. The clothing images in this dataset are mainly women's dresses, including complex patterns and diverse colors, and each image has a corresponding text description. In the experiment, Photo-Sketching is used to convert the clothing images into corresponding sketches.

[0067] 4.2 Experimental settings

[0068] In the first stage, the sizes of the global and local basis vector spaces of the clothing are set to 512 dimensions, and the input image size is 256×256 pixels. For clothing images, the local feature map Z local is set to 32×32, and the global feature map Z global is set to 16×16; for sketches, they are directly compressed into a discrete space of Z = 16×16. The experiment chooses to train the model on 4 V-100 GPUs (32GB), the batch size is 128, the number of iterations is 500 times, the parameters are updated by the Adam optimizer, and the learning rate is set to 0.0003. In the second stage, the number of Transformer layers is set to 24 layers, each layer contains 16 heads, the batch size is set to 32, and the number of iterations is 500 times.

[0069] 4.3 Evaluation Metrics

[0070] The experiments of the present invention mainly use the following three evaluation metrics:

[0071] · IS (Inception Score): IS mainly uses the InceptionNet-v3 network to evaluate the clarity and diversity of the generated images. A higher IS value indicates higher fidelity and diversity.

[0072] · FID (Fréchet Inception Distance): FID is used to measure the distance between the Inception feature vectors of real images and generated images in the same domain. A lower FID means that the generated images are closer to the distribution of real images.

[0073] · LPIPS (Learned Perceptual Image Patch Similarity): LPIPS is also known as "perceptual loss" and is used to measure the difference between two images. The lower the LPIPS value, the more similar the two images are and the lower the diversity, and vice versa.

[0074] 4.4 Generation from Clothing Sketches to Images

[0075] 4.4.1 Comparison with Traditional Unimodal Methods

[0076] The experiment selected four classic unimodal conditional image generation methods, including the original conditional generation network Pix2Pix, the unsupervised generation model CycleGAN that does not require paired data, the generation model MUNIT with relatively optimal generated image diversity, and the unsupervised generation model U-GAT-IT based on the attention mechanism. For these unimodal methods, the text description information is not considered, and only the mapping ability of each model between the two image domains is evaluated. The results of these methods are reported in Table 1:

[0077] Table 1 Benchmarks of Different Methods

[0078]

[0079]

[0080] The scores of IS and LPIPS of the clothing images generated by the method of the present invention are 3.30 and 0.374 respectively, which are increased by 10.73% and 10.8% compared with MUNIT; the FID score is 26.08, which is increased by 24.19 compared with the state-of-the-art U-GAT-IT. The generated results are shown in Figure 4 . Compared with several other methods, the images generated by the method of the present invention are more diverse and contain more colors; in addition, the results of the present invention have very high fidelity and can generate wrinkles of real clothing.

[0081] 4.4.2 Experimental Results Based on Multimodal Information Generation

[0082] The experiment uses the FEIDEGGER dataset to train the model and tests it on the test set. The results are as Figure 5 shown. The first column is the input clothing sketch, the second column is the text description information, and the third, fourth, and fifth columns are the generated clothing images. The results show that the method proposed by the present invention fully combines the sketch information and the text description information and can generate clothing images with specified attributes. In particular, the model can generate multiple styles of clothing images based on the same sketch and text, which significantly improves its diversity. However, due to the limitation of the dataset scale, some semantic attribute information has not been fully learned, such as the pattern types on the clothing and the material types of the clothing.

[0083] 4.4.3 Comparison with Single-Layer Encoding Methods

[0084] To fully prove the superiority of the clothing encoding method of the present invention, three currently very excellent single-layer encoding models are selected, mainly including VQ-VAE, PeCo, and VQ-GAN. The experiment uses the same configuration to train the three encoding models on the VITON dataset and the FEIDEGGER dataset. In the test stage, clothing images with multiple wrinkles or complex textures are selected and reconstructed using the trained models respectively. The reconstruction results are as Figure 6As shown, it can be found that on the VITON dataset, the reconstruction ability of VQ-VAE at the clothing folds is relatively poor. The pattern reconstructed by PeCo has lost some color information. There is relatively little difference between VQ-GAN and the method of the present invention. On the FEIDEGGER dataset, the texture of the reconstructed image of VQ-VAE is chaotic, and there are many blurred patches at the edges of the clothing images. The images reconstructed by PeCo contain a lot of noise, indicating that neither of them has fully learned the texture and color of the clothing. Although VQ-GAN shows strong reconstruction ability on the VITON dataset, it still cannot faithfully reproduce when faced with clothing with complex textures and mixed colors. In contrast, the method of the present invention preserves the color and texture information of the image as much as possible, and can achieve a relatively good restoration effect even when faced with intricate lines and spots. In addition, the reconstruction FID was evaluated on two datasets in the experiment, and the results are reported in Table 2. On the VITON dataset, the method of the present invention has a slight improvement compared to VQ-GAN. On FEIDEGGER, the method of the present invention has an improvement of 1.65 on the training set and 4.43 on the test set.

[0085] Table 2 Reconstruction FID of Different Encoding Methods

[0086]

[0087] 4.5 Ablation Experiments

[0088] In this section, the ablation studies of the feature matching loss and the hierarchical encoding method proposed in the model are mainly discussed.

[0089] 4.5.1 Feature Matching Loss

[0090] To explore the effectiveness of the feature matching loss, the experiment uses the same configuration to train two models on the FEIDEGGER dataset, one of which includes the feature matching loss and the other is the original loss. Figure 7 The mean square error (MSE) during training is shown. The initial MSE is very low, about 0.4, because the clothing images contain many blank areas and the MSE only focuses on the differences between pixel levels. It can be found that the model with the feature matching loss converges faster than other models, almost twice as fast as the original model, and the overall MSE is also smaller, indicating that the feature matching loss can initially focus on the low-level features of the clothing. At 390 epochs, the images reconstructed by the model of the present invention can already faithfully reproduce the texture and color information of the clothing. Table 3 shows the reconstruction FID of this model on FEIDEGGER, and the feature matching loss has a significant improvement on the reconstruction FID. Figure 8Examples of image reconstruction results for different loss types (the first column is the input image, the second column is the original loss reconstruction image, and the third column is the reconstruction image with only feature loss added).

[0091] 4.5.2 Hierarchical Encoding

[0092] To verify the advantages of hierarchical encoding, in the experiment, the feature matching loss was removed from the model of the present invention, and the local clothing information encoding was removed and changed to single-layer encoding. Figure 8 Shows an example of a reconstruction result (the fourth column is the reconstruction image with only the hierarchical encoding mechanism added, and the fifth column is the reconstruction image of the method of the present invention). The encoding method of the present invention can reconstruct the image texture more clearly, especially at the edges of the clothing. Single-layer encoding will produce ghosting and noise, while the method of the present invention shows better reconstruction results. The experiment evaluated the reconstruction FID on FEIDEGGER and reported the results in Table 3. The experiment proves that both hierarchical encoding and feature matching loss play a relatively large role in improving the encoding ability of the model. Combining the two can enable the model to fully focus on the most subtle textures in clothing images.

[0093] Table 3 Reconstruction FID on FEIDEGGER

[0094]

[0095] 5 Conclusion

[0096] The present invention uses the collaborative feature representation of multi-modal information to propose a multi-modal generation model for clothing sketches to images. The model combines the advantages of CNN multi-scale feature extraction with the powerful masked attention mechanism of Transformer. For clothing images, a hierarchical encoding model is proposed and a feature matching loss is introduced to make the generated clothing image texture clearer. At the same time, a Duplicate-Transformer is proposed to learn the association between different modal information and collaboratively guide the generation of clothing images with specified attributes. In the experiment, the method of the present invention can generate highly realistic clothing images and has greater flexibility in attribute control. After comparison, the FID of the images generated by the method of the present invention is reduced by 24.19 compared with the state-of-the-art attention-guided single-modal generation method U-GAT-IT, the IS value is increased by 10.73% compared with the state-of-the-art MUNIT, and the LPIPS value is increased by 10.8%. Both diversity and fidelity have been significantly improved.

[0097] The above is an example of the best implementation mode of the present invention, and the parts not described in detail are all common general knowledge of those skilled in the art. The protection scope of the present invention shall be subject to the content of the claims, and any equivalent transformation based on the technical inspiration of the present invention is also within the protection scope of the present invention.

Claims

1. A method for generating an image from a clothing sketch based on multimodal information, characterized in that, it includes the following steps: Step 1: Input the clothing image into the trained clothing image hierarchical encoding model. Use the clothing image hierarchical encoding model to process the clothing image to obtain clothing local information and clothing global information, and use the corresponding basis vector spaces to vectorize the clothing local information and clothing global information respectively to obtain the global information encoding of the clothing and the local information encoding of the clothing; Input the clothing sketch into the sketch encoding model to obtain the sketch encoding; Input the text into the input text encoding model to obtain the text encoding; The training method of the clothing image hierarchical encoding model includes the following steps: Step 1.1: Input the image x into the encoder E. The encoder E downsamples the image x to 1 / 4 of its original size to obtain a local feature map containing detailed information Then continue to downsample it to 1 / 8 of its original size to obtain a global feature map containing global information The method for obtaining Z global is as follows: For each vector at position (i, j) in the global feature map Z global , the quantizer q(·) finds the vector e global closest to it in V k through the nearest neighbor algorithm and replaces it to obtain Z global ; The method for obtaining Z local is as follows: Upsample Z global to the same size as Z local and perform a residual connection with it to obtain a new fused feature, and then find the vector e local closest to each position in this fused feature map in V k through the nearest neighbor algorithm and replace it to obtain Z local ; Step 1.2: Input the local feature map Z local into the local basis vector space V local to obtain the global information encoding Z of the clothing local ; then input the global feature map Z global into the global basis vector space V global to obtain the global information encoding Z of the clothing global ; Step 1.3: Encode the global information Z of the clothing local and the global information Z of the clothing global are input into the decoder D to obtain the reconstructed image: x' = D(Z local , Z global ) Step 1.

4. Calculate the total loss function \(L = L\) encoding +L fm , where \(L\) fm is the average of the loss values of all layers of the discriminator. The formula for calculating the loss value of the \(k\)-th layer of the discriminator is: where \(N\) is the number of features of the \(k\)-th layer of the discriminator, and \(\) is the \(i\)-th feature of the image \(x\) extracted by the \(k\)-th layer discriminator; Among them To calculate the L2 norm, ||·|| 2 To calculate the mean squared error loss, sg[] represents stopping the gradient propagation and stopping updating the parameters of this part; by minimizing the total loss function, the parameters of the encoder E, the local basis vector space V global , the global basis vector space V global , the decoder D and the discriminator are trained, and after training, a trained hierarchical encoding model of clothing images is obtained; Step 2: Combine the sketch encoding and the text encoding to form conditional information, connect it with the global information encoding of the clothing, and jointly input it into Transformer-G. The masked attention mechanism in Transformer-G automatically generates the global information sequence of the clothing; use the global information encoding of the clothing image as conditional information, connect it with the local information encoding of the clothing, and jointly input it into Transformer-L to automatically generate the local information sequence of the clothing; Among them, Transformer-L is used to generate the clothing detail information encoding, and Transformer-G is used to generate the clothing global information encoding; Step 3: Find the corresponding vectors of the generated clothing global information sequence and clothing local information sequence in the corresponding basis vector spaces respectively and convert them into a global two-dimensional feature map and a local two-dimensional feature map, and input the global two-dimensional feature map and the local two-dimensional feature map into the trained decoder D to generate the final image.

2. The method for generating an image from a clothing sketch based on multimodal information according to claim 1, characterized in that, in step 3, Transformer-G and Transformer-L are used to generate the clothing sequence by means of autoregressive prediction.

Citation Information

Cited By

  • 3D garment template generation method based on sketch driving

    CN121527312A