A Fine-Grained Fashion Text-Guided Garment Image Generation Method
By constructing the text-image feature encoding module, the fashion text and texture features are extracted for fine-grained fashion feature learning, the problem of missing fine-grained information in clothing image generation is solved, and the clothing image generation with more fashionable style and details is achieved.
Patent Information
- Application Number
- CN202410563661.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-05-08
AI Technical Summary
Existing clothing image generation methods are difficult to effectively capture the fine-grained key information in fashion text descriptions, resulting in the generated clothing images lacking fashion style details and the generated clothing images lack diversity and accuracy.
By constructing a text-image feature encoding module containing global-local encoder, text encoder, pattern encoder and image encoder, fashion text features and texture pattern features are extracted, fine-grained fashion feature learning, and clothing image generation is generated by combining global fashion descriptors, fine-grained text-image features and coarse-grained image features.
The diversity and accuracy of clothing image generation is improved, and the generated clothing image can better reflect fashion semantics and include fine-grained information such as style, color and accessories.
Smart Images

Figure CN118334160B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for generating clothing images guided by fine-grained fashion texts, belonging to the fields of computer vision and artificial intelligence. Background Art
[0002] Text-guided clothing image generation is widely used in the fashion industry, virtual fitting, online shopping and other fields. However, due to the complex semantics of fashion texts and the texture of clothing patterns, clothing image generation has become a difficult point. Based on the generative adversarial network method, a pyramid multi-level generation network architecture is proposed based on style transfer, which can unsupervised separate and generate high-level attributes in images, and can adjust the "style" of images in each convolutional layer according to latent codes during the process of text-guided clothing image generation, directly controlling the intensity of image features at different scales. For example, Lin <SIGGRAPH Special Interest Group for Computer GRAPHICS, 2023) uses two modalities of text and texture to change the appearance of clothing in the latent direction through a pre-trained generative adversarial model, and then controls the generation of model poses, clothing types and fabrics. Pernus <Conference on Computer Vision and Pattern Recognition, 2023) uses the Clip pre-trained language-image contrast model to embed the given input image into the latent space of the production adversarial model, and then performs text-conditioned operations based on the latent space to finally generate images related to the text. Zhang <ACM Multimedia, 2023) combines multiple modal information and the contrastive language-image pre-trained Clip model, and uses an encoder-decoder to guide the generation of clothing images. The above well-known methods mainly focus on multi-modal human pose recovery and the generation of clothing images of basic styles.
[0003] Due to the characteristics of fashionable clothing, such as diverse style types, complex color pattern clothing, and many fine-grained attributes, the existing well-known methods cannot effectively capture the fine-grained key information in fashion text descriptions, such as "collar", "long sleeves", "stripes and spots", etc., and it is difficult to ensure that the generated images can fully and accurately reflect the text descriptions. The present invention extracts fine-grained fashion text features from local fashion descriptors through a text encoder to obtain style, color, accessory and clothing detail information, ensuring the relevance of fine-grained fashion text-guided clothing image generation; at the same time, combined with fine-grained image features, fine-grained text-image features are obtained to ensure that the generated images are consistent in fashion semantics and have the fine-grained attributes of clothing. Summary of the Invention
[0004] The technical problem to be solved by the present invention is as follows: The present invention provides a method for generating clothing images guided by fine-grained fashion texts, which solves the problem that the generated clothing images lack fine-grained fashion style details, improves the diversity of clothing image generation, enables the generated clothing images to have more clothing detail information such as style, color, accessories, etc., and enhances the accuracy of clothing image generation.
[0005] The technical solution of the present invention is: A method for generating clothing images guided by fine-grained fashion texts. First, perform text-image feature extraction on the input multi-modal fashion clothing dataset to obtain fashion text features, texture pattern, and clothing image features respectively. Secondly, by constructing a text-image feature encoding module including a global-local encoder, a text encoder, a pattern encoder, and an image encoder, encode the fashion text, texture pattern, and clothing image features respectively to obtain a global fashion descriptor, fine-grained fashion text features, fine-grained image features, and coarse-grained image features. Then, perform fine-grained fashion feature learning on the fine-grained fashion text features and fine-grained image features to obtain fine-grained text-image features with fashion clothing detail information such as style, color, accessories, etc. Finally, combine the global fashion descriptor, fine-grained text-image features, and coarse-grained image features to generate clothing images and obtain new clothing images.
[0006] The method includes the following steps:
[0007] Step1. For the clothing images I k ,…,D d} and fashion texts T k ∈D k and fashion texts T k ∈D k in the input multi-modal fashion clothing dataset D = {D1,…,D text , where d represents the number of text-image pairs in the dataset, perform text-image feature extraction to obtain fashion text features f patch , texture pattern features f image and clothing image features f
[0008] Step2. By constructing a text-image feature encoding module including a global-local encoder, a text encoder, a pattern encoder, and an image encoder, encode the fashion text features f text , texture pattern features f patch , and clothing image features f image respectively to obtain a global fashion descriptor t global , fine-grained fashion text features f fine-text , fine-grained image features f fine-image and coarse-grained image features f coarse-image ;
[0009] Step3. Perform fine-grained fashion feature learning on the fine-grained fashion text feature f fine-text and the fine-grained image feature f fine-image to obtain a fine-grained text-image feature f Δs′ with fashion clothing detail information such as style, color, and accessories;
[0010] Step4. Combine the global fashion descriptor t global , the fine-grained text-image feature f Δs′ and the coarse-grained image feature f coarse-image to generate a new clothing image I'.
[0011] The specific content of Step2 is as follows:
[0012] First, perform descriptor tokenization on the fashion text feature f text and fill it with word embeddings of c∈R C×768 . Then, obtain the local fashion text descriptor t local =t {1:C},EOT ∈R (C-1)×768 and the global fashion text descriptor through the constructed global-local encoder, where the EOT ("end of text") component of t aggregates the global fashion text information, and C is the number of word texts;
[0013] Next, define the fashion text attributes as including three types of attributes: style, color, and accessories. Among them, the style attribute t local_style ={s1,…,s i1}, i1 represents the number of style attributes, and i1 = 20 (including 20 styles such as long sleeves, stand-up collars, short sleeves, and sleeveless); the color attribute t local_color ={c1,…,c i2}, i2 represents the number of color attributes, and i2 = 20 (including 20 colors such as blue, white, red, and black); the accessory attribute t local_accessory ={a1,…,a i3}, i3 represents the number of accessory attributes, and i3 = 17 (including 17 accessories such as buttons, pockets, and belts);
[0014] Finally, encode the local text descriptor t local through a text encoder that includes three types of text attributes to obtain the fine-grained fashion text feature f fine-text ;
[0015] The specific content of Step3 is as follows:
[0016] First, define the text-image cross-attention layer: Focus on the fashion text feature ftext and the fashion image feature f image , where the query Q = W Q ·Norm(f image ), W Q ∈R (h·d)×d , the key-value pair [K, V] = W KV ·Norm(f text ), W KV ∈R (h·d·2)×d , W Q and W KV are the text-image attention layer force parameters, h is the number of attention heads, and d is the dimension of each attention head;
[0017] Secondly, multiple text-image attention layers are used to perform text-image feature fusion on the fashion text feature f text and the fashion image feature f image , and the calculation formula is as follows:
[0018] where and are the attention parameters of the i-th text-image attention layer, and the text-image fusion feature f Δs is obtained;
[0019] Then, taking the fusion feature f Δs as the embedding of the conditional text-image feature where c, m, f represent coarse, medium, and fine granularities, i represents the conditional module, and the fine-grained fashion text feature f fine-text and the fine-grained image feature f fine-image are used as the embedding of the fusion text-image feature where c, m, f represent coarse, medium, and fine granularities, and s represents fusion;
[0020] Finally, by calculating the fine-grained text-image feature f Δs′ is obtained, where the linear transformation function Concat = σ n (W n σ n-1 (...σ1(W1 + b1)...)+b n ), W i and b i are the weight and bias of the i-th layer respectively, and σ i is the activation function of the i-th layer.
[0021] The beneficial effects of the present invention are:
[0022] 1. The known methods mainly generate clothing types and fabrics through two modalities of text and texture, without considering more detailed information about clothing, resulting in a single generated clothing type and blurred texture patterns. The present invention first extracts local fashion descriptors based on text features, then defines them as different fashion attributes, and finally obtains fine-grained fashion text features including three types of attributes: style, color, and accessories, so as to solve the problem of lacking detailed fashion style features in generated clothing images and improve the diversity of clothing image generation.
[0023] 2. The known methods adopt the Clip pre-trained language-image contrast model, which improves the accuracy of clothing images, but fails to combine image features well, resulting in the ability to generate only clothing images with single solid-color texture patterns. The present invention defines a text-image cross-attention layer to pay attention to fashion text features and clothing image features, and obtains fine-grained text-image features through fine-grained fashion feature learning, which can generate different texture patterns and accessory details, and improve the accuracy of clothing image generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flowchart of the present invention;
[0025] Figure 2 is a flowchart of the text-image feature encoding of the present invention;
[0026] Figure 3 is a flowchart of the fine-grained text-image features of the present invention;
[0027] Figure 4 is a clothing generation image guided by the fine-grained fashion text of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0028] Example 1: As Figures 1 - 4 shown, a method for generating a clothing image guided by fine-grained text, the method includes:
[0029] Step1. For the clothing images I k ,…,D d} and fashion texts T k ∈D k and fashion texts T k ∈D k in the input multi-modal fashion clothing dataset D = {D1,…,D text , where d represents the number of text-image pairs in the dataset, perform text-image feature extraction to obtain fashion text features f patch , texture pattern features f image and clothing image features f
[0030] Step 2: By constructing a text-image feature encoding module that includes a global-local encoder, a text encoder, a pattern encoder, and an image encoder, respectively encode the fashion text feature f text , the texture pattern feature f patch , and the clothing image feature f image to obtain the global fashion descriptor t global , the fine-grained fashion text feature f fine-text , the fine-grained image feature f fine-image , and the coarse-grained image feature f coarse-image ;
[0031] Step 3: Perform fine-grained fashion feature learning on the fine-grained fashion text feature f fine-text and the fine-grained image feature f fine-image to obtain the fine-grained text-image feature f Δs′ with fashion clothing detail information such as style, color, and accessories;
[0032] Step 4: Combine the global fashion descriptor t global , the fine-grained text-image feature f Δs′ , and the coarse-grained image feature f coarse-image to generate a clothing image and obtain a new clothing image I'.
[0033] The specific content of Step 2 is as follows:
[0034] First, perform descriptor tokenization on the fashion text feature f text , pad it into word embeddings of c ∈ R C×768 , and respectively obtain the local fashion text descriptor t local = t {1:C},EOT ∈ R (C-1)×768 and the global fashion text descriptor where the EOT ("end of text") component of t aggregates the global fashion text information, and C is the number of word texts.
[0035] Then, define the fashion text attributes as including three types of attributes: style, color, and accessories. Among them, the style attribute t local_style = {s1,..., s i1}, i1 represents the number of style attributes, and i1 = 20 (including 20 styles such as long sleeves, stand-up collars, short sleeves, and no sleeves); the color attribute t local_color = {c1,..., c i2}, i2 represents the number of color attributes, and i2 = 20 (including 20 colors such as blue, white, red, and black); the accessory attribute t local_accessory = {a1,..., a i3Let \(i_3\) represent the number of accessory attributes, where \(i_3 = 17\) (including 17 kinds of accessories such as buttons, pockets, belts, etc.);
[0036] Finally, by constructing a text encoder that includes three types of text attributes, the local text descriptor \(t\) local is encoded to obtain the fine-grained fashion text feature \(f\) fine-text .
[0037] The specific steps of Step 3 are as follows:
[0038] First, by defining a text-image cross-attention layer: Attention is paid to the fashion text feature \(f\) text and the clothing image feature \(f\) image , where the query \(Q = W\) Q ·Norm(\(f\) image ), \(W\) Q ∈\(R\) (h·d)×d , the key-value \([K, V]=W\) KV ·Norm(\(f\) text ), \(W\) KV ∈\(R\) (h·d·2)×d , \(W\) Q and \(W\) KV are the parameters of the text-image attention layer, \(h\) is the number of attention heads, and \(d\) is the dimension of each attention head.
[0039] Secondly, multiple text-image attention layers are used to fuse the fashion text feature \(f\) text and the clothing image feature \(f\) image to perform text-image feature fusion. The calculation formula is as follows:
[0040] where and are the attention parameters of the \(i\)-th text-image attention layer, and the text-image fusion feature \(f\) Δs is obtained.
[0041] Then, using the fusion feature \(f\) Δs as the embedding of the conditional text-image feature where \(c\), \(m\), \(f\) represent coarse, medium, and fine granularity, \(i\) represents the conditional module, and the fine-grained fashion text feature \(f\) fine-text and the fine-grained image feature \(f\) fine-image are used as the embedding of the fused text-image feature where \(c\), \(m\), \(f\) represent coarse, medium, and fine granularity, and \(s\) represents fusion.
[0042] Finally, by calculating the fine-grained text-image feature \(f\) Δs′ is obtained, where the linear transformation function Concat = \(\sigma\) n(W n σ n-1 (...σ1(W1 + b1)...) + b n ),W i and b i are the weight and bias of the i-th layer respectively, and σ i is the activation function of the i-th layer.
[0043] Example 2: As Figure 1 shown, a fine-grained fashion text-guided clothing image generation method, the specific steps of this method are as follows:
[0044] Step1. First, for the clothing images I k , …, D d} in the input multi-modal fashion clothing dataset D = {D1, …, D k ∈ D k and the fashion text T k ∈ D k , where d represents the number of text-image pairs in the dataset, the fashion text feature f text is obtained through the pre-trained Clip text encoder composed of Bert. Secondly, the clothing image passes through the pre-trained VGGNet and Clip-VisionTransformer to obtain the texture pattern feature f patch , and the clothing image feature f image .
[0045] Step2. First, as Figure 2 shown, according to the obtained text fashion text feature f text , the Clip text encoder is used to perform descriptor tokenization on it and pad it into word embeddings of c ∈ R C×768 , where C = 77 represents the maximum number of words selected. Features are extracted through the pre-trained global-local encoder, and the attention layer is used to process the word embeddings to obtain the local fashion text descriptor t local = t {1:C},EOT ∈ R (C-1)×768 and the global fashion text descriptor t global ∈ R 768 , where the EOT ("end of text") component of t aggregates the global fashion text information, and C is the number of word texts. As described in Table 1, the fashion text attributes of the local fashion text descriptor t local are defined as including 3 types of attributes: style, color, and accessories. Among them, the style attribute t local_style = {s1, …, s i1}, i1 represents the number of style attributes, and i1 = 20 (including 20 styles such as long sleeves, stand-up collars, short sleeves, and sleeveless); the color attribute t local_color={c1,…,c i2}, i2 represents the number of color attributes, where i2 = 20 (including 20 colors such as blue, white, red, black, etc.); the accessory attribute t local_accessory ={a1,…,a i3} i3 represents the number of accessory attributes, where i3 = 17 (including 17 accessories such as buttons, pockets, belts, etc.);
[0046] By constructing a Clip text encoder that includes 3 types of text attributes, tokenize the local text descriptor t local to get X = E token (t local ), where X is the token embedding vector matrix, and E token is the token embedding function that maps it to the tokenized space, and after using the transformer encoding and normalization operation X' = LN(Transformer(X)), where LN is the layer normalization operation, to obtain the fine-grained fashion text feature where represents the projection matrix of the text feature, which is used to project the encoded feature into the desired dimensional space.
[0047] Table 1
[0048]
[0049]
[0050] Secondly, the clothing texture pattern feature f patch obtains the fine-grained image feature f fine-image through the pre-trained VGG-19 network. To emphasize the reference clothing texture pattern, introduce the calculation of the correlation of the RGB texture image feature where the Gram matrix function
[0051] Finally, the clothing image feature f image passes through the masked Clip encoder, and during the training process, add a masked-level contrastive language-image pre-training module to explicitly learn the correspondence between the word embedding and the visual part through the mask, so as to obtain the coarse-grained image feature f coarse-image .
[0052] Step3. First, as Figure 3 shown, combine the obtained fashion text feature f text and the clothing image feature f image through a custom text-image cross layer: where the query Q = W Q ·Norm(f image ), W Q ∈R(h ·d)×d The key-value pair [K, V] = W KV ·Norm(f text ), W KV ∈R (h·d·2)×d , W Q and W KV are the text-image attention layer force parameters, h is the number of attention heads, and d is the dimension of each attention head, to focus on text-image. The multi-head format is adopted in the custom text-image cross layer for the operation of the attention mechanism: Q = rearrange(Q', b(hd)xy -> (bh)(xy)d′), K = rearrange(K', bn(hd) -> (bh)nd′), V = rearrange(V', bn(hd) -> (bh)nd′), where rearrange is the rearrangement function.
[0053] Secondly, multiple text-image attention layers are used to fuse the fashion text feature f text and the clothing image feature f image for text-image feature fusion, and the calculation form is as follows:
[0054] where and are the attention parameters of the i-th text-image attention layer. In this process, the text-image similarity calculation is introduced so that this module can better capture features from text and images, using the adaptive layer where d1 is the adaptive text and image, norm is the normalization function, and the similarity between text and image is measured by the modified cosine distance: where B is the size of the batch, v b is the b-th adapted feature vector, and the text-image fusion feature f Δs is obtained.
[0055] Then, with the fusion feature f Δs through the inversion e4e encoder containing 3 sub-modules of coarse, medium, and fine, each sub-module contains 3 fully connected layers, and the fusion feature f Δs is converted into the embedding of the conditional text-image feature where c, m, f represent coarse, medium, and fine granularity, and i represents the conditional module. The fine-grained fashion text feature f fine-text and the fine-grained image feature f fine-image are also used as the embedding of the fusion text-image feature through the inversion e4e encoder containing 3 sub-modules of coarse, medium, and fine, and each sub-module contains 6 fully connected layers where c, m, f represent coarse, medium, and fine granularity, and s represents fusion.
[0056] Finally, by calculating the fine-grained text-image feature f is obtained Δs′ , where the linear transformation function Concat = σ n (W n σ n-1 (...σ1(W1 + b1)...) + b n ), W i and b i are the weights and biases of the i-th layer respectively, and σ i is the activation function of the i-th layer. To enhance fine-grained fashion feature learning, a complete objective function containing two losses is added, which can be expressed as:
[0057] L = L rec + L sim = f Δs′ - f Δs + 1 - cos(f Δs′ , f Δs ), where L rec is the L-2 distance reconstruction loss adding the supervised learning editing direction f Δs′ , and L sim is the cosine similarity loss explicitly encouraging the network to minimize the cosine direction of the predicted embedding direction f Δs′ and f Δs in the latent space.
[0058] Step4. First, the obtained global text descriptor, fine-grained text-image feature, and coarse-grained image feature are used as latent codes and input into the generator G(t global , f Δs′ , f coarse-image ) of the generative adversarial network. On each attention block, a separate cross-attention mechanism is added where the query Q = W Q ·Norm(f image ), W Q ∈R (h·d)×d , the key-value pair [K, V] = W KV ·Norm(t global ), W KV ∈R (h·d·2)×d , W Q and W KV are the text-image attention layer force parameters, h is the number of attention heads, and d is the dimension of each attention head, to focus on the local text descriptor, and through the attention layer:
[0059] where represents the cross-attention, self-attention, and weight modulation layers of the l-th layer.
[0060] Secondly, the generator output is a multi-scale image pyramid with l = 5 levels, and the multi-scale image pyramid is defined as The resolution of the generated image for each layer is The discriminator consists of two parts: text processing t D and image processing φ. The Clip text encoder is used to extract t from the fashion text D to process the text, and the Transformer image encoder φ is used to process the image. The function is used: where V GAN is the standard non-saturating generative adversarial network loss, D ij (I, T) = ψ j (φ i→j (x i ), t D ) + Conv 1×1 (φ i→j (x i ))), ψ j is a 4-layer 1×1 modulated convolution, Conv 1×1 is the skip connection, V match is the discriminator conditional function, which compares the features of the two functions. The function t D processing the text is integrated into the discriminator as a modulating effect
[0061] Finally, each layer of the processing pyramid structure is taken on the image, and at each level x i makes true / false predictions on the multi-scales where i < j ≤ L. To extract features at different scales, through image processing each sub-network φ i→j is a subset of the entire network , where i > 0 indicates entering the later stage, and j < L indicates early exit. Each layer in φ consists of self-attention and convolution. The last layer flattens the spatial extent into a 1×1 tensor, producing the output resolution, as Figure 4 shown, and finally a more realistic and natural-looking clothing image is obtained through continuous upsampling processing
[0062] In summary, compared with the related well-known methods, the method of the present invention generates more realistic and natural clothing with a finer-grained fashion style
[0063] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Without departing from the spirit of the present invention, various changes can be made within the scope of knowledge possessed by those of ordinary skill in the art
Claims
1. A fine-grained fashion text-guided clothing image generation method, characterized in that: Step 1: Extract text-image features from the input multi-modal fashion clothing dataset to obtain fashion text features, texture pattern features, and clothing image features respectively; Step 2: By constructing a text-image feature encoding module including a global-local encoder, a text encoder, a pattern encoder, and an image encoder, encode the fashion text, texture pattern, and clothing image features respectively to obtain a global fashion text descriptor, fine-grained fashion text features, fine-grained image features, and coarse-grained image features; Step 3: Perform fine-grained fashion feature learning on the fine-grained fashion text features and fine-grained image features to obtain fine-grained text-image features with fashion clothing detail information, where the fashion clothing detail information includes style, color, and accessories; Step 4: Combine the global fashion text descriptor, fine-grained text-image features, and coarse-grained image features to generate a clothing image to obtain a new clothing image; The specific steps of Step 1 are as follows: First, for the clothing images and fashion texts in the input multi-modal fashion clothing dataset, the fashion texts are passed through a pre-trained Clip text encoder composed of Bert to obtain the fashion text feature f text , secondly, the clothing images are passed through a pre-trained VGGNet and Clip-VisionTransformer to obtain the texture pattern feature f patch and the clothing image feature f image ; The specific steps of Step 2 are as follows: First, the obtained fashion text feature f text is tokenized into descriptors using a Clip text encoder, padded into word embeddings, features are extracted through a pre-trained global-local encoder, and an attention layer is used to process the word embeddings to obtain the local fashion text descriptor t local and the global fashion text descriptor t global . The fashion text attributes of the local fashion text descriptor t local are defined as including three types of attributes: style, color, and accessories; By constructing a Clip text encoder that includes three types of text attributes, tokenize the local fashion text descriptor t local perform tokenization, map it to the tokenized space, and use the transformer encoding and normalization operations to obtain fine-grained fashion text features; Secondly, the texture pattern feature f patch obtains the fine-grained image feature f through the pre-trained VGG-19 network fine-image , and in order to emphasize the reference clothing texture pattern, the correlation of calculating the RGB texture image feature is introduced; Finally, the clothing image feature f image Through the masked Clip encoder, a masked contrastive language-image pre-training module is added during training to explicitly learn the correspondence between word embeddings and the visual part through masking, thereby obtaining the coarse-grained image feature f coarse-image ; The specific steps of Step 3 are as follows: First, define a text-image cross-attention layer to focus on fashion text features and clothing image features, and adopt a multi-head format in the text-image cross-attention layer for the operation of the attention mechanism; Secondly, multiple text-image cross-attention layers are adopted to fuse the fashion text feature f text and the clothing image feature f image to obtain the text-image fusion feature f Δs ; Then, with the fused feature f Δs Through the inversion e4e encoder containing 3 sub-modules of coarse, medium, and fine, each sub-module containing 3 fully connected layers, the fused feature f Δs is transformed into the embedding of conditional text-image features, the fine-grained fashion text feature f fine-text and the fine-grained image feature f fine-image are also used as the embedding of fused text-image features through the inversion e4e encoder containing 3 sub-modules of coarse, medium, and fine, with each sub-module containing 6 fully connected layers; Finally, the fine-grained text-image feature f is obtained by calculating through the linear transformation function Concat Δs′ .
Citation Information
Patent Citations
A fine-grained classification method for fashion women's wear images based on component detection and visual features
CN109145947A
Sectional fine-grained commodity image description generation method and device and medium
CN115953590A