A heatmap-guided semantic disentanglement generative adversarial network and its clothing inspiration design method

By introducing a semantic de-entanglement generation adversarial network guided by heat maps, the problem of unsupervised generation of fashion clothing images is solved, high-fidelity fashion image generation and effective retention of attribute textures is achieved, and the inspiration transfer ability of intelligent clothing design is improved.

CN114970194BActive Publication Date: 2025-05-16HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210666001.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-05-16
Estimated Expiration
2042-06-14

AI Technical Summary

Technical Problem

The prior art has difficulty in adaptively learning and generating ‘inspired’ fashion clothing images in an unsupervised manner, especially in retaining the properties and texture of fashion image.

Method used

A semantic de-entanglement generation adversarial network guided by heat maps is proposed, including fashion clothing image encoder, generator, discriminator and local clothing image discrimination network. Through unsupervised learning, the fashion clothing image style characteristics of the source and target domains are integrated, and the mixed style fashion clothing image image is generated, and the texture matching degree is evaluated through heat maps.

Benefits of technology

It realizes the generation of high-fidelity fashion clothing images under unsupervised conditions, which can effectively retain the attributes and texture information of fashion images, and improves the inspiration transfer ability of intelligent clothing design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114970194B_ABST
    Figure CN114970194B_ABST
Patent Text Reader

Abstract

Disclosed is a heat map-guided semantic disentanglement generative adversarial network and a clothing inspiration design method thereof, belonging to the field of generative adversarial models and clothing design. The heat map-guided semantic disentanglement generative adversarial network includes a fashion clothing image encoder, a fashion clothing image generator, a fashion clothing image discriminator and a local clothing image discriminant network; the fashion clothing image encoder is used to capture the most distinguishing features of different input fashion items and disentangle the features into two key factors, namely attributes and textures; the fashion clothing image generator generates mixed-style fashion clothing images by using the attributes and textures encoded by the encoder; the fashion clothing image discriminator discriminates the authenticity of the generated image by using the clothing generated by the fashion clothing image generator; the local clothing image discriminant network introduces a local loss based on a heat map to evaluate the visual semantic matching degree between the generated fashion clothing image texture and the input fashion clothing image texture information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of generative adversarial models and clothing auxiliary design, and relates to a semantic disentanglement generative adversarial network guided by a heat map and a clothing auxiliary design method based on the network, wherein the method uses fashion clothing as the most original input. Background Art

[0002] Fashion designers can be inspired by anything visual, from ancient Roman architecture to a plate of baked beans. Talented designers seek inspiration from these visual objects and then design details and textures to incorporate into their fashion design collections. Designers often use raster or vector software such as Photoshop or Adobe Illustrator to complete the entire "inspiration" design process. However, these traditional tools can only produce perfect design images in the hands of experienced designers, but cannot automatically create "inspiration" clothing designs from existing visual objects. Therefore, it is necessary to design a tool that can automatically complete clothing design to assist fashion designers in completing the "inspiration" design process. Fortunately, the latest advances in big data analysis and deep learning technology have provided powerful tools for the realization of this "inspiration" design, making intelligent assisted design feasible.

[0003] Recent advances in deep learning techniques have made this kind of intelligent assisted design feasible, among which generative adversarial networks are a powerful tool for fashion image generation. Most current fashion work focuses on the research of conditional generative adversarial networks. These methods use a large amount of manually annotated data as conditional information for generative adversarial networks to design clothing. These conditional models rely heavily on the quality and quantity of labeled training data. Therefore, it is necessary to develop a method that adaptively learns to complete the design of fashion clothing images in an unsupervised manner. Unsupervised generative adversarial networks are mainly used to transfer low-level information of images and usually cannot represent high-level semantic information of images, resulting in difficulty in retaining semantic information of fashion image attributes and accurately capturing texture information. However, effectively disentangling the features of fashion clothing images, such as attributes and texture, is crucial for inspirational design in real life. In practice, intelligent fashion design should retain the intrinsic attributes of fashion clothing and automatically adapt to the learning to transfer the texture of any given fashion clothing image during the stylization process. Summary of the invention

[0004] The present invention relies on the existing generative adversarial network model and proposes a semantic disentangled generative adversarial network guided by heat map and its clothing image inspiration design method. The semantic disentangled generative adversarial network guided by heat map includes a fashion clothing image encoder, a fashion clothing image generator, a fashion clothing image discriminator and a local clothing image discriminant network, which aims to learn to integrate the feature representation of the style of fashion clothing images from the source domain and the fashion clothing images from the target domain in an unsupervised manner; the fashion clothing image encoder is used to capture the most distinguishing features of different input fashion items and disentangle the features into two key factors, namely attributes and textures; the fashion clothing image generator generates mixed-style fashion clothing images by using the attributes and textures encoded by the encoder; the fashion clothing image discriminator discriminates the authenticity of the generated image by using the clothing generated by the fashion clothing image generator; the local clothing image discriminant network introduces a local loss based on heat map to evaluate the visual semantic matching degree between the texture of the generated fashion clothing image and the texture information of the input fashion clothing image; the heat map refers to the area containing texture and attribute information of the clothing displayed in a highlighted form.

[0005] Furthermore, the fashion clothing image encoder includes an image feature extraction module and an image semantic disentanglement module; the image feature extraction module is used to deeply extract image features and extract effective pixel information; the image semantic disentanglement module is used to disentangle the image into attributes and textures, and generate a heat map for auxiliary information.

[0006] Furthermore, the image feature extraction module and the semantic disentanglement module include an image downsampling module, the first 47 residual blocks in Resnet152, a "sum" operation for generating a heat map of an input image, the last 3 residual blocks in Resnet 152 for generating textures and a global average pooling operation, and a convolution operation for generating attributes; the image semantic disentanglement module uses a global average pool to output the spatial average of the feature map of each unit after being convolved by 47 residual blocks, and uses global maximum pooling to output the spatial maximum of the feature map to evaluate the importance of the image in different areas; the fashion clothing image encoder adopts an encoding method based on a semantic disentanglement module, which decomposes the input fashion clothing image into independent factors and auxiliary information for generating a heat map, and the independent factors refer to attributes and textures.

[0007] Furthermore, the fashion clothing image generator adopts the generator structure of StyleGAN2, uses the texture and attributes generated by the fashion clothing image encoder, takes the attribute code as the constant input of StyleGAN2, and the texture code as the input of each StyleBlock of StyleGAN2 to synthesize the fashion clothing image.

[0008] Furthermore, the fashion clothing image discriminator adopts the discriminator architecture of StyleGAN2 to discriminate whether the generated image has the corresponding clothing semantics and the authenticity of the generated image. Further, the fashion clothing image discriminator is used to discriminate the score of the image feature matching between the image result generated by the fashion clothing image generator and the fashion image input, and the matching score result is used to update the fashion clothing image encoder, the fashion clothing image generator and the fashion clothing image discriminator to improve the authenticity of the generated result.

[0009] Furthermore, the local clothing image discrimination network is composed of a feature block encoder and a feature block discriminator; the feature block encoder is composed of five down-sampled residual blocks, a residual block for channel amplification and a convolution layer with a kernel size; the feature block encoder first randomly samples image blocks on the image results generated by the fashion clothing image generator and the input image, and then sequentially sends these randomly sampled image blocks to five down-sampled residual blocks, a residual block for channel amplification and a convolution layer with a kernel size to encode the randomly sampled image blocks into feature vectors; the residual block uses the same configuration as the fashion clothing image discriminator; the feature block discriminator adopts the discriminator architecture of StyleGAN2, and uses the feature blocks sampled by the feature block encoder to calculate the joint feature statistics to obtain the perceptual similarity value.

[0010] The clothing inspiration design method of the semantic disentanglement generative adversarial network guided by the heat map proposed in the present invention comprises the following steps: A. constructing a fashion clothing image data set, wherein the data set includes different fashion clothing types, textures and structures; B. designing a fashion clothing image encoder, wherein the fashion clothing image encoder is used to capture the most distinguishing features of different input fashion items, and disentangle the most distinguishing features into two key factors, namely attributes and textures; C. designing a fashion clothing image generator, wherein the fashion clothing image generator generates a mixed-style fashion clothing image through the attributes and textures; D. designing a fashion clothing image discriminator, wherein the fashion clothing image discriminator discriminates the authenticity of the generated image by using the clothing generated by the fashion clothing image generator; E. designing a local clothing image discriminant network, wherein the local clothing image discriminant network evaluates the visual semantic matching degree between the generated fashion clothing image texture and the input fashion clothing image texture information based on the local loss of the heat map; the heat map refers to the area where the clothing contains texture and attribute information in a highlighted form.

[0011] Furthermore, step A includes: A1, constructing a clothing image dataset of different categories and styles, integrating keyword search items of clothing e-commerce, including category, texture, style, color and detail information, deleting items with complex backgrounds, and constructing fashion item data; A2, constructing five categories of fashion clothing images for training and testing, including tops, bottoms, shoes, bags and hats; and randomly dividing the five categories of fashion clothing images into a training dataset and a test dataset.

[0012] The input of the semantically disentangled generative adversarial network guided by the heat map is a random fashion clothing image in the constructed fashion database. In one iteration, the fashion clothing image encoder first encodes the input fashion image of the source domain and the fashion image of the target domain into the attributes and texture of the source domain and the attributes and texture of the target domain respectively; then, the attributes and texture of the source domain are sent to the fashion clothing image generator to reconstruct the fashion image of the source domain. At the same time, the attributes of the source domain and the texture of the target domain are sent to the fashion clothing image generator to generate a mixed fashion clothing image that can maintain the attributes of the source domain and the texture of the target domain; the reconstructed source domain fashion image and the mixed fashion clothing picture are respectively sent to the fashion clothing image discriminator to learn the authenticity of the generated texture; at the same time, the target domain fashion clothing image and the mixed fashion clothing picture are also sent to the local clothing image discriminant network to evaluate the visual semantic matching degree between the texture of the generated mixed fashion clothing picture and the texture information of the input target domain fashion clothing image. In particular, reconstruction loss is used to ensure that the overall structure of the reconstructed fashion clothing image and the input source domain fashion clothing image remain consistent; perceptual loss is used to ensure that the semantic information of the reconstructed fashion clothing image is consistent with the semantic information of the source domain fashion clothing image; and style loss is used to avoid the occurrence of the checkerboard effect in the semantic information of the reconstructed fashion clothing image.

[0013] The beneficial effects of the present invention are as follows: the present invention proposes a fashion clothing image assisted design generation method of a semantic disentangled generative adversarial network guided by a heat map, which is intended to assist in the process of inspiration design. In particular, in order to utilize the attributes and texture information carried by the fashion image itself, a fashion clothing image encoder is introduced to complete the disentanglement of fashion clothing, and the attributes and textures of fashion clothing are disentangled respectively; in order to generate high-fidelity fashion images, a fashion clothing image generator and a fashion clothing image discriminator designed based on the StyleGAN2 architecture are used to generate reconstructed source domain fashion clothing images and mixed fashion clothing images. In order to make the texture of the mixed fashion clothing image consistent with the texture of the input target domain fashion clothing image, a local clothing image discriminant network is designed to introduce a local loss based on a heat map to evaluate the visual semantic matching degree between the generated mixed fashion clothing image texture and the input target domain fashion clothing image texture information. This framework has great potential in practical application fields such as fashion inspiration design and interactive design with users. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a flow chart of the method for designing and generating a semantic disentanglement generative adversarial network guided by a heat map of the present invention.

[0015] Figure 2 This is a model framework diagram of the semantic disentanglement generative adversarial network design generation method guided by a heat map of the present invention.

[0016] Figure 3 It is a design result diagram generated by the method of the present invention. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0018] Attached Figure 1 The flowchart of the fashion clothing image design generation method based on the generative adversarial network provided by the present invention is shown, which is described in detail as follows:

[0019] Step S1: Construct a dataset of clothing images of different categories and styles. The training data used in the present invention comes from the website www.ployvore.com. Users on the website can upload their own images, share and modify their own images, and other users can score and evaluate the images created by users. All images on the website have a clean picture background. The present invention mainly constructs information on fashion items of different categories, textures and styles. The training dataset contains five fashion-related categories, including tops, bottoms, shoes, bags and hats.

[0020] Step S2: Designing a fashion image encoder .like Figure 2 As shown in (a), based on the semantic disentanglement module ( Figure 2 The fashion clothing image encoder of the SDA module in (a) includes an image downsampling module and the first 47 residual blocks in Resnet152 for deep image feature extraction; the proposed semantic disentanglement module uses global average pooling to output the spatial average of the feature map of each unit after being convolved by 47 residual blocks; and uses global maximum pooling to output the spatial maximum of the feature map; by projecting the weights of the output layer onto the convolution feature map, it aims to evaluate the importance of the image in different regions, so that the input fashion clothing image is decomposed into independent factors, namely attributes and textures, as well as auxiliary information for generating heat maps. Specifically, Figure 2 As shown in (a), Indicates the source domain images, represents the first k Features (see Figure 2 (a) Feature tensor generated before the SDA module), Indicates the spatial position The value of the feature map at the position; then, for the k feature maps, perform global average pooling ( )get ,in and The range is Among them and Respectively represent k The height and weight of the feature map. Use global maximum pooling No. k The result of the feature map is expressed as ,in Indicates solution Then, the fully connected layer is trained to learn the weight of the kth feature map of the source image and ; Source domain image The attention score can be expressed as follows:

[0021] (1)

[0022] in, Indicates a connection operation; and Use and Represents a fully connected layer.

[0023] Similarly, the target domain Image inside The attention score can be expressed as:

[0024] (2)

[0025] For each fashion item, the model should focus on learning the properties of the source image and the texture of the target image, rather than the redundant information conveyed by the image background. By leveraging the attention scores from the source domain fashion apparel images to the target domain source fashion apparel images, the SDA-based encoder should identify the most discriminative regions of the image. The classification function for the source and target domains based on SDA is:

[0026] (3)

[0027] in, Representing images Randomly sampled from the original domain ; Representing images Randomly sampled from the target domain .

[0028] from Figure 2 As can be observed in (a), by using the semantic disentanglement module, the source domain fashion clothing image and the target domain fashion clothing images The difference areas are highlighted. To further generate the attention feature map for disentanglement of fashion clothing images, several additional residual blocks and global average pooling operations are used to generate texture (or ), and use the convolutional layer to generate attributes (or ). The feature maps generated by these attention mechanisms are embedded into the encoder E to focus on the most distinctive parts in different domains, thereby effectively migrating the texture of the target domain to the source domain and preserving the original attributes of the fashion clothing images in the source domain.

[0029] Step S3: Design a fashion clothing image generator Fashion Clothing Image Generator The goal is to map texture and attributes to a fashion image. Fashion Image Generator Adopting the generator framework of StyleGAN2, using the fashion clothing image encoder Generated textures and attributes, using attributes as a fashion clothing image generator Constant input of texture as fashion clothing image generator The input of each StyleBlock is used to synthesize fashion clothing images. In order to make the generated image retain the attributes of the source domain fashion image and learn the hybrid design result of the texture of the target domain fashion image, the attribute and texture To synthesize mixed fashion items, i.e. ,The reconstruction loss function of the source domain fashion clothing image can be expressed as:

[0030] (4)

[0031] In order to learn the perceptual similarity between images, the learned perceptual image patch similarity loss is used to optimize the SDA-based encoder and generator. The perceptual loss of the source domain fashion clothing images can be expressed as:

[0032] (5)

[0033] in, represents a perceptual feature extractor.

[0034] Step S4: Design a fashion clothing image discriminator : To distinguish the difference between generated images and real images, fashion clothing image discriminator The discriminator architecture of StyleGAN2 is used to judge whether the generated image has the corresponding clothing semantics and the authenticity of the generated image; for attribute code-based and the texture code The generated samples of synthetically reconstructed fashion clothing images, the objective function of the discriminator can be expressed as:

[0035] (6)

[0036] In addition, for the source domain fashion clothing images Properties and the target domain fashion clothing images Texture Mixed design results The adversarial loss can be expressed as:

[0037] (7)

[0038] Step S5: Design a local clothing image discrimination network. As shown in Figure (2) (b), the local clothing image discrimination network consists of a feature block encoder With feature patch discriminator Composition; Feature Block Encoder First, random sampling of image blocks is performed on the image results generated by the fashion clothing image generator and the input image, and then these randomly sampled image blocks are sequentially sent to five downsampling residual blocks, a residual block for channel amplification, and a convolutional layer with a kernel size to encode the randomly sampled image blocks into feature vectors; the feature block discriminator Using the discriminator architecture of StyleGAN2, using the feature block encoder The encoded feature vectors calculate the joint feature statistics of the random features to obtain the perceptual similarity values ​​of these features; the perceptual similarity values ​​are used to update the fashion image encoder, fashion image generator and local clothing image discriminant network to improve the authenticity of the generated results; the objective function based on the local clothing image discriminant network is defined as follows:

[0039] (8)

[0040] in, represents a combination of local blocks selected from a mixed fashion project, Indicates items from fashion The local block combination selected in .

[0041] The generator objective function that needs to be satisfied in the end is:

[0042] (9)

[0043] The design results generated by the method of the present invention are as follows Figure 3 shown.

[0044] The main contributions of the present invention are summarized as follows: a semantic disentangled generative adversarial network guided by heat map is proposed, which can automatically complete inspiration transfer interactively with users to realize intelligent clothing design. The semantic disentangled generative adversarial network guided by heat map includes a fashion clothing image encoder, a fashion clothing image generator, a fashion clothing image discriminator and a local clothing image discriminant network, which aims to learn to integrate the feature representation of the style of fashion clothing images from the source domain and the fashion clothing images from the target domain in an unsupervised manner; the fashion clothing image encoder is used to capture the most discriminative features of different input fashion items and disentangle the features into two key factors, namely attributes and textures; the fashion clothing image generator generates mixed-style fashion clothing images by using the attributes and textures encoded by the encoder; the fashion clothing image discriminator discriminates the authenticity of the generated images by using the clothing generated by the fashion clothing image generator; the local clothing image discriminant network introduces a patch loss based on heat map to evaluate the visual semantic matching degree between the generated fashion clothing image texture and the input fashion clothing image texture information; the proposed semantic disentangled generative adversarial network guided by heat map opens up a huge research space for later fashion auxiliary design.

[0045] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A heat map guided semantic disentanglement generative adversarial network, comprising a fashion image encoder, a fashion image generator, a fashion image discriminator and a local fashion image discriminant network, wherein the generative adversarial network learns the feature representation of the style of fashion images from a source domain and a fashion image from a target domain in an unsupervised manner based on unmatched data pairs; characterized in that, The fashion clothing image encoder is used to capture the most distinguishing features of different input fashion items, and disentangle the most distinguishing features into two key factors, namely attributes and textures; the fashion clothing image generator generates a mixed style fashion clothing image through the attributes and textures; the fashion clothing image discriminator discriminates the authenticity of the generated image by using the clothing generated by the fashion clothing image generator; the local clothing image discriminant network evaluates the visual semantic matching degree between the generated fashion clothing image texture and the input fashion clothing image texture information based on the local loss of the heat map; the heat map refers to the area where the clothing contains texture and attribute information in a highlighted form; The fashion clothing image encoder includes an image feature extraction module and an image semantic disentanglement module; the image feature extraction module is used to perform deep extraction of image features and extract effective pixel information; the image semantic disentanglement module is used to disentangle the image into attributes and textures, and generate a heat map for auxiliary information; The image feature extraction module and the semantic disentanglement module include an image downsampling module, the first 47 residual blocks in Resnet152, a "sum" operation for generating a heat map of an input image, the last 3 residual blocks in Resnet 152 for generating textures and a global average pooling operation, and a convolution operation for generating attributes; the image semantic disentanglement module uses a global average pool to output the spatial average of the feature map of each unit after being convolved by 47 residual blocks, and uses global maximum pooling to output the spatial maximum of the feature map to evaluate the importance of images in different areas; the fashion clothing image encoder adopts an encoding method based on a semantic disentanglement module, which decomposes the input fashion clothing image into independent factors and auxiliary information for generating a heat map, and the independent factors refer to attributes and textures.

2. The heat map guided semantic disentanglement generative adversarial network according to claim 1, characterized in that: The fashion clothing image generator adopts the generator structure of StyleGAN2, uses the texture and attributes generated by the fashion clothing image encoder, takes the attribute code as the constant input of StyleGAN2, and the texture code as the input of each StyleBlock of StyleGAN2 to synthesize the fashion clothing image.

3. The heat map guided semantic disentanglement generative adversarial network according to claim 1, characterized in that: The fashion clothing image discriminator adopts the discriminator architecture of StyleGAN2 to judge whether the generated image has the corresponding clothing semantics and the authenticity of the generated image.

4. The heat map guided semantic disentanglement generative adversarial network according to claim 1, characterized in that: The local clothing image discrimination network consists of a feature block encoder and a feature block discriminator; the feature block encoder consists of five down-sampling residual blocks, a residual block for channel amplification and a convolutional layer with a kernel size; the residual block uses the same configuration as the fashion clothing image discriminator; the feature block discriminator adopts the discriminator architecture of StyleGAN2 and uses the feature blocks sampled by the feature block encoder to calculate the joint feature statistics to obtain the perceptual similarity value.

5. A clothing inspiration design method based on a heat map-guided semantic disentangled generative adversarial network, wherein the generative adversarial network learns the feature representation of the style of fashion clothing images from a source domain and a fashion clothing image from a target domain in an unsupervised manner based on unmatched data pairs, characterized in that: Includes steps: A. Construct a fashion clothing image dataset, wherein the dataset includes different fashion clothing types, textures, and structures; B. Designing a fashion apparel image encoder, the fashion apparel image encoder is used to capture the most discriminative features of different input fashion items and disentangle the most discriminative features into two key factors, namely, attributes and textures; the fashion apparel image encoder includes an image feature extraction module and an image semantic disentanglement module; The image feature extraction module is used to deeply extract image features and extract effective pixel information; the image semantic disentanglement module is used to disentangle the image into attributes and textures, and generate a heat map for auxiliary information; C. designing a fashion clothing image generator, wherein the fashion clothing image generator generates a fashion clothing image of mixed style by using the attributes and textures; D. Designing a fashion clothing image discriminator, wherein the fashion clothing image discriminator discriminates the authenticity of the generated image by using the clothing generated by the fashion clothing image generator; E. Designing a local clothing image discrimination network, wherein the local clothing image discrimination network evaluates the degree of visual semantic matching between the generated fashion clothing image texture and the input fashion clothing image texture information based on the local loss of the heat map; the heat map refers to the area containing texture and attribute information of the clothing displayed in a highlighted form; Wherein, the step B comprises: B1. The fashion clothing image encoder includes an image feature extraction module and an image semantic disentanglement module; the image feature extraction module includes an image downsampling module and the first 47 residual blocks in Resnet152, which are used to deeply extract image features and extract main pixel information; B2. The image semantic disentanglement module uses global average pooling to output the spatial average of the feature map of each unit after being convolved by 47 residual blocks; uses global maximum pooling to output the spatial maximum of the feature map; and projects the weights of the output layer onto the convolutional feature map to evaluate the importance of the image in different regions, so that the input fashion clothing image is decomposed into independent factors, namely attributes and textures, as well as auxiliary information to generate heat maps.

6. The clothing inspiration design method according to claim 5, characterized in that: The step A comprises: A1. Build a dataset of clothing images of different categories and styles, integrate keyword search items of clothing e-commerce, including category, texture, style, color and detail information, delete items with complex backgrounds, and build fashion item data; A2. Construct five categories of fashion clothing images for training and testing, including tops, bottoms, shoes, bags and hats; randomly divide the five categories of fashion clothing images into a training data set and a test data set.

7. The clothing inspiration design method according to claim 5 or 6, characterized in that: In step C, the fashion clothing image generator adopts the generator structure of StyleGAN2, uses the texture and attributes generated by the fashion clothing image encoder, takes the attributes as constant inputs of StyleGAN2, and uses the texture as input of each StyleBlock of StyleGAN2 to synthesize the fashion clothing image.

8. The clothing inspiration design method according to claim 5 or 6, characterized in that: In step D, the fashion clothing image discriminator adopts the discriminator architecture of StyleGAN2 to discriminate the image result generated by the fashion clothing image generator and the image feature matching score of the fashion image input, and the matching score result is used to update the fashion clothing image encoder, the fashion clothing image generator and the fashion clothing image discriminator to improve the authenticity of the generated result.

9. The clothing inspiration design method according to claim 5 or 6, characterized in that: The step E comprises: E1, the local clothing image discrimination network is composed of a feature block encoder and a feature block discriminator; the feature block encoder first randomly samples image blocks on the image results generated by the fashion clothing image generator and the input image, and then sequentially sends these randomly sampled image blocks to five downsampling residual blocks, a residual block for channel amplification and a convolution layer with a kernel size to encode the randomly sampled image blocks into feature vectors; E2. The feature block discriminator adopts the discriminator architecture of StyleGAN2, and uses the feature vector encoded by the feature block encoder to calculate the joint feature statistics of random features to obtain the perceptual similarity values ​​of these features; the perceptual similarity values ​​are used to update the fashion clothing image encoder, the fashion clothing image generator and the local clothing image discriminant network to improve the authenticity of the generated results.

Citation Information

Patent Citations

  • Image-text cross-modal feature unentanglement method based on depth mutual information constraint

    CN110807122A

  • Clothing image generation system and method based on de-entanglement network

    CN113052230A