A real scene image editing method based on hierarchical classification text guidance
Patent Information
- Application Number
- CN202310793941.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-06-30
AI Technical Summary
大多方法对于真实应用来说,操作过于繁琐,需要通过输入具体的文本描述实现对于图像内容的操控
[0013]1.通过训练层级多标签分类模型,对输入的风格描述文本进行层级分类,将抽象的词汇转为具体的文本描述,一是作为映射网络训练的选择,二是作为CLIP模型的文本输入,达到使模型更加自动化训练的目的,即不需要太多人为的操控。
Smart Images

Figure CN116912362B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of hierarchical semantic representation of generative adversarial networks (GANs), inverse mapping of images, hierarchical text classification, and text-guided image editing. Specifically, it refers to a method in image generation models that, by inputting abstract-style text, edits the latent vectors of inverse-mapped real-world indoor scene images using a hierarchical text classification model and a contrastive language image pre-training model (CLIP). Background Technology
[0002] Existing image editing methods include using a pre-trained classifier to learn a boundary, combined with a generative adversarial network (GAN) to manipulate the image by shifting the latent vector along a certain direction. However, this method largely relies on the assumption of complete decoupling of the latent space and requires manual adjustment of parameters such as manipulation intensity. Other methods propose editing specific regions of an image by manipulating matching positions in a style map; that is, selecting a location in the desired image and using an image generation network to synthesize a new image. However, this method requires manually selecting the area to be modified, which is cumbersome. Recently, some methods have also achieved good results by controlling facial image changes through text, due to the relatively simple structure of the human face.
[0003] Recent advancements and increased attention have been made in text-guided image editing. TediGAN maps images and text to a shared StyleGAN latent space, leveraging text to manipulate the latent vectors of the image. FEAT introduces an attention module that matches input text with the image, learns an attention mask, and utilizes a generative adversarial network to achieve text-guided image editing. Due to the popularity of diffusion models, several text-guided image editing methods based on denoising diffusion models have also achieved good results, such as DALLE and DiffusionCLIP, further improving the performance of text-to-image generation.
[0004] In recent years, Generative Adversarial Networks (GANs) have developed rapidly and achieved great success in the field of high-quality image generation. Specifically, StyleGAN is one of the well-known GAN models capable of generating high-fidelity images. Furthermore, research has found that StyleGAN provides a semantically rich latent space, with different network layers corresponding to different semantics. As mentioned in HiGAN, in scene image generation, the lower layers of StyleGAN control the composition of layout, followed by objects, then attributes, while the higher layers control color. Moreover, these hierarchical latent spaces possess disentanglement properties. This allows us to use pre-trained models to perform editing operations on both synthetic and real images.
[0005] In summary, current text-guided image editing methods still have some problems. Most methods are too cumbersome for real-world applications, requiring the input of specific text descriptions to manipulate image content. Furthermore, due to the complexity and diversity of scene images, research on text-guided editing of real-world scene images is relatively limited, with most methods focusing on facial images. Therefore, this invention proposes a method for real-world scene image editing based on hierarchical classification and text guidance. This method utilizes the recently proposed CLIP model to achieve intuitive text-based image manipulation, which requires neither pre-training of the operation direction nor manual selection of the image position to be manipulated. The CLIP model is a model pre-trained using 400 million pairs of image-text data from the internet. Since natural language can express a wider range of visual concepts, this method combines the hierarchical semantic features of CLIP and StyleGAN, using a hierarchical text classification model to classify the input text description hierarchically. The classification results are then applied to the hierarchical training of the StyleGAN mapping network and the semantic control of real-world images, achieving more automated manipulation of real-world scene images through abstract text descriptions. Summary of the Invention
[0006] To address the aforementioned problems, this invention proposes a method for editing real-world images based on hierarchical classification text guidance. Starting from the hierarchical semantic representation of StyleGAN and the manipulation of images through text, a method for editing real-world indoor scenes across modalities is designed. This method can use an abstract style text description to imbue real-world indoor scenes with the characteristics of that style while preserving their inherent attributes, making it applicable to practical applications such as interior design. The technical solution of this invention includes the following steps:
[0007] Step 1: Select a hierarchical multi-label text classification model, input the first-level vocabulary t1 into the model, and perform hierarchical classification of the interior style description. The model output has three levels: the first-level vocabulary t1 is an abstract style description, the second-level vocabulary t2 is a compositional description of the scene image, and the third-level vocabulary t3 is a detailed description corresponding to the abstract style;
[0008] The composition description includes layout, objects, attributes, and colors;
[0009] The detailed description includes specific descriptions of the layout, objects, attributes, and colors;
[0010] Step 2: Use the e4e inversion model to obtain the latent vector w, w∈W+, of the indoor images trained in the LSUN dataset, where W+ represents the vector space; and based on the semantic hierarchical characteristics of StyleGAN, combine the second-level vocabulary t2 obtained in Step 1 to segment the latent vector w.
[0011] Step 3: Train multiple latent space residual mappers. Since different StyleGAN layers are known to generate different levels of detail in scene images, the multiple latent space residual mappers are divided into four groups, each with a separate mapper. These four groups correspond to the generation of layout, objects, attributes, and color details in the scene image, respectively. Using the three-level vocabulary t3 and CLIP model obtained in Step 1, intuitive abstract text manipulation of real-world scene images is achieved.
[0012] The beneficial effects of this invention are as follows:
[0013] 1. By training a hierarchical multi-label classification model, the input style description text is classified hierarchically, transforming abstract words into concrete text descriptions. This serves two purposes: first, as a choice for training the mapping network; and second, as text input for the CLIP model, thereby achieving the goal of making the model more automated in training, i.e., requiring less human intervention.
[0014] 2. By leveraging the hierarchical semantic representation of StyleGAN, different mapping networks are trained for the different semantics of the scene images corresponding to different layers. This achieves the goal of training only the mapping network that needs to change the semantics of the input image, while keeping other elements of the input image unchanged, thereby improving the efficiency of model training and reducing the resources required. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating the implementation of the method of this invention;
[0016] Figure 2 This is a schematic diagram of the present invention. Detailed Implementation
[0017] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0018] like Figure 1 and 2 As shown, this invention proposes a real-scene image editing method based on hierarchical classification text guidance. This method uses a hierarchical multi-label text classification model to convert abstract text descriptions into concrete text descriptions. Based on the hierarchical semantic representation features of StyleGAN, an automated text-manipulated image editing network model is designed.
[0019] This invention first selects a hierarchical multi-label text classification model to hierarchically classify the input style description text, obtaining an expansion from abstract words to concrete words. The e4e inversion model is used to obtain latent vectors of indoor scene images. Based on the semantic hierarchical characteristics of StyleGAN, the latent vectors are divided. A latent space residual mapper is trained, divided into four groups representing the generation of layout, objects, attributes, and color details in the scene image. The mapping model can be selectively trained using the second-level vocabulary obtained from the text hierarchical model. The third-level vocabulary obtained from the text classification model is input into the CLIP network, and the CLIP loss is used to control the training of the mapping network. After the latent vectors are hierarchically input into the mapping network, a bias vector is obtained. This bias vector is summed with the original vector and then input into StyleGAN to obtain the edited image.
[0020] The specific implementation steps of this invention are as follows:
[0021] Step 1: Select a hierarchical multi-label text classification model. The input text is a description of interior design styles, such as Scandinavian style, Chinese style, minimalist style, etc. After training the model, the results are as follows: Figure 2 The three-level text classification structure shown has the following layers: the first-level vocabulary is the abstract style description t1 of the input; the second-level vocabulary t2 consists of different constituent elements of the scene image, which in this method include "layout", "object", "attribute" and "color"; and the third-level vocabulary t3 is the specific description of the abstract style, such as "screen", "raw wood" and "reddish brown" corresponding to the Chinese style.
[0022] EURLEX57K is a large hierarchical multi-label text classification dataset containing 57k English EU legislative documents and approximately 4.3k European vocabulary tags. The tag set is divided into zero-shot tags, few-shot tags, and frequent tags. Few-shot tags are those with a frequency of 50 or less in the training set, while frequent tags are those with a frequency greater than 50 in the training set. This method utilizes the EURLEX57K dataset to train a hierarchical multi-label text classification model.
[0023] The specific implementation steps for step 1 are as follows:
[0024] Experimental setup: We trained the model using the EURLEX57K dataset, with a training set size of 45,000, a validation set size of 6,000, and a test set size of 6,000. The learning rate was set to 1e-4.
[0025] 1-1. Based on graph convolutional networks, using a text encoder and a label encoder, text semantics S are extracted by sharing the hierarchical structural relation representation E learned in the label set. t and tag semantics S lAs shown in the formula below, where V t The set of nodes representing the hierarchical structure is obtained by taking the text description as input, passing the text features T obtained through bidirectional GRU and CNN layers, and then performing a linear transformation; V l The set of label nodes is obtained by averaging the pre-trained label inputs with the label set as input, and σ is the activation function ReLU.
[0026] S t =σ(E·V) t )
[0027] S l =σ(E·V) l )
[0028] 1-2. Transform the text semantics S t and tag semantics S l Projecting to a joint embedding space, the joint embedding loss controls the text semantics S t and tag semantics S l The similarity.
[0029] 1-3 Through matching learning loss, fine-grained label semantics, coarse-grained label semantics, and incorrect label semantics are obtained through training. Among them, the fine-grained label semantics are closest to the input third-level vocabulary t3; that is, the fine-grained label semantics are t3, the coarse-grained label semantics are t2, and the other incorrect label semantics are far away from the first-level vocabulary t1.
[0030] The overall loss function is:
[0031] L = L cls (y,y′)+σ1L joint +σ2L match
[0032] Among them, L cls Let y be the cross-entropy class loss, and y′ represent the true label and the probability of the label, respectively. jojnt For joint embedding loss, L match For hierarchical aware matching loss, σ1 and σ2 are hyperparameters used to balance the joint embedding loss and hierarchical aware matching loss.
[0033] 1-4. Using a trained hierarchical multi-label text classification model, input the first-level vocabulary t1 to obtain the required third-level vocabulary t3 and second-level vocabulary t2.
[0034] Step 2: Based on the e4e model trained on the LSUN dataset, the real images are mapped to the w latent space (w∈W+). Taking advantage of the hierarchical semantic interpretability of StyleGAN, that is, the generated semantics in the images corresponding to different network layers are different (as verified in HiGAN), the inverted latent space is divided.
[0035] LSUN is a large-scale image dataset for scene understanding, containing images for 10 scene categories and 20 object categories. The scene categories primarily include images of bedrooms, living rooms, and classrooms. The training data contains a large number of images for each category, ranging from approximately 120,000 to 3,000,000. The validation data includes 300 images, and the test data contains 1000 images for each category.
[0036] The specific implementation steps for step 2 are as follows:
[0037] 2-1. Using the e4e model trained on the LSUN dataset, obtain the inversion latent vector w of the real indoor scene in .pt format, which will be used as input to StyleGAN.
[0038] 2-2. Based on the semantic hierarchical characteristics of StyleGAN, the obtained latent vector w is divided into layers [0, 2) of the generator network, layers [2, 6) of the generator network for the object, layers [6, 12) of the generator network for the attribute, and layers [12, 14) of the generator network for the color.
[0039] Step 3: Train the latent space residual mapper, combine the CLIP model and StyleGAN2 to obtain a new image after editing the real scene image through abstract text description.
[0040] Experimental parameter settings: The training set size was set to 5000, and the test set size was set to 1000. The batch size during training was set to 2, and the batch size during testing was set to 1. The learning rate was set to 0.5.
[0041] The specific implementation steps are as follows:
[0042] 3-1. Since it has been shown that different StyleGAN layers are responsible for generating different levels of detail in the scene image, the four latent space residual mappers are divided into four groups, respectively for layout, object, attribute, and color, and a different part of the latent vector w is provided to each group.
[0043] Based on the second-level vocabulary t2 obtained in step 1, each group of latent space residual mappers is selectively trained. The latent space residual mappers corresponding to words not included in the second-level vocabulary t2 do not need to be trained.
[0044] 3-2. Represent the latent vector of the input image as w = (w l w o w p w c ,w0), where w l w o w p w c w0 represents the partitioning of w based on different layers, where w l This corresponds to the vector part of the layout layer, w o This corresponds to the vector part of the object layer, w. p This corresponds to the vector part of the attribute layer, w c This corresponds to the vector part of the color layer. w0 represents the remaining part after dividing the latent vector w. Since the StyleGAN network has 18 layers, the grouping is the first 14 layers. After passing through the latent space residual mapper, we can obtain M(w) = (M1(w)). l M2(w) o M3(w) p M4(w) c ), w0), where M1, M2, M3, and M4 represent the groups of the mapping network, respectively.
[0045] 3-3. After training with CLIP loss, the latent space residual mapper multiplies the obtained bias vector Δ with the initial latent vector w of the image, thereby editing the latent vector w while preserving other semantic content in the input image. CLIP loss minimizes the cosine distance between the generated image and the text prompt.
[0046] L CLIP (w)=D CLIP (G(w+M(w)), t3),
[0047] Here, G represents the StyleGAN generator. To preserve some visual attributes of the original input image, the latent space after mapping is constrained using the L2 norm. Therefore, the final loss function is as follows:
[0048] L(w)=αL CLIP (w)+β||M(w)||2,
[0049] Where α is the weight of the CLIP loss and β is the weight of the L2 norm loss. In the experiment, α = 1 and β = 0.8.
[0050] 3-4. Input the edited latent vector w+M(w) into the StyleGAN network, and finally output the edited image.
[0051] This invention utilizes the semantic hierarchical characteristics of StyleGAN and its hierarchical multi-label text classification model to map abstract words to concrete words, enabling automated editing of text-guided images and reducing manual intervention. Selectively training the mapping network also improves training efficiency, shortening training time and avoiding unnecessary resource waste.
Claims
1. A method for editing realistic scene images based on hierarchical classification text guidance, characterized in that, Includes the following steps: Step 1: Select a hierarchical multi-label text classification model and classify the first-level vocabulary. Input this model to perform hierarchical classification of interior style descriptions; the model's output has three levels: first-level vocabulary. For abstract style description, second-level vocabulary Description of the composition of scene images and third-level vocabulary A detailed description corresponding to the abstract style; The composition description includes layout, objects, attributes, and colors; The detailed description includes specific descriptions of the layout, objects, attributes, and colors; Step 2: Use the e4e inversion model to obtain the latent vector w of the indoor images trained in the LSUN dataset. +, + represents the vector space; and based on the semantic hierarchical characteristics of StyleGAN, combined with the second-level vocabulary obtained in step 1. Segment the potential vector w; Step 3: Train multiple latent space residual mappers; since different StyleGAN layers are known to generate different levels of detail in the scene image, the multiple latent space residual mappers are divided into four groups, one for each group, with the four groups corresponding to the generation of layout, objects, attributes, and color details in the scene image, respectively; and the three-level vocabulary obtained in Step 1 is used. Using the CLIP model, we can intuitively manipulate real-world scene images through abstract text. The specific implementation is as follows: 3-1. Since it has been shown that different StyleGAN layers are responsible for generating different levels of detail in the scene image, the four latent space residual mappers are divided into four groups, respectively for layout, object, attribute and color, and different parts of the latent vector w are provided to each group. Based on the secondary vocabulary obtained in step 1 For each set of latent space residual mappers, selective training is performed, specifically for second-level vocabulary. The latent space residual mapper groups corresponding to words not included in the list do not need to be trained; 3-2. Represent the latent vector of the input image as follows: ,in These are respectively represented as different layers for The division, in which This corresponds to the vector part of the layout layer. This corresponds to the vector part of the object layer. This corresponds to the vector part of the attribute layer. This corresponds to the vector part of the color layer. Indicates will The remaining portion after the latent vector partitioning, since the StyleGAN network has 18 layers, is divided into groups of the first 14 layers; after passing through the latent space residual mapper, it can be obtained... ,in , , , These represent the groups of the mapping network; 3-3. After the latent space residual mapper is trained under the influence of CLIP loss, the resulting bias vector is... With the initial latent vector of the image Multiplication achieves the processing of latent vectors in an image. The CLIP loss minimizes the cosine distance between the generated image and the text prompt, while preserving other semantic content in the input image. ; Here, G represents the StyleGAN generator; to preserve some visual attributes of the original input image, the latent space after mapping is constrained using the L2 norm; therefore, the final loss function is as follows: ; 3-4. Edit the latent vectors The image is input into the StyleGAN network and the final output is the edited image.
2. The method for editing realistic scene images based on hierarchical classification text guidance according to claim 1, characterized in that, The specific method for step 1 is as follows: 1-1. Based on graph convolutional networks, using a text encoder and a label encoder, hierarchical structural relationship representations learned in the label set are shared. Extracting semantic information from the text respectively and tag semantics As shown in the formula below, where Represents a set of nodes in a hierarchical structure; Represents a collection of label nodes. The activation function is ReLU; ; ; 1-2. Semantic interpretation of text and tag semantics Projected into a joint embedding space, the joint embedding loss controls the text semantics. and tag semantics Similarity; 1-3 Using a matching learning loss, fine-grained label semantics, coarse-grained label semantics, and incorrect label semantics are obtained through training. Among these, the fine-grained label semantics are closest to the input level 3 vocabulary. That is, fine-grained tag semantics are Coarse-grained tag semantics are Other incorrect label semantics are far removed from the first-level vocabulary. ; 1-4. Using a trained hierarchical multi-label text classification model, input the first-level vocabulary. Get the required level 3 vocabulary and Level 2 vocabulary .
3. A method for editing realistic scene images based on hierarchical classification text guidance according to claim 1 or 2, characterized in that, Step 2 is explained in the following steps: 2-1. Using the e4e model trained on the LSUN dataset, obtain the inversion latent vector w of the real indoor scene in .pt format, which will be used as the input of StyleGAN; 2-2. Based on the semantic hierarchical characteristics of StyleGAN, the obtained latent vector w is divided into layers [0,2) of the generator network, layers [2,6) of the generator network for the object, layers [6,12) of the generator network for the attribute, and layers [12,14) of the generator network for the color.