Natural Language-Based Image Modification and Generation Method

Through the image modification and generation method based on natural language, using image generation models and text encoders, non-professional image creative output and refined control are realized, solving the problem of limited image generation capabilities and providing rapid image modification and generation solutions.

CN114140666BActive Publication Date: 2025-07-08SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111474605.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-06
Publication Date
2025-07-08
Estimated Expiration
2041-12-06

AI Technical Summary

Technical Problem

In the prior art, non-professionals cannot use computer-aided tools to efficiently produce images, and the generation ability of image generation algorithms is limited by the training range and cannot be finely adjusted.

Method used

Using natural language-based image modification and generation methods, the image generation model, image encoder, text encoder and contrast language image model are used to set the target generation strategy and layer update weights to achieve refined control and generation of images.

Benefits of technology

实现了通过自然语言进行可精细化控制的图像修改或生成,生成效果好且速度快,适用于非专业人士的图像创意产出。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140666B_ABST
    Figure CN114140666B_ABST
Patent Text Reader

Abstract

The present invention provides a method for image modification and generation based on natural language, including: calculating and obtaining an initial image latent vector based on the input image according to the task type; inputting target text information and calculating a target text embedding vector; inputting the target text and calculating a target text embedding vector based on the target text; setting different target generation strategies and calculating the layer update weights of the corresponding image generation pre-trained model based on the target generation strategies; training and optimizing the parameters of the image generation pre-trained model and the image latent vector according to the initial image latent vector, the target text embedding vector and the layer update weights, so as to obtain the latent vector of the updated synthesized image and the image generation pre-trained model; and obtaining and outputting the synthesized target image based on the latent vector of the updated synthesized image and the image generation pre-trained model. The present invention fills the gap in the task of image modification or generation that can be refined and controlled through natural language, and has good image modification and generation effects, and can obtain the output result in a short time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image modification and generation method, in particular to an image modification and generation method based on natural language. Background Art

[0002] With the development of computer hardware computing power and deep learning algorithms, computer-aided intelligent image design has gradually become a key tool in designers' work, including automatic coloring, automatic filling, etc. These algorithms give reference suggestions based on the designers' existing work, or supplement missing information, improving the work efficiency of design professionals.

[0003] However, for non-professionals, the lack of professional knowledge makes it difficult to produce creative works themselves, and they cannot use auxiliary design tools to produce image creativity.

[0004] Traditional algorithms for image modification and generation based on text information within a certain range mostly obtain the text-image generation ability by jointly training an image generation model and a language model, and their generation ability is limited to the text range provided during training. Due to the complexity of the image generation model, this range is usually relatively limited, and fine-tuning cannot be performed during the process of generating images. Summary of the Invention

[0005] In view of the above deficiencies in the prior art, the purpose of the present invention is to provide an image modification and generation method based on natural language, which fills the gap in the task of image modification or generation that can be finely controlled through natural language, has good image modification and generation effects, and can obtain output results in a relatively short time.

[0006] The present invention solves the above technical problems through the following technical solutions.

[0007] An image modification and generation method based on natural language includes the following steps:

[0008] Step S1, calculate and obtain an initial image latent vector based on the input image according to the task type;

[0009] Step S2, input the target text and calculate the target text embedding vector based on the target text;

[0010] Step S3, set different target generation strategies and calculate the layer update weights of the corresponding image generation pre-trained model based on the target generation strategies;

[0011] Step S4, calculate the initial image latent vector, the target text embedding vector, and the layer update weights based on the input image, and train and optimize the parameters of the image generation pre-trained model and the image latent vector to obtain the updated latent vector of the synthesized image and the image generation pre-trained model;

[0012] Step S5: Based on the latent vector of the updated synthetic image and the image generation pre-training model, obtain and output the synthesized target image.

[0013] Preferably, step S1 includes the following steps:

[0014] Step S1.1: Obtain the user input and determine whether there is an image in the input;

[0015] Step S1.2: If the judgment in step S1.1 is yes, the current task is to modify the image. Use the image encoder to calculate the latent vector corresponding to the input image, and use the calculated latent vector as the initial image latent vector;

[0016] Step S1.3: If the judgment in step S1.1 is no, the current task is to generate an image. Randomly sample a latent vector in the latent space of the input layer as the initial image latent vector;

[0017] Among them, the image encoder is an encoder that has the reverse calculation of the input latent vector of the corresponding image generator, such as the ReStyle encoder corresponding to StyleGAN.

[0018] Preferably, step S2 includes the following steps:

[0019] Step S2.1: Obtain the target text input by the user;

[0020] Step S2.2: Split the target text into a symbol set through a tokenizer;

[0021] Step S2.3: Use the pre-trained text encoder to calculate the target text embedding vector of the symbol set.

[0022] Among them, the tokenizer is a codebook that has the function of splitting natural language text into words and converting symbols. The pre-trained text encoder is a text model that has the function of embedding the text symbol set into a vector space. The tokenizer and the text encoder are usually used in pairs, such as the tokenizer and the text encoder of GPT2.

[0023] Preferably, step S3 includes the following steps:

[0024] Step S3.1: Set different target generation strategies. The target generation strategy includes the setting of degrees of freedom, and the setting of degrees of freedom includes: the setting of shape degree of freedom, texture degree of freedom, and content degree of freedom;

[0025] Step S3.2: Calculate the layer update weights corresponding to the image generation pre-training model according to the set target generation strategy. Among them, the layer update weights are used to determine the trainability of each layer of the image generation pre-training model

[0026] Among them, the degree of freedom is a hyperparameter that controls the effect of the generated image. The higher the degree of freedom, the wider the generation range, but the greater the probability of distortion; the lower the degree of freedom, the narrower the generation range, but the smaller the probability of distortion. The image generation pre-training model is a pre-trained image generator with layer decoupling ability, such as StyleGAN. The layer update weight determines the trainability of each layer of the generation model.

[0027] Preferably, the step S4 includes the following steps:

[0028] Step S4.1, input the initial image latent vector into the image generation pre-training model to obtain the output synthesized image;

[0029] Step S4.2, input the output synthesized image into the pre-trained visual embedding model to obtain the embedding vector of the synthesized image;

[0030] Step S4.3, input the embedding vector of the synthesized image and the embedding vector of the target text symbol set into the contrastive language-image pre-training model, and calculate the semantic distance as the contrastive loss value for model training;

[0031] Step S4.4, backpropagate the contrastive loss value to each node of the network, scale the loss value of each node according to the layer update weight, and then update the latent vector of the synthesized image and the parameters of the image generation pre-training model through the optimizer.

[0032] Among them, the contrastive language-image pre-training model is a model pre-trained based on text-image pairs, which has the ability to calculate the semantic distance between text and image. It has a wider image coverage range than the image generation pre-training model, and the corpus source is rich, such as the CLIP model.

[0033] Preferably, the step S5 includes the following steps:

[0034] Step S5.1, input the latent vector of the updated synthesized image into the updated image generation pre-training model to obtain the synthesized target image;

[0035] Step S5.2, output the synthesized target image to the display screen and display the result.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] 1. The present invention utilizes pre-training models including an image generation model, an image encoder model, a text encoder model, and a contrastive language-image model to decouple the modification and generation tasks of natural language and images, enabling the corpus to be freely amplified or replaced, and realizing image modification and generation based on natural language.

[0038] 2. The present invention introduces the target strategy setting for image synthesis, and controls features such as the shape, texture, and content of the synthesized image by defining the layer update weights, featuring refined control.

[0039] 3. The present invention does not require long - time model training, and only needs to optimize the image embedding vector and model parameters.

[0040] 4. The present invention unifies image modification and generation within the same framework, having better generality.

[0041] 5. The present invention fills the gap in the task of image modification or generation with refined control through natural language. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] By reading the following detailed description of non - restrictive embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:

[0043] Figure 1 is the algorithm framework diagram of the method for image modification and generation based on natural language according to the present invention;

[0044] Figure 2 is the schematic diagram of the result of image modification based on the present invention;

[0045] Figure 3 is the schematic diagram of the result of image generation based on the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all fall within the protection scope of the present invention.

[0047] Aiming at the problems of limited corpus, low controllability, and slow speed in the prior art, the purpose of the present invention is to provide a method for image modification and generation based on natural language. This method fills the gap in the task of image modification or generation with refined control through natural language, has good image modification and generation effects, and can obtain output results in a relatively short time.

[0048] As Figure 1 shown, this embodiment provides a method for image modification and generation based on natural language, including:

[0049] Step S1, based on the task type, calculate and obtain the initial image latent vector based on the input image;

[0050] Step S2, input the target text, and calculate the target text embedding vector based on the target text;

[0051] Step S3, set different target generation strategies, and calculate the layer update weights of the corresponding image generation pre-training model based on the target generation strategies;

[0052] Step S4, calculate the initial image latent vector image, the target text embedding vector text, and the layer update weight generation strategy strategy based on the input image, and train and optimize the parameters of the image generation pre-training model and the image latent vector to obtain the updated latent vector of the synthesized image and the image generation pre-training model;

[0053] Step S5, based on the updated latent vector of the synthesized image and the image generation pre-training model, obtain and output the synthesized target image.

[0054] The said Step S1 includes the following steps:

[0055] Step S1.1, obtain the user input, and determine whether there is an image in the input;

[0056] Step S1.2, if the judgment in Step S1.1 is yes, then the current task is to modify the image, use the image encoder to calculate the latent vector corresponding to the input image, and use the calculated latent vector as the initial image latent vector;

[0057] Step S1.3, if the judgment in Step S1.1 is no, then the current task is to generate an image, and randomly sample a latent vector in the latent space of the input layer as the initial image latent vector.

[0058] Among them, the image encoder is an encoder with the ability to inversely calculate the input latent vector of the corresponding image generator, such as the ReStyle encoder corresponding to StyleGAN.

[0059] The said Step S2 includes the following steps:

[0060] Step S2.1, obtain the target text input by the user;

[0061] Step S2.2, split the target text into a symbol set through a tokenizer;

[0062] Step S2.3, use the pre-trained text encoder to calculate the target text embedding vector of the symbol set.

[0063] Among them, the tokenizer is a codebook with the ability to split natural language text into words and convert symbols. The pre-trained text encoder is a text model with the ability to perform vector space embedding on the text symbol set. The tokenizer and the text encoder are usually used in pairs, such as the tokenizer and the text encoder of GPT2.

[0064] The step S3 includes the following steps:

[0065] Step S3.1, setting different target generation strategies, where the target generation strategy includes setting degrees of freedom, and the setting of degrees of freedom includes: setting shape degrees of freedom, texture degrees of freedom, and content degrees of freedom;

[0066] Step S3.2, calculating layer update weights corresponding to the image generation pre-trained model according to the set target generation strategy, where the layer update weights are used to determine the trainability of each layer of the image generation pre-trained model.

[0067] Among them, the degree of freedom is a hyperparameter for controlling the generated image effect. The higher the degree of freedom, the wider the generation range, but the greater the distortion probability; the lower the degree of freedom, the narrower the generation range, but the smaller the distortion probability. The image generation pre-trained model is a pre-trained image generator with layer decoupling ability, such as StyleGAN. The layer update weights determine the trainability of each layer of the generation model.

[0068] The step S4 includes the following steps:

[0069] Step S4.1, inputting the initial image latent vector into the image generation pre-trained model to obtain the output synthesized image;

[0070] Step S4.2, inputting the output synthesized image into the pre-trained visual embedding model to obtain the embedding vector of the synthesized image;

[0071] Step S4.3, inputting the embedding vector of the synthesized image and the embedding vector of the target text symbol set into the contrastive language-image pre-trained model, and calculating the semantic distance as the contrastive loss value for model training;

[0072] Step S4.4, backpropagating the contrastive loss value to each node of the network, scaling the loss value of each node according to the layer update weights, and then updating the latent vector of the synthesized image and the parameters of the image generation pre-trained model through the optimizer.

[0073] Among them, the contrastive language-image pre-trained model is a model pre-trained according to text-image pairs, with the ability to calculate the semantic distance between text and image, having a wider image coverage range than the image generation pre-trained model, and rich corpus sources, such as the CLIP model.

[0074] The step S5 includes the following steps:

[0075] Step S5.1, inputting the updated latent vector of the synthesized image into the updated image generation pre-trained model to obtain the synthesized target image;

[0076] Step S5.2, outputting the synthesized target image to the display screen and presenting the result.

[0077] This embodiment fills the gap in image modification or generation tasks that can be finely controlled through natural language, with good image modification and generation effects, and the output results can be obtained in a relatively short time.

[0078] Figure 2 It is a schematic diagram of the result of image modification based on the present invention; Figure 3 It is a schematic diagram of the result of image generation based on the present invention.

[0079] The above specific embodiments further elaborate on the technical problems to be solved, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for image modification and generation based on natural language, characterized in that, It includes the following steps: Step S1 includes: Step S1.1, obtaining user input and determining whether there is an image in the input; Step S1.2, if the determination in Step S1.1 is yes, the current task is to modify the image, calculating the latent vector corresponding to the input image using an image encoder, and using the calculated latent vector as the initial image latent vector; Step S1.3, if the determination in Step S1.1 is no, the current task is to generate an image, randomly sampling a latent vector in the latent space of the input layer as the initial image latent vector; Step S2, inputting the target text and calculating the target text embedding vector based on the target text; Step S3, setting different target generation strategies and calculating the layer update weights of the corresponding image generation pre-trained model based on the target generation strategy; wherein, Step S3 includes the following steps: Step S3.1, setting different target generation strategies, the target generation strategy including the setting of degrees of freedom, and the setting of degrees of freedom including: the setting of shape degrees of freedom, texture degrees of freedom, and content degrees of freedom; Step S3.2, calculating the layer update weights of the corresponding image generation pre-trained model according to the set target generation strategy, wherein the layer update weights are used to determine the trainability of each layer of the image generation pre-trained model; Step S4, training and optimizing the parameters of the image generation pre-trained model and the image latent vector according to the initial image latent vector, the target text embedding vector, and the layer update weight generation strategy calculated from the input image to obtain the updated latent vector of the synthesized image and the image generation pre-trained model; wherein, Step S4 includes the following steps: Step S4.1, inputting the initial image latent vector into the image generation pre-trained model to obtain the output synthesized image; Step S4.2, inputting the output synthesized image into the pre-trained visual embedding model to obtain the embedding vector of the synthesized image; Step S4.3, inputting the embedding vector of the synthesized image and the target text embedding vector into the contrastive language-image pre-trained model to calculate the semantic distance as the contrastive loss value for model training; Step S4.4, backpropagating the contrastive loss value to each node of the network, scaling the loss value of each node according to the layer update weights, and then updating the latent vector of the synthesized image and the parameters of the image generation pre-trained model through an optimizer; Step S5, obtaining and outputting the synthesized target image based on the updated latent vector of the synthesized image and the image generation pre-trained model.

2. The method for image modification and generation based on natural language according to claim 1, wherein The image encoder is an encoder that reversely calculates the input latent vector of the corresponding image generator.

3. The method for image modification and generation based on natural language according to claim 1, wherein Step S2 includes the following steps: Step S2.1, obtaining the target text input by the user; Step S2.2, splitting the target text into a symbol set through a tokenizer; Step S2.3, calculating the target text embedding vector of the symbol set using a pre-trained text encoder.

4. The method for image modification and generation based on natural language according to claim 3, wherein, The tokenizer is a codebook that splits natural language text into words and converts symbols; the pre-trained text encoder is a text model that embeds the text symbol set into a vector space; the tokenizer and the text encoder are used in pairs.

5. The method for image modification and generation based on natural language according to claim 1, characterized in that, The degree of freedom is a hyperparameter that controls the generated image effect. The higher the degree of freedom, the wider the generation range, but the greater the probability of distortion; the lower the degree of freedom, the narrower the generation range, but the smaller the probability of distortion.

6. The method for image modification and generation based on natural language according to claim 1, wherein The image generation pre-training model is a pre-trained image generator with layer decoupling ability.

7. The method for image modification and generation based on natural language according to claim 1, wherein The contrastive language-image pre-training model is a model pre-trained based on text images and has the ability to calculate the semantic distance between text and images.

8. The method for image modification and generation based on natural language according to claim 1, characterized in that Step S5 includes the following steps: Step S5.1, input the latent vector of the updated synthetic image into the updated image generation pre-training model to obtain the synthesized target image; Step S5.2, output the synthesized target image to the display screen and display the result.

Citation Information

Patent Citations

  • Network training method and device and image generation method and device

    CN111223040A

  • Increasing inclusiveness of search result generation through tuned mapping of text and images into the same high-dimensional space

    US20190266262A1