A multi-purpose image redrawing method, device, and medium based on prompt learning

Through a multi-purpose image redrawing method based on prompt word learning, using a text encoder and a diffusion model, combined with a data set of expansion ratio training model, the problem of limitations in the training strategy of the image local redrawing model in the prior art is solved, and the harmony between the multi-purpose image redrawing and generated content and the overall content of the image is achieved.

CN117710524BActive Publication Date: 2025-06-24SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311566976.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-06-24
Estimated Expiration
2043-11-23

AI Technical Summary

Technical Problem

The existing image local redrawing model has limitations in training strategies, and it is difficult to eliminate and add objects in the mask area at the same time. It is easy to generate random objects when there are many surrounding objects, and the text semantics do not match the image content, so the generated content is in harmony with the overall image.

Method used

A multi-purpose image redrawing method based on prompt word learning is adopted. By obtaining task prompt words, a text encoder is used to generate conditional encoding and non-conditional encoding, and inputting them to the diffusion model for redrawing, constructing task prompt words for overall content redrawing and specific object redrawing, and combining the data set of expansion ratio training model to ensure that the generated content is harmonious with the overall content of the image.

Benefits of technology

Multi-purpose image redrawing is realized, and the overall content redrawing and specific object redrawing can be performed according to task prompt words, which improves the success rate of object erasing, ensures that the generated content is harmonious with the overall content of the image, and solves the problem of overall dissonance between the generated content and the image in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117710524B_ABST
    Figure CN117710524B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-purpose image redrawing method, device, and medium based on prompt learning. The method includes the following steps: obtaining at least one task prompt, an input image, and mask data; processing the text including the task prompt using a text encoder to obtain conditional encoding and unconditional encoding; encoding the input image using an image encoder; using the mask data, conditional encoding, unconditional encoding, and the encoded input image as inputs to a diffusion model for redrawing, and decoding the output of the diffusion model to obtain the redrawn output image. Compared with the prior art, the present invention performs text encoding based on task prompts to guide the diffusion model to redraw input pictures for various tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a multi-purpose image redrawing method, device, and medium based on prompt learning. Background Art

[0002] An image local redrawing model obtains an output picture redrawn in a mask area based on an input image of an object, an input mask, and other optional control conditions, such as text. As Figure 3 shown. If text is used as a control condition, the output picture needs to conform to the content described by the text in the mask area.

[0003] The existing local redrawing solutions have the following disadvantages:

[0004] (1) Due to the limitations of the training strategy, the existing local redrawing models usually cannot well complete the elimination of objects in the mask area and the addition of objects with one model.

[0005] (2) When the existing local redrawing model takes the elimination of objects in the mask as the training goal, the training strategy usually uses the surrounding information of the image as a condition to restore the randomly covered image area. When there are many surrounding objects, the model will generate random objects in the covered area, which is contrary to the goal of eliminating objects. As Figure 2 (a) shown.

[0006] (3) Some training strategies of the existing local redrawing models restore randomly sized mask areas based on the global image description text, which may lead to a mismatch between the text semantics and the image content in the mask area. As Figure 2 (b) shown.

[0007] (4) Some of the existing local redrawing models restore the covered image area containing objects based on the local object description text, without precisely considering the size and proportional relationship between the mask area and the actual generated content area, which may lead to disharmony between the generated content and the overall image content. As Figure 2 (c) shown.

[0008] In summary, there is currently a lack of an image local redrawing method to solve or partially solve the foregoing problems. Summary of the Invention

[0009] The purpose of the present invention is to overcome the defects of the existing technologies described above, and to provide a multi-purpose image redrawing method, device, and medium based on prompt learning to optimize the effect of image redrawing.

[0010] The purpose of the present invention can be achieved by the following technical solutions:

[0011] One aspect of the present invention provides a multi-purpose image redrawing method based on prompt learning, comprising the following steps:

[0012] Obtain at least one task prompt, an input image, and mask data;

[0013] Process the text including the task prompt using a text encoder to obtain conditional encoding and unconditional encoding;

[0014] Encode the input image using an image encoder;

[0015] Use the mask data, conditional encoding, unconditional encoding, and the encoded input image as inputs to a diffusion model for redrawing, and decode the output of the diffusion model to obtain the redrawn output image.

[0016] As a preferred technical solution, for the task of overall content redrawing construction, the task prompt includes an overall content redrawing construction task prompt. Using the content redrawing construction task prompt and empty text as inputs to the text encoder, conditional encoding and unconditional encoding are obtained respectively.

[0017] As a preferred technical solution, for the task of specific object redrawing construction, the task prompt includes a specific object redrawing construction task prompt. Using the specific object redrawing construction task prompt superimposed with the obtained input text as the input to the text encoder to obtain conditional encoding, and using empty text as the input to the text encoder to obtain unconditional encoding.

[0018] As a preferred technical solution, for the task of erasing objects in the mask area, the task prompt includes an overall content redrawing construction task prompt and a specific object redrawing construction task prompt. Using the overall content redrawing construction task prompt and the specific object redrawing construction task prompt as inputs to the text encoder respectively to obtain conditional encoding and unconditional encoding.

[0019] As a preferred technical solution, the training process of the diffusion model includes the following steps:

[0020] Generate multiple dilated masks by changing the ratio of the size of the original mask area to the size of the dilated mask area, and train the diffusion model.

[0021] As a preferred technical solution, for the task of redrawing an object proportionally, the task prompt words include an overall content redrawing construction task prompt word and a proportional redrawing prompt word. The obtained input text is respectively superimposed with the overall content redrawing construction task prompt word and the proportional redrawing prompt word, and is respectively used as the input of the text encoder. The outputs are superimposed proportionally to obtain a conditional encoding, and an empty text is used as the input of the text encoder to obtain an unconditional encoding.

[0022] As a preferred technical solution, the task prompt words are trainable vectors.

[0023] As a preferred technical solution, the conditional encoding is used to encourage the diffusion model to generate relevant concepts, and the unconditional encoding is used to suppress the diffusion model from generating relevant concepts.

[0024] Another aspect of the present invention provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the foregoing multi-purpose image redrawing method based on prompt word learning.

[0025] Another aspect of the present invention provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device. The one or more programs include instructions for executing the foregoing multi-purpose image redrawing method based on prompt word learning.

[0026] Compared with the prior art, the present invention has the following beneficial effects:

[0027] Realize multi-purpose image redrawing: By obtaining task prompt words during the redrawing process, using the text encoder to generate text encodings including different task prompt words, and using them as conditional encoding and unconditional encoding of the diffusion model, so as to achieve other effects and perform text encoding based on task prompts to guide the diffusion model to redraw the input picture for multiple tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a schematic diagram of the overall framework of the redrawing operation in the embodiment;

[0029] Figure 2 It is a schematic diagram when redrawing using the existing image local redrawing technology;

[0030] Figure 3 It is a schematic diagram of the image local redrawing process;

[0031] Figure 4 It is a schematic diagram of the text encoding process under the overall content redrawing construction task in the embodiment;

[0032] Figure 5Schematic diagram of the text encoding process for the specific object redrawing construction task in the embodiment;

[0033] Figure 6 Effect comparison diagram of the overall content redrawing and specific object redrawing in the embodiment;

[0034] Figure 7 Schematic diagram of the object erasing task in the masked area in the embodiment;

[0035] Figure 8 Schematic diagram of the dilation operation on the mask in the embodiment;

[0036] Figure 9 Schematic diagram of the text encoding process for the object redrawing task according to a ratio in the embodiment;

[0037] Figure 10 Output schematic diagram of obtaining different ratios of object content in the mask by adjusting the α value in the embodiment. Detailed implementation manners

[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] Embodiment 1

[0040] In view of the problems existing in the foregoing prior art, this embodiment provides a multi-purpose image redrawing method based on prompt learning. In this embodiment, the input text is encoded using a CLIP text encoder, and at the same time, the input image is encoded using a VAE image encoder. Together with the mask, a diffusion model is used to obtain the encoded output image, and then the output image is decoded through a VAE image decoder. Among them, the mask includes information such as size, shape, and position.

[0041] (1) In this method, special task prompts are constructed for different tasks. Task prompts are constructed for the overall content redrawing and specific object redrawing respectively, so that the image local redrawing model has the capabilities of both overall content redrawing and specific object redrawing.

[0042] In this solution, task prompts are constructed for the overall content redrawing and specific object redrawing respectively. As Figure 4 shown, when training the diffusion model, we construct the task prompt "#Inpainting" for the overall content local redrawing. As Figure 5As described above, a task prompt "#Insert" is constructed for the local redrawing of a specific object. After passing through the text encoder, the vector obtained from the prompt can be used to control the diffusion model to generate images within the masked area. Both the task prompt and the model are trainable.

[0043] Through such training, when the model is only input with the task prompt "#Inpainting", the model does not need to consider text guidance at this time and only needs to consider the overall content of the image for redrawing, making the content of the redrawn area harmonious with other areas of the image; while when the model is input with text containing "#Insert", the model focuses on generating the object that conforms to the prompt within the masked area. Figure 6 The output results are for inserting the overall content redrawing prompt and the specific object redrawing prompt respectively. It can be found that through such a design, the content harmony of the overall content redrawing and the text effectiveness of the specific object redrawing have been significantly improved.

[0044] (2) This method uses the overall content redrawing prompt as the positive prompt and the specific object redrawing prompt as the negative prompt, making the local redrawing model have a higher object erasure success rate compared to the existing solutions.

[0045] In the previous step, after enabling the model to have the capabilities of both overall content redrawing and specific object redrawing by adding two prompts, in this step, it is further combined with the unconditional guidance sampling strategy of the diffusion model for use. Generally speaking, in the unconditional guidance sampling strategy of the diffusion model, the positive prompt encourages the model to generate relevant concepts, and the negative prompt inhibits the model from generating relevant concepts.

[0046] Here, we use the encoded vector of the overall content redrawing as the conditional encoding input of the diffusion model, and at the same time use the encoded vector of the specific object redrawing as the unconditional encoding input of the diffusion model, as Figure 7 shown. Since the encoded vector of the specific object redrawing always guides the model to generate the object, by using such an encoded vector as the negative guidance, we can guide the diffusion model to inhibit the generation of the object, so that the model can focus on erasing the object within the masked area.

[0047] (3) This method constructs a dataset where the size of the dilated mask is proportional to the actual size of the object within the mask, for improving the training strategy of the local redrawing model.

[0048] In this step, in order to enable the local redrawing model to generate objects within the mask area and allow the user to control the degree of fitting of the generated objects to the mask shape, a training dataset of "object-mask-fitting degree" is first constructed. Specifically, first select a segmentation dataset where the objects in each image of this dataset have segmentation masks that exactly fit the objects. Dilate such an exact object segmentation mask through image dilation processing. The dilation operation uses the convolution kernel-based dilation function in OpenCV. By specifying the key parameters in the function, namely the convolution kernel sizes of 3, 5, and 7, the segmentation mask can be dilated to different degrees. As Figure 8 shown, let the size of the white area (i.e., the original mask area) of the dilated mask be S(white), and the size of the gray area be S(gray). We can calculate the area ratio of the two areas as follows:

[0049] S(white):S(gray)=α:(1 - α)

[0050] where α represents the degree of fitting of the mask to the object. The larger α is, the more the object and the mask shapes fit; the smaller α is, the looser the degree of fitting between the object and the mask shapes. When α = 1, it means that the shapes of the object and the mask are completely fitted. Thus, we obtain the "object-mask-fitting degree" dataset, and such a dataset is used for the next step of training the local redrawing model to enable it to have the ability to adjust the degree of fitting between the generated object and the mask.

[0051] (4) This method establishes the connection between the dilation ratio of the mask and the encoding ratio of the text vector with two task prompt words, realizing further shape guidance for the generated object.

[0052] In the training of this step, our goal is to use the "object-mask-fitting degree" dataset constructed in the previous step to train the local redrawing model so that the model has the ability to adjust the degree of fitting between the generated object and the mask shape. The specific approach is as follows.

[0053] First, introduce a new task prompt word "#Shape_Insert", and at the same time use the proportional dataset constructed in (3) to train the local redrawing model. During training, our input is divided into four parts: the image χ0 containing the object content, the text prompt word y, the dilation ratio a of the mask, and the mask m dilated proportionally in the object area. Use the following loss function formula L to train the diffusion model:

[0054] L = ||∈ - ∈ θ (χ t , x0, y, α, m, t)|| 2

[0055] where ∈ is the Gaussian noise added by the diffusion model to the input x0 at time step t, and x t is the image obtained by adding noise to the input x0 at time step t. θ is the unet model in the diffusion model, and y contains the description of the generation target and the task prompt. Specifically, as Figure 9 shown, the vectors obtained by passing the redrawing prompt "#Shape_Insert" and the input text through the text encoder and the vectors obtained by passing the task prompt "#Inpainting" and the input text through the text encoder are weighted and summed according to α and 1 - α respectively to obtain the interpolated vector, which is used as the conditional encoding y of the diffusion model for input. Through such training, a corresponding relationship is established between the mask fitting degree and the interpolation ratio of the two task prompts "#Shape_Insert" and "#Inpainting". After training, when the user needs to perform local redrawing of the object within the mask area according to the size ratio, the user inputs the α value, and the two encoded vectors based on the prompts "#Shape_Insert" and "#Inpainting" are weighted and summed to obtain the interpolated encoded vector, which can guide the local redrawing model to generate the object with the corresponding fitting degree according to α. For example, when the proportion of the object we need to redraw in the mask area is 0.8, we can set the α value to 0.8, Figure 10 as shown is the output with different proportions of the object content in the mask area.

[0056] In summary, the present invention:

[0057] (1) Task prompts are constructed in the local redrawing model, enabling the control of the content in the mask area by the task prompts. The model can either redraw according to the overall picture content or only redraw the object in the text prompt in the mask area. By using the task prompts, the overall content redrawing and the specific object redrawing are combined. When performing the task of local content redrawing using the model of the present invention, the same model can be used for both overall content redrawing and specific object redrawing, and only specific prompts need to be inserted into the input text.

[0058] (2) The text encodings of different task prompts can be combined and input into the conditional encoding and unconditional encoding of the diffusion model to achieve other effects. For example, using the encodings of the overall content redrawing prompt and the specific object redrawing prompt as the conditional encoding and unconditional encoding can make the local redrawing model focus on erasing the object and improve the situation where redundant objects are generated in the mask area affected by the original image during the erasing process.

[0059] (3) A method for constructing a mask using the dilation ratio for training the local redrawing model is provided, improving the training strategy of the local redrawing model to make the generated content in the mask area match the input text.

[0060] (4) Train the model by combining the mask data with different inflation ratios and the text vector encoding ratios of the two task prompts, so that the local redrawing model has the ability to generate specific objects proportionally in the masked area. Precisely consider the size ratio of the masked area and the generated object, so that the size of the generated content in the masked area can be controlled by the ratio, thereby making the generated content more harmonious with the overall content of the image.

[0061] It should be noted that the task prompts in this application can be any English or Chinese words, not limited to words such as "#Inpainting" and "#Insert", and can be created as needed. In addition, the method of using task prompts to adjust the local redrawing model task in this embodiment can be applied to tasks other than tasks such as overall content redrawing and specific object content redrawing, such as specific color redrawing tasks.

[0062] Embodiment 2

[0063] This embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the multi-purpose image redrawing method based on prompt word learning as described in Embodiment 1.

[0064] Embodiment 3

[0065] This embodiment provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device. The one or more programs include instructions for executing the multi-purpose image redrawing method based on prompt word learning as described in Embodiment 1.

[0066] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A multi-purpose image redrawing method based on prompt learning, characterized in that, It includes the following steps: Obtain at least one task prompt, an input image, and mask data; Process the text including the task prompt using a text encoder to obtain conditional encoding and unconditional encoding; Encode the input image using an image encoder; Use the mask data, conditional encoding, unconditional encoding, and the encoded input image as inputs to a diffusion model for redrawing, and decode the output of the diffusion model to obtain the redrawn output image. The training process of the diffusion model includes the following steps: Generate multiple dilated masks by changing the ratio of the size of the original mask region to the size of the dilated mask region, and train the diffusion model. For the task of redrawing an object proportionally, the task prompt includes a global content redrawing construction task prompt and a proportional redrawing prompt. Superimpose the obtained input text with the global content redrawing construction task prompt and the proportional redrawing prompt respectively, and use them as inputs to the text encoder. Add the outputs proportionally to obtain conditional encoding, and use an empty text as the input to the text encoder to obtain unconditional encoding.

2. The multi-purpose image redrawing method based on prompt learning according to claim 1, wherein, For the global content redrawing construction task, the task prompt includes a global content redrawing construction task prompt. Use the content redrawing construction task prompt and an empty text as inputs to the text encoder to obtain conditional encoding and unconditional encoding respectively.

3. A multi-purpose image redrawing method based on prompt learning according to claim 1, characterized in that, For the specific object redrawing construction task, the task prompt includes a specific object redrawing construction task prompt. Superimpose the specific object redrawing construction task prompt with the obtained input text and use it as the input to the text encoder to obtain conditional encoding, and use an empty text as the input to the text encoder to obtain unconditional encoding.

4. A multi-purpose image redrawing method based on prompt learning according to claim 1, characterized in that, For the task of erasing an object in a mask region, the task prompt includes a global content redrawing construction task prompt and a specific object redrawing construction task prompt. Use the global content redrawing construction task prompt and the specific object redrawing construction task prompt as inputs to the text encoder respectively to obtain conditional encoding and unconditional encoding.

5. A multi-purpose image redrawing method based on prompt learning according to claim 1, characterized in that, The task prompt is a trainable vector.

6. A multi-purpose image redrawing method based on prompt learning according to any one of claims 1-4, characterized in that The conditional encoding is used to encourage the diffusion model to generate relevant concepts, and the unconditional encoding is used to inhibit the diffusion model from generating relevant concepts.

7. An electronic device, characterized in that, It includes: One or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the multi-purpose image redrawing method based on prompt learning as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, It includes one or more programs for execution by one or more processors of an electronic device. The one or more programs include instructions for executing the multi-purpose image redrawing method based on prompt learning as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Text intelligent generation 3D virtual teaching resource system and working method thereof

    CN115757850A

  • Image processing method and device, computer and storage medium

    CN116310046A