A text-driven subject personalized completion method based on an image inpainting model

By introducing a separate repair framework and attribute decoupling mechanism in text-driven subject image completion, the trade-off problem between data and accuracy and noise interference problems in the existing methods are solved, and high-quality and consistent image completion effect is achieved.

CN119722872BActive Publication Date: 2025-05-27DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510227784.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-27
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing methods have trade-offs between data and precision in text-driven subject image completion, and there is a problem of global noise interference and insufficient local information integration during object editing.

Method used

A generation method based on diffusion model is proposed. By introducing efficient parameter image customization technology and separation repair framework, the image completion process is divided into two stages: local content generation and global context coordination, and the model overfitting problem is solved through attribute decoupling mechanism and text attribute replacement module.

Benefits of technology

It realizes accurate object insertion and attribute-level modification under text guidance, improves the quality and consistency of image completion, and overcomes the problems of noise interference and insufficient information integration in existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722872B_ABST
    Figure CN119722872B_ABST
Patent Text Reader

Abstract

A text-driven subject personalized completion method based on an image inpainting model, belonging to the field of image generation and editing. Step one is to perform personalized fine-tuning on the image inpainting model; step two is two-stage image completion. The present invention further optimizes the personalized fine-tuning method based on DreamBooth to solve the problem of subject feature attribute coupling, proposes vector decomposition to further solve the coupling problem, and proposes a two-stage image personalized completion framework to improve the quality of image inpainting, and finally realizes a high-quality text-driven subject personalized completion method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image generation and editing. Specifically, it particularly relates to a text-driven subject personalized completion method based on an image repair model. Background Art

[0002] With the continuous development of diffusion models, image generation technologies based on diffusion models have made remarkable progress and further promoted the evolution of the image editing field. Among them, subject-driven image personalization generation has attracted much attention because it can generate realistic images of specified subjects in various scenarios. Existing methods such as DreamBooth and IP-Adapter mainly use different adaptation strategies to enable the pre-trained text-to-image diffusion model to achieve various editing effects while maintaining the object identity. However, most of these methods focus on the reconstruction of the entire scene, and it faces great challenges when users hope to insert a specific object in a specific area of a given scene. To meet practical needs, such as rendering of effect drawings, poster design, and virtual try-on, subject-driven image completion has gradually become an important task.

[0003] For this task, existing methods such as AnyDoor and MimicBrush extract example information from reference images to conditionally constrain the pre-trained diffusion model. However, these reference visual features often contain overly specific content, restricting the generation ability of the pre-trained text-to-image model and making it difficult to flexibly edit the desired effects or attributes of the target object. To overcome this limitation, some recent works such as LAR-Gen have tried to improve the effect by combining image and text-guided attention mechanisms. However, these methods still face challenges in the trade-off between training data efficiency and visual editing quality. Summary of the Invention

[0004] The present invention proposes a generation method based on a diffusion model for realizing text-driven subject image completion, effectively alleviating the trade-off problem between data and accuracy. In this method, by introducing parameter-efficient image customization techniques (such as DreamBooth) for model adaptation, the goal can be achieved with a small number of template images. However, there are still significant limitations in directly applying such methods to achieve precise object insertion and modification. For example, the denoising process in the repair model may introduce global noise, interfering with the integration of local information and affecting the synthesis quality of the target area. In addition, the small-sample fine-tuning strategy couples the appearance of the subject to a single identity marker ([sks]), which is prone to appearance overfitting, thus weakening the editing effect of text instructions.

[0005] To address the above problems, the present invention proposes an innovative generation framework to achieve high precision and efficiency at the attribute level. The present invention introduces a separate repair framework, which divides the repair process into two stages: local content generation and global context coordination, enhancing the precise integration of local objects and the coordination and consistency of the overall vision respectively. At the same time, this method introduces an attribute decoupling mechanism, which solves the problem of model overfitting through diverse text descriptions and image pairs of subject attributes. In addition, the text attribute replacement module introduced by the present invention in the inference stage uses an orthogonal decomposition strategy to separate interference information from text guidance, further improving the editing quality of the target object.

[0006] The technical solution adopted by the present invention:

[0007] A text-driven subject personalization completion method based on an image inpainting model, the steps are as follows:

[0008] Step 1, perform personalized fine-tuning on the image inpainting model;

[0009] (1.1) Obtain a set of images with a common target subject and the corresponding text descriptions of the images. Each image and its corresponding text description form a subject data pair, which constitutes the training data for fine-tuning the image inpainting model, so that the image inpainting model can learn the features of the target subject. The text description corresponding to the image refers to a concise description of the target subject containing identity markers.

[0010] (1.2) Based on the target problem proposed by the user, use a multimodal large language model to describe the attributes of the target subject in the set of images with a common target subject in step (1.1), and obtain key-value pairs including all feature descriptions. Based on the key-value pairs, use the multimodal large language model again, and based on the further target problem proposed by the user, obtain more regular text about the description of the target subject. The regular text is a detailed description that does not contain identity markers and includes some target subject features in the key-value pairs.

[0011] (1.3) With the help of a text-controlled image generation model, generate a specified number of regular images based on the regular text obtained in step (1.2), and obtain regular data pairs including the regular text and the regular images.

[0012] (1.4) Use an image segmentation model to obtain the masks of all the images in the subject data pairs and the regular data pairs, and add the masks to the corresponding subject data pairs and regular data pairs for fine-tuning the image inpainting model. All the images include the regular images obtained in step (1.3) and the images of the target subject in step (1.1).

[0013] (1.5) Use the masked body data pairs and regular data pairs obtained in step (1.4) to fine-tune the image inpainting model. Different sampling probabilities should be adopted for different body data pairs and regular data pairs, and fine-tune the image inpainting model based on the LoRA method to learn the features of the target subject.

[0014] Preferably, the image inpainting model adopts a diffusion inpainting model:

[0015] Randomly add noise to all images to time step t, and send the noisy images into the UNet of the diffusion inpainting model. After passing through the initial convolutional layer, at the same time, dilate the mask obtained in step (1.4) to a random size and encode it together with the corresponding image. Add the encoded whole to the result after passing through the initial convolutional layer, and perform subsequent downsampling and upsampling processing in the UNet to predict the added noise. During the training process of the diffusion inpainting model, it is necessary to fine-tune the key linear transformation parameter matrix and the value linear transformation parameter matrix of the attention layer of the diffusion inpainting model based on the LoRA method; it is necessary to calculate the reconstruction loss and weight the loss using the dilated mask to update the parameters of LoRA.

[0016] Step two: Two-stage image completion;

[0017] (2.1) During the process of image completion, the input of the fine-tuned image inpainting model includes the background image to be completed, the binary mask image corresponding to the background image, the completion area specified by the user on the binary mask image, and the personalized text description containing identity markers for the target subject set by the user. The completion area is used to complete the target subject. Crop out a specified-size area at the same position in the background image and the binary mask image that contains and is larger than the completion area to obtain a local image and a local binary mask; send the personalized text description containing identity markers for the target subject set by the user, the background image to be completed, the binary mask image corresponding to the background image, the local image, and the local binary mask into the fine-tuned image inpainting model.

[0018] (2.2) Encode the personalized text description containing identity markers for the target subject set by the user using a text encoder to obtain , encode the key-value pairs obtained in step (1.2) using a text encoder to obtain , and then apply vector decomposition to decouple the features to obtain the decoupled feature vectors. The decoupling formula is as follows:

[0019]

[0020] Among them, represents the modulus of the vector;

[0021] (2.3) Encode the local image to obtain the latent variable of the local image. Initialize random noise of the same dimension with a standard normal distribution, and perform forward noise addition of the diffusion repair model on the latent variable of the local image to time step T, where T is the total number of noise addition steps of the diffusion repair model, to obtain the latent variable of the noisy local image. Then, simultaneously input the latent variable of the noisy local image and the decoupled feature vector obtained in step (2.2) into UNet for denoising to obtain the initial predicted noise;

[0022] (2.4) The diffusion repair model uses the initial predicted noise obtained in step (2.3) to denoise the latent variable of the noisy local image before denoising in step (2.3) to obtain the latent variable of the denoised local image. Next, use the random noise initialized with the standard normal distribution in step (2.3) to perform forward noise addition on the latent variable of the local image obtained in step (2.3) to the current time step to obtain the latent variable of the noisy local image at the current time step. Fuse the latent variable of the denoised local image and the latent variable of the noisy local image at the current time step using a local binary mask to obtain the latent variable of the noisy local image that needs to be input into UNet, and simultaneously input it into UNet with the decoupled feature vector obtained in step (2.2) for denoising to obtain the predicted noise at the current time step, which is used as the denoising input for the diffusion repair model at the next time step. Repeat this step until , the predicted noise at the last time step is obtained. Where is a hyperparameter set manually, and T is the total number of noise addition steps of the diffusion repair model. The fusion formula is as follows:

[0023]

[0024] Where represents the latent variable of the denoised local image, is the local binary mask, represents the latent variable of the noisy local image at time step t; represents the element-wise multiplication operation of matrices.

[0025] (2.5) In the diffusion repair model, use the predicted noise at the last time step to thoroughly denoise the latent variable of the denoised local image after fusion at the last time step, directly predict the clean latent variable, and then use the decoder to decode the clean latent variable to obtain the local clean image. Next, according to the cropping area in step (2.1), fuse the local clean image with the input background image to obtain the global clean image, and then use the encoder to encode to obtain the latent variable corresponding to the global clean image.

[0026] (2.6) Encode the globally clean image obtained in step (2.5) together with the binary mask image corresponding to the background image input in step (2.1), and send it into UNet. At the same time, in the diffusion repair model, initialize random noise of the same dimension using the standard normal distribution in step (2.3), and forward add noise to the latent variable corresponding to the globally clean image obtained in step (2.5) to , obtaining the latent variable after adding noise to the clean image, and then also send it into UNet to obtain the globally predicted noise.

[0027] (2.7) Use the globally predicted noise obtained in step (2.6) to denoise the latent variable after adding noise to the clean image in step (2.6), obtaining the latent variable of the denoised global image;

[0028] (2.8) The diffusion repair model initializes random noise of the same dimension using the standard normal distribution in step (2.3), and forward adds noise to the latent variable corresponding to the globally clean image obtained in step (2.5) to the current time step, obtaining the latent variable of the noisy global image at the current time step; fuse the latent variable of the denoised global image obtained in step (2.7) and the latent variable of the noisy global image at the current time step using the binary mask image corresponding to the background image, obtaining the latent variable of the noisy global image to be input into UNet, and send it into UNet together with the decoupled feature vector obtained in step (2.2) for denoising, obtaining the predicted noise at the current time step, and use the predicted noise to denoise the latent variable of the noisy global image input into UNet, obtaining the latent variable of the denoised global image at the next time step, which is used as the denoising input of the diffusion repair model at the next time step; repeat this step until , obtaining the latent variable of the globally clean image without noise at the final time step. The fusion formula is as follows:

[0029]

[0030] Where represents the latent variable of the denoised global image, is the binary mask image corresponding to the background image, represents the latent variable of the noisy global image at time step t.

[0031] (2.9) Use the image decoder of the diffusion repair model to decode the latent variable of the globally clean image without noise obtained in step (2.8), obtaining the personalized completed image required by the user.

[0032] The beneficial effects of the present invention:

[0033] 1. The present invention realizes precise object insertion and attribute-level modification through text guidance, overcomes the problems of global noise interference and insufficient local information integration in the existing methods during object editing, and significantly improves the practical application ability in multiple scenarios.

[0034] 2. The present invention innovatively proposes a separated image inpainting framework, which divides the image inpainting process into two stages: local content generation and global context harmony, enhancing both the local feature details of the object and the overall visual consistency.

[0035] 3. The present invention effectively alleviates the problem of object appearance overfitting by introducing an attribute decoupling mechanism and a text-attribute substitution module, supporting diverse attribute editing and higher-quality text-driven object editing.

[0036] 4. The present invention outperforms mainstream methods in multiple quantitative metrics such as CLIP-T, CLIP-I, and DINO in identity preservation and attribute editing tasks. At the same time, it also performs outstandingly in the subjective evaluation of user research, and can generate patched results with identity preservation and context harmony.

[0037] 5. Compared with few-shot fine-tuning methods such as DreamBooth, DreamMix avoids phenomena such as filling in irrelevant content and losing main body features; compared with other large-scale training methods, the present invention achieves high-quality editing effects with less data requirements, and has significant advantages in identity preservation and attribute editing effects, with broad potential for technical promotion. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is the overall framework diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.

[0040] The invention is a text-driven subject personalization inpainting method based on an image inpainting model, and proposes a technical solution called DreamMix. DreamMix designs a personalized fine-tuning method for feature decoupling, designs a vector decomposition method for further decoupling, and designs a two-stage image inpainting framework to improve the image quality.

[0041] For the feature learning of the image subject, similar to DreamBooth, DreamMix fine-tunes the image inpainting model based on the Lora method to enable the model to learn the features of the subject. On this basis, in order to weaken the feature coupling, first, based on the target problem proposed by the user, a multimodal large language model is used to describe the attributes of the image subject, obtaining a key-value pair of attribute descriptions. Then, with the help of the multimodal large language model and the key-value pair, multiple description statements are generated as regular text. Next, with the help of the text-controlled image generation model, the regular images corresponding to the regular text are generated. Then, in the fine-tuning stage, the subject image and text description, as well as the regular image and text description, are simultaneously used to train the image inpainting model, where the regular text does not contain attribute words [sks] to avoid the [sks] words from learning decoupled features. During the fine-tuning process, different from other methods that use randomly generated masks, this method uses the masks of the subject images and regular image subjects generated by the image segmentation model for model fine-tuning.

[0042] To further weaken the attribute coupling, DreamMix performs vector decomposition on the text features. Specifically, for the input text, first, the text feature vector is obtained through the text encoder. Then, the attributes associated with the input text are selected from the attribute key-value pairs obtained from the multimodal large language model and fed into the text encoder to obtain their corresponding feature vectors. Then, the orthogonal decomposition formula is applied to obtain the decoupled feature vectors, which are used to control the generation of the image inpainting model as text conditions.

[0043] To improve the quality of image inpainting and enhance the detailed features of the subject and the overall harmony of the image, this method proposes a two-stage image inpainting framework. First, to improve the detailed features of the subject, according to the completion area specified by the user on the binary mask image, the background image is locally cropped and image inpainting is performed for a certain number of steps. Next, single-step denoising is performed and decoded by the image decoder to obtain a clean local image. Then, the clean local image is re-fused into the background image and encoded by the image encoder, and the remaining steps of image inpainting are performed on the global image. As Figure 1 shown, a text-driven subject personalized completion method based on an image inpainting model is as follows:

[0044] Step 1: Personalize and fine-tune the image inpainting model;

[0045] (1.1) Obtain a set of images with a common target subject and the corresponding text descriptions of the images. Each image and its corresponding text description form a subject data pair, which constitutes the training data for fine-tuning the image inpainting model to enable the image inpainting model to learn the features of the target subject. The corresponding text description of the image refers to a concise description of the target subject containing an identity marker.

[0046] (1.2)Based on the target question proposed by the user, use a multimodal large language model to describe the attributes of the target subject in the set of images with a common target subject described in step (1.1), and obtain key-value pairs including all the feature descriptions. Based on the key-value pairs, use the multimodal large language model again, and based on the further target question proposed by the user, obtain more regular text descriptions of the target subject. The regular text is a detailed description that does not contain identity markers and includes some of the target subject features in the key-value pairs (see Figure 1 in (a)).

[0047] (1.3)With the help of an image generation model based on text control, generate a specified number of regular images based on the regular text obtained in step (1.2), and obtain a regular data pair including the regular text and the regular images.

[0048] (1.4)With the help of an image segmentation model, obtain the masks of all the images in the subject data pairs and the regular data pairs, and add the masks to the corresponding subject data pairs and regular data pairs for fine-tuning of the image restoration model. All the images include the regular images obtained in step (1.3) and the images of the target subject in step (1.1) (see Figure 1 in (b)).

[0049] (1.5)Use the subject data pairs and regular data pairs with masks obtained in step (1.4) to fine-tune the image restoration model. Different sampling probabilities should be used for different subject data pairs and regular data pairs, and fine-tuning training of the image restoration model should be carried out based on the LoRA method to learn the features of the target subject.

[0050] Preferably, the image restoration model uses a diffusion restoration model:

[0051] Randomly add noise to all the images to time step t, and send the noisy images into the UNet of the diffusion restoration model. After passing through the initial convolutional layer, at the same time, dilate the mask obtained in step (1.4) to a random size, and encode it together with the corresponding image. Add the encoded whole to the result after passing through the initial convolutional layer, and perform subsequent downsampling and upsampling processing in the UNet to predict the added noise; during the training process of the diffusion restoration model, it is necessary to fine-tune the key linear transformation parameter matrix and the value linear transformation parameter matrix of the attention layer of the diffusion restoration model based on the LoRA method; it is necessary to calculate the reconstruction loss, and use the dilated mask to weight the loss to update the parameters of LoRA.

[0052] Step Two: Two-stage image completion;

[0053] During the process of image completion, the input of the fine-tuned image inpainting model includes the background image to be completed, the binary mask image corresponding to the background image, the completion area specified by the user on the binary mask image, and the personalized text description containing identity markers for the target subject set by the user. The completion area is used to complete the target subject, and local images and local binary masks are obtained by cropping areas of the same position in the background image and the binary mask image that contain and are larger than the completion area to a specified size; the personalized text description containing identity markers for the target subject set by the user, the background image to be completed, the binary mask image corresponding to the background image, the local image, and the local binary mask are fed into the fine-tuned image inpainting model.

[0054] (2.2)The personalized text description containing identity markers for the target subject set by the user is encoded using a text encoder to obtain ,and the key-value pairs obtained in step (1.2) are encoded using a text encoder to obtain ,Next, vector decomposition is applied to decouple the features, and the decoupled feature vectors are obtained (see (c) in Figure 1 ), and the decoupling formula is as follows:

[0055]

[0056] where represents the norm of the vector;

[0057] (2.3)The latent variable of the local image is obtained by encoding the local image, random noise of the same dimension is initialized with a standard normal distribution, and the latent variable of the local image is forward-denoised by the diffusion inpainting model to time step T, where T is the total number of denoising steps of the diffusion inpainting model, to obtain the latent variable of the noisy local image; then the latent variable of the noisy local image and the decoupled feature vectors obtained in step (2.2) are fed into UNet for denoising to obtain the initial predicted noise;

[0058] (2.4) The diffusion repair model uses the initial predicted noise obtained in step (2.3) to denoise the latent variables of the noisy local image before denoising in step (2.3), obtaining the latent variables of the denoised local image. Next, random noise of the same dimension is initialized using the standard normal distribution in step (2.3), and the latent variables of the local image obtained in step (2.3) are forward-noised to the current time step, obtaining the latent variables of the noisy local image at the current time step. The latent variables of the denoised local image and the latent variables of the noisy local image at the current time step are fused using a local binary mask to obtain the latent variables of the noisy local image to be input into UNet, which are sent into UNet for denoising together with the decoupled feature vectors obtained in step (2.2), obtaining the predicted noise at the current time step as the denoising input for the diffusion repair model at the next time step. Repeat this step until , the predicted noise at the last time step is obtained. Among them is a hyperparameter set manually, and T is the total number of noise addition steps of the diffusion repair model. The fusion formula is as follows:

[0059]

[0060] Among them represents the latent variables of the denoised local image, is the local binary mask, represents the latent variables of the noisy local image at time step t; represents the element-wise multiplication operation of matrices.

[0061] (2.5) In the diffusion repair model, the predicted noise at the last time step is used to thoroughly denoise the latent variables of the denoised local image after fusion at the last time step, directly predicting the clean latent variables, and then the decoder is used to decode the clean latent variables to obtain the local clean image. Next, according to the cropping area in step (2.1), the local clean image is fused with the input background image to obtain the global clean image, and then the encoder is used to encode to obtain the latent variables corresponding to the global clean image.

[0062] (2.6) The global clean image obtained in step (2.5) and the binary mask image corresponding to the background image input in step (2.1) are encoded together and sent into UNet. At the same time, in the diffusion repair model, random noise of the same dimension is initialized using the standard normal distribution in step (2.3), and the latent variables corresponding to the global clean image obtained in step (2.5) are forward-noised to , obtaining the latent variables after noise addition to the clean image, and then they are also sent into UNet to obtain the global predicted noise.

[0063] (2.7) Denoise the latent variables of the clean image after adding noise in step (2.6) using the global prediction noise obtained in step (2.6) to obtain the latent variables of the denoised global image;

[0064] (2.8) The diffusion repair model initializes random noise of the same dimension using the standard normal distribution in step (2.3), and forward adds noise to the latent variables corresponding to the global clean image obtained in step (2.5) to the current time step to obtain the latent variables of the noisy global image at the current time step; fuse the latent variables of the denoised global image obtained in step (2.7) and the latent variables of the noisy global image at the current time step using the binary mask image corresponding to the background image to obtain the latent variables of the noisy global image to be input into UNet, and send them into UNet together with the decoupled feature vectors obtained in step (2.2) for denoising to obtain the prediction noise at the current time step. Use the prediction noise to denoise the latent variables of the noisy global image input into UNet to obtain the latent variables of the denoised global image at the next time step, which is used as the denoising input of the diffusion repair model at the next time step; repeat this step until , the latent variables of the global image without noise at the final time step are obtained. The fusion formula is as follows:

[0065]

[0066] where represents the latent variables of the denoised global image, is the binary mask image corresponding to the background image, represents the latent variables of the noisy global image at time step t.

[0067] (2.9) Use the image decoder of the diffusion repair model to decode the latent variables of the global image without noise obtained in step (2.8) to obtain the personalized completed image required by the user (see (d) in Figure 1 ).

[0068] Example:

[0069] In this embodiment, the framework of the present invention is built based on the repair model of Fooocus, and the DreamBench personalized image dataset is used for fine-tuning. The LoRA with rank 4 is applied to and the attention layer. During the fine-tuning process, the image is adjusted to 1024×1024, the AdamW optimizer is used, and the learning rate is 1e-4. On a single RTX 4090 GPU, the fine-tuning time for each object is about 20 minutes.

[0070] The present invention is compared with existing excellent methods in two tasks of personalized image generation and text-driven attribute editing. Specifically, the present invention selects thirty subjects of DreamBench and designs 4,000 image-mask pairs for testing on COCO-val2017. The comparison results of each index are as follows:

[0071]

[0072] Among them, CLIP-I and DINO evaluate the image similarity between the generated image and the subject, and CLIP-T calculates the relevance between the generated image and the text condition.

[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A text-driven subject personalized completion method based on an image restoration model, characterized in that: Here are the steps: Step 1: Personalized fine-tuning of image restoration model; Step 2: Two-stage image completion; The step 1 specifically includes: (1.1) Obtain a set of images with a common target subject and text descriptions corresponding to the images, wherein each image and its corresponding text description serve as a subject data pair, constituting training data for a fine-tuned image restoration model, so that the image restoration model can learn the characteristics of the target subject; the text description corresponding to the image refers to a concise description of the target subject containing an identity tag; (1.2) Based on the target question raised by the user, a multimodal large language model is used to describe the attributes of the target subject in the set of images with a common target subject described in step (1.1), and a key-value pair including all feature descriptions is obtained; based on the key-value pair, the multimodal large language model is used again, and based on the target question further raised by the user, more regular text about the target subject description is obtained; the regular text is a detailed description of some target subject features in the key-value pair without identity tags; (1.3) using a text-controlled image generation model, generating a specified number of regular images based on the regular text obtained in step (1.2) to obtain regular data pairs including regular text and regular images; (1.4) Obtain masks of all images in the subject data pair and the regular data pair with the help of the image segmentation model, and add the masks to the corresponding subject data pair and the regular data pair for fine-tuning the image restoration model; all images include the regular image obtained in step (1.3) and the image of the target subject in step (1.1); (1.5) Using the masked subject data pairs and regular data pairs obtained in step (1.4), fine-tune the image restoration model. Different sampling probabilities are used for different subject data pairs and regular data pairs. The image restoration model is fine-tuned based on the Lora method to learn the characteristics of the target subject. The step 2 comprises: (2.1) During the image completion process, the input of the fine-tuned image inpainting model includes the background image to be completed, the binary mask image corresponding to the background image, the completion area specified by the user on the binary mask image, and the personalized text description containing the identity mark of the target subject set by the user; the completion area is used to complete the target subject, and the area of ​​the specified size is cropped from the same position in the background image and the binary mask image that is contained in and larger than the completion area to obtain a local image and a local binary mask; the personalized text description containing the identity mark of the target subject set by the user, the background image to be completed, the binary mask image corresponding to the background image, the local image, and the local binary mask are sent to the fine-tuned image inpainting model; (2.2) Encode the personalized text description containing identity tags set by the user for the target subject using a text encoder , use the text encoder to encode the key-value pairs obtained in step (1.2) , then apply vector decomposition to decouple the features and obtain the decoupled feature vector. The decoupling formula is as follows: in, represents the magnitude of a vector; (2.3) Encode the local image to obtain the latent variable of the local image, initialize the random noise of the same dimension with the standard normal distribution, and perform forward noise addition of the diffusion repair model on the latent variable of the local image to time step T, where T is the total number of noise addition steps of the diffusion repair model, to obtain the latent variable of the noisy local image; then, the latent variable of the noisy local image and the decoupled feature vector obtained in step (2.2) are simultaneously sent to UNet for denoising to obtain the initial prediction noise; (2.4) The diffusion repair model uses the initial predicted noise obtained in step (2.3) to denoise the latent variable of the noisy local image before denoising in step (2.3) to obtain the latent variable of the denoised local image; next, the random noise of the same dimension is initialized using the standard normal distribution in step (2.3), and the latent variable of the local image obtained in step (2.3) is forward-noised to the current time step to obtain the latent variable of the noisy local image at the current time step; the latent variable of the denoised local image and the latent variable of the noisy local image at the current time step are fused using a local binary mask to obtain the latent variable of the noisy local image that needs to be input into UNet, and the latent variable is sent to UNet for denoising together with the decoupled feature vector obtained in step (2.2) to obtain the predicted noise of the current time step as the denoising input of the diffusion repair model for the next time step; repeat this step until , and get the prediction noise of the last time step; where is a manually set hyperparameter, T is the total number of noise addition steps of the diffusion repair model; the fusion formula is as follows: in represents the latent variable of the denoised local image, is a local binary mask, represents the latent variable of the noisy local image at time step t; Represents the element-by-element multiplication operation of a matrix; (2.5) In the diffusion inpainting model, the latent variables of the denoised local image after fusion at the last time step are thoroughly denoised using the predicted noise at the last time step, and the direct prediction is The clean latent variable is then decoded by the decoder to obtain a local clean image. Next, according to the cropped area in step (2.1), the local clean image is fused with the input background image to obtain a global clean image, and then the encoder is used to encode the latent variable corresponding to the global clean image. (2.6) Encode the global clean image obtained in step (2.5) and the binary mask image corresponding to the background image input in step (2.1) together and send them to UNet; at the same time, in the diffusion repair model, use the standard normal distribution in step (2.3) to initialize the random noise of the same dimension, and add the noise to the latent variable corresponding to the global clean image obtained in step (2.5) , get the latent variable after adding noise to the clean image, and then send it to UNet to get the global prediction noise; (2.7) Using the global prediction noise obtained in step (2.6), denoise the latent variable of the clean image after adding noise in step (2.6) to obtain the latent variable of the denoised global image; (2.8) The diffusion repair model uses the standard normal distribution in step (2.3) to initialize random noise of the same dimension, and adds noise to the latent variable corresponding to the global clean image obtained in step (2.5) to the current time step to obtain the latent variable of the noisy global image at the current time step; the latent variable of the denoised global image obtained in step (2.7) and the latent variable of the noisy global image at the current time step are fused using the binary mask image corresponding to the background image to obtain the latent variable of the noisy global image that needs to be input into UNet, and the latent variable is sent to UNet for denoising together with the decoupled feature vector obtained in step (2.2) to obtain the predicted noise of the current time step, and the predicted noise is used to denoise the latent variable of the noisy global image input into UNet to obtain the latent variable of the denoised global image at the next time step, which is used as the denoising input of the diffusion repair model at the next time step; repeat this step until , get the latent variable of the global picture without noise at the final time step; the fusion formula is as follows: in represents the latent variable of the denoised global picture, is the binary mask image corresponding to the background image, The latent variable representing the noisy global picture at time step t; (2.9) Use the image decoder of the diffusion inpainting model to decode the latent variable of the noise-free global image obtained in step (2.8) to obtain the personalized completed image required by the user.

2. According to the text-driven subject personalized completion method based on the image restoration model of claim 1, it is characterized in that: The image restoration model in step 1 adopts a diffusion restoration model: randomly add noise to all images to time step t, send the noisy images to the UNet of the diffusion restoration model, pass through the initial convolution layer, and at the same time, dilate the mask obtained in step (1.4) with a random size and encode it together with the corresponding image, add the encoded whole to the result after the initial convolution layer, and perform subsequent downsampling and upsampling processing of UNet to predict the added noise; During the training of the diffusion repair model, it is necessary to linearly transform the key parameter matrix of the attention layer of the diffusion repair model based on the lora method Sum value linear transformation parameter matrix Perform fine-tuning training; it is necessary to calculate the reconstruction loss and use the expanded mask to weight the loss to update the parameters of LoRa.

Citation Information

Patent Citations

  • Text image editing method and system based on character attribute guidance

    CN114863441A

  • Decoupling self-enhancement-based detail-controllable personalized image generation method and system

    CN117876522A