Image target repair method, device, equipment, and storage medium

By using vague codes and text prompts in the image repair method to infer the repair features and guide the diffusion model to generate target objects, the inconsistency and artifact problems of objects and text in the prior art are solved, and the high-fidelity image repair effect is achieved.

CN118967524BActive Publication Date: 2025-05-13BEIJING ZHIXIANG FUTURE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411082632.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2025-05-13
Estimated Expiration
2044-08-08

AI Technical Summary

Technical Problem

The existing image repair method based on diffusion model-based text-guided has the problem that the generated objects are inconsistent with the text, and there are obvious artifacts in the missing areas.

Method used

By obtaining images including repaired areas, text prompts and binary masks, the mask image of the image is determined, and the repaired features of the target object are inferred based on the text prompts and mask images. The repaired features are used as visual hints to guide the diffusion model to generate the target object.

Benefits of technology

The generated target object is consistent with the text prompt and there is no obvious artifact, achieving high-fidelity image repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118967524B_ABST
    Figure CN118967524B_ABST
Patent Text Reader

Abstract

The present application provides an image target restoration method, device, equipment, and storage medium, the method comprising: obtaining an image including a restoration area, a text prompt, and a binary mask code; the text prompt is used to describe the target object in the restoration area; the binary mask code is used to indicate the area to be restored; the mask image of the image is determined according to the binary mask code; the restored features of the target object are inferred according to the text prompt and the mask image; the restored features are used as visual prompts to guide the diffusion model to generate the target object. The method provided by the present application infers the restored features of the target object based on the image, text prompt, and binary mask code of the restoration area, and then uses the restored features as visual prompts to guide the diffusion model to generate the target object, so that the generated target object is consistent with the text prompt, and there are no obvious artifacts, with high fidelity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image target restoration method, device, equipment, and storage medium. Background Art

[0002] At present, the text-guided image restoration method based on the diffusion model shows the best performance. It can be divided into two categories. The first category is the method that directly continues the existing text-to-image diffusion model. In each step of the denoising process, the new objects generated from random noise are mixed with the background. It is a method that does not require training. The other method is based on training. By randomly erasing part of the content of the picture and giving the erased content a title to obtain a dataset, this dataset is used to fine-tune a pre-trained text-to-image diffusion model so that the model can restore the image through the corresponding title content.

[0003] Although the first type of methods show promising results, they only use simple region-by-region blending, which makes the images they produce have obvious artifacts in the missing areas.

[0004] The second type of method has the problem that the generated objects are inconsistent with the text. This is because the text describes the entire image, not just the missing local area. This may result in the generation of new areas being out of the control of the text if there is content related to the text in the background. Summary of the invention

[0005] In order to solve one of the above-mentioned technical defects, the present application provides an image target repair method, device, equipment, and storage medium.

[0006] In a first aspect, the present application provides an image object restoration method, the method comprising:

[0007] Acquire an image including a repaired area, a text prompt and a binary mask code; the text prompt is used to describe a target object in the repaired area; the binary mask code is used to indicate the area to be repaired;

[0008] Determine a mask image of the image according to the binary mask code;

[0009] Infer the restored features of the target object based on the textual hint and the mask image;

[0010] The restored features are used as visual cues to guide the diffusion model to generate the target object.

[0011] Optionally, determining a mask image of the image according to the binary mask code includes:

[0012] Determine the mask image x'=x⊙(1-m) of the image;

[0013] Among them, x' is the mask image, x is the image, m is the binary mask code, and ⊙ is the multiplication operator.

[0014] Optionally, based on the textual hint and the mask image, infer the restored features of the target object, including:

[0015] Extracting visual features from mask images in, is the feature vector, i is the feature identifier, N is the feature length, d is the feature dimension;

[0016] According to the text prompts and Reconstructing semantic features

[0017] Based on semantic features Determine the repaired features of the target object.

[0018] Optionally, based on text prompts and Reconstructing semantic features include:

[0019] Will Connect with the text prompt to get the connection feature, where PE is the position code, PE∈R N×d , ME is the mask code, ME∈R N×d ;

[0020] Reconstructing semantic features from concatenated features

[0021] Optionally, based on semantic features Determine the post-repair features of the target object, including:

[0022] Determine the restored features of the target object

[0023] Among them, m is a binary mask code, ⊙ is a multiplication operator, and DownSample(·) is downsampling.

[0024] Optionally, the diffusion model includes a U-Net module; the U-Net module includes a plurality of attention blocks;

[0025] Each attention block consists of a self-attention layer, a reference adapter layer, and a cross-attention layer;

[0026] In each attention block, the output of the self-attention layer is the input of the reference adapter layer; the output of the reference adapter layer is the input of the cross-attention layer; the output of the cross-attention layer is the input of the next attention block;

[0027] The reference adapter layer is a multi-head cross-attention layer.

[0028] Optionally, the output H of the self-attention layer s =SelfAttn(H,C,C)+H; H is the input of the attention block where the self-attention layer is located, C is the text feature extracted from the text prompt, and SelfAttn(·) is the self-attention layer;

[0029] Output of the reference adapter layer in, is the repaired feature, RefAdapter(·) is the reference adapter layer;

[0030] The output H of the cross attention layer x =CrossAttn(H r ,C,C)+H r .

[0031] In a second aspect of the present application, an image object restoration device is provided, the device comprising:

[0032] An acquisition module is used to acquire an image including a repair area, a text prompt and a binary mask code; the text prompt is used to describe a target object in the repair area; and the binary mask code is used to indicate the area to be repaired;

[0033] A determination module, used for determining a mask image of the image according to the binary mask code obtained by the acquisition module;

[0034] An inference module, used to infer the restored features of the target object according to the text prompt obtained by the acquisition module and the mask image determined by the determination module;

[0035] The generation module is used to use the repaired features obtained by the inference module as visual cues to guide the diffusion model to generate the target object.

[0036] In a third aspect of the present application, an electronic device is provided, including:

[0037] Memory;

[0038] Processor; and

[0039] Computer programs;

[0040] The computer program is stored in the memory and is configured to be executed by the processor to implement the method as described in the first aspect above.

[0041] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored; the computer program is executed by a processor to implement the method described in the first aspect above.

[0042] The present application provides an image target restoration method, device, equipment, and storage medium, the method comprising: obtaining an image including a restoration area, a text prompt, and a binary mask code; the text prompt is used to describe the target object in the restoration area; the binary mask code is used to indicate the area to be restored; the mask image of the image is determined according to the binary mask code; the restored features of the target object are inferred according to the text prompt and the mask image; the restored features are used as visual prompts to guide the diffusion model to generate the target object. The method provided by the present application infers the restored features of the target object based on the image, text prompt, and binary mask code of the restoration area, and then uses the restored features as visual prompts to guide the diffusion model to generate the target object, so that the generated target object is consistent with the text prompt, and there are no obvious artifacts, with high fidelity. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0044] Figure 1 A schematic diagram of a process flow of an image object restoration method provided in an embodiment of the present application;

[0045] Figure 2 A schematic diagram of an implementation architecture of an image object restoration method provided in an embodiment of the present application;

[0046] Figure 3 A comparison diagram of the effect of an image object restoration method provided by an embodiment of the present application and the effect of an existing segmentation mask method;

[0047] Figure 4 A comparison diagram of the effect of an image object restoration method provided by an embodiment of the present application and the effect of an existing bounding box mask method;

[0048] Figure 5 A schematic diagram of the structure of an image object restoration device provided in an embodiment of the present application;

[0049] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0050] In order to make the technical solutions and advantages in the embodiments of the present application more clearly understood, the exemplary embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than an exhaustive list of all the embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0051] In the process of realizing the present application, the inventors found that the current image restoration method guided by text based on diffusion model shows the best performance. It is mainly divided into two categories. The first category is a method that directly continues the existing text-to-image diffusion model. In each step of the denoising process, the new objects generated from random noise are mixed with the background. It is a method that does not require training. The other method is based on training. By randomly erasing part of the content of the picture and giving the erased content a title to obtain a data set, and using this data set to fine-tune a pre-trained text-to-image diffusion model, the model can restore the image through the corresponding title content. Although the first method shows promising results, it only uses a simple regional mixing, so that the image it generates has obvious artifacts in the missing area. The second method has the problem of inconsistency between the generated object and the text. This is because the text describes the entire picture, not just the missing local area, which may result in the generation of new areas being out of the control of the text if there is text-related content in the background part.

[0052] In response to the above problems, the present application provides an image target repair method, device, equipment, and storage medium. The method includes: obtaining an image including a repair area, a text prompt, and a binary mask code; the text prompt is used to describe the target object in the repair area; the binary mask code is used to indicate the area to be repaired; the mask image of the image is determined according to the binary mask code; according to the text prompt and the mask image, the repaired features of the target object are inferred; the repaired features are used as visual cues to guide the diffusion model to generate the target object. The method provided by the present application infers the repaired features of the target object based on the image, text prompt, and binary mask code of the repair area, and then uses the repaired features as visual cues to guide the diffusion model to generate the target object, so that the generated target object is consistent with the text prompt, and there are no obvious artifacts, with high fidelity.

[0053] See also Figure 1 The implementation process of the image target restoration method provided in this embodiment is as follows:

[0054] 101, obtaining an image including a repaired area, a text prompt, and a binary mask code.

[0055] The text hint is used to describe the target object in the repair area, and the binary mask code is used to indicate the area to be repaired.

[0056] Where x is an image (including the inpainted region), c is a textual cue describing the new object, and m is a binary mask indicating the region to be inpainted, current text-guided object inpainting diffusion models simply optimize the alignment of the textual cue c with the desired object x⊙m in the latent space, which leads to suboptimal results. To improve this problem, two main issues need to be addressed: (1) For text-guided image inpainting, it is not sufficient to rely solely on a single U-Net to achieve visual-semantic alignment in all denoising time steps; (2) It is challenging to stably generate high-fidelity objects in complex sampling spaces without additional semantic guidance.

[0057] It should be noted that ⊙ is a multiplication operator, and x⊙m means the point-by-point multiplication of x and m.

[0058] In order to alleviate the deficiency of relying solely on a single U-Net to align text prompts and visual objects in the entire denoising process, the image target repair method provided in this embodiment obtains the repaired features of the target object through steps 102 and 103 outside the U-Net, and the features can enhance the visual-semantic correspondence.

[0059] 102, determining a mask image of the image according to the binary mask code.

[0060] For example, a mask image x'=x⊙(1-m) of the image is determined.

[0061] Among them, x' is the mask image, x is the image, m is the binary mask code, and ⊙ is the multiplication operator.

[0062] 103, based on the textual hint and the mask image, the restored features of the target object are inferred.

[0063] The inference process of step 103 is:

[0064] 103-1, Extracting visual features from mask images

[0065] in, is the feature vector, i is the feature identifier, N is the feature length, d is the feature dimension.

[0066] In the specific implementation, the visual features can be extracted from the mask image x⊙(1-m) through the CLIP (Contrastive Language-Image Pre-Training) image encoder

[0067] For example,

[0068] Among them, ImageEnc(·) is a related method for realizing visual feature extraction.

[0069] is a feature vector of a local image.

[0070] 103-2, according to the text prompts and Reconstructing semantic features

[0071] Given the characteristics of the damage and text prompt c, step 103-2 needs to predict the semantic features that can stably reconstruct the target object in the CLIP space This allows visual-semantic alignment to be naturally achieved throughout the entire denoising process of the diffusion model.

[0072] In step 103-2, Connect with text prompts to obtain connection features, and reconstruct semantic features from the connection features

[0073] Among them, PE is the learnable position code, PE∈R N×d , ME is a learnable mask encoding, ME∈R N×d If the part is 1, it means that the visual feature is a mask area, and if it is 0, it is a non-mask area.

[0074] For example,

[0075] Among them, SemInpainter(·) is a related method for reconstructing semantic features, and [·] indicates a connection operation.

[0076] 103-3, based on semantic features Determine the repaired features of the target object.

[0077] For example, determine the restored features of the target object

[0078] Among them, m is a binary mask code, ⊙ is a multiplication operator, and DownSample(·) is downsampling.

[0079] Through steps 102 and 103, the restored features of the target object can be predicted based on the unmasked image context and textual cues.

[0080] Step 103 can be implemented by a semantic repairer, which can be implemented by a 24-layer transformer structure (such as an image encoder structure similar to CLIP). Through step 103, a repaired feature This feature is highly semantically aware and context-aware of non-repaired areas, and the result is mixed with the semantic features of unmasked areas.

[0081] The semantic repairer uses an effective knowledge distillation objective in the multimodal feature space to guide the learning of the semantic repairer, which can naturally align the repaired semantic features with the textual cues and visual objects. In addition, the output of the semantic repairer is used as a semantically aligned visual cue and is input into the object repair diffusion model through an additional reference adapter layer, which can enhance the controllability of the diffusion model for object repair (this process is implemented in step 104).

[0082] 104, the restored features are used as visual cues to guide the diffusion model to generate the target object.

[0083] Among them, the diffusion model includes the U-Net module.

[0084] In step 104, in order to ensure that the generated target object features and textual hints are aligned, multimodal knowledge can be transferred from the teacher model (such as the above-mentioned CLIP) to the semantic repairer. That is, the semantic repairer is trained by performing the text-guided mask feature prediction task to recover the true semantic features of the masked position in the space of CLIP.

[0085] At the same time, since controlling the diffusion model to achieve high-fidelity text-to-image generation remains challenging, the prediction results of the semantic repairer are used as additional conditions, called visual cues, in the implicit space of the diffusion model to guide it to perform controllable repairs.

[0086] like Figure 2 As shown in the lower part of , a new reference adapter layer is introduced in the basic diffusion model. The reference adapter layer is a multi-head cross attention layer. The reference adapter layer is alternately inserted between the original self-attention layer and the cross attention layer of the U-Net model.

[0087] That is to say, the U-Net module includes multiple attention blocks. Each attention block consists of a self-attention layer, a reference adapter layer, and a cross-attention layer.

[0088] In each attention block, the output of the self-attention layer is the input of the reference adapter layer. The output of the reference adapter layer is the input of the cross-attention layer. The output of the cross-attention layer is the input of the next attention block. The embedding of semantic features is completed, achieving precise control in complex spaces.

[0089] set up is the intermediate hidden state of the U-Net model, where N h Varies with the resolution of different layers. are the text features of the text cue c extracted by the text encoder in the diffusion model.

[0090] The output of the self-attention layer H s =SelfAttn(H,C,C)+H.

[0091] Among them, H is the input of the attention block where the self-attention layer is located, C is the text feature extracted from the text prompt, and SelfAttn(·) is the self-attention layer.

[0092] Output of the reference adapter layer

[0093] in, is the repaired feature, and RefAdapter(·) is the reference adapter layer.

[0094] The output H of the cross attention layer x ·CrossAttn(H r ,C,C)+H r .

[0095] In order to facilitate data interaction between layers, a shared module can be defined for different types of attention layers, and the data (such as output) can be stored in the shared module, and the corresponding data can be read from the shared module as input later.

[0096] Shared modules such as Output = Module (Query, Key, Value). Among them,

[0097] Query is the query word of the shared module, Key is the key extracted from the shared module, Value is the value extracted from the shared module, and Module(·) is the query method of the shared module.

[0098] The image object restoration method provided in this embodiment can generate higher quality image restoration effects, and decomposes the typical single-stage text-guided image restoration process into two cascade stages: 1) firstly, semantic pre-repair is performed, that is, the restored features of the target object are inferred in the multimodal feature space through steps 102 and 103; 2) then object generation is performed, that is, in step 104, the inferred semantic features are used as visual cues to guide the diffusion model, thereby achieving high-fidelity image restoration.

[0099] The effect of the image target restoration method provided by this embodiment is compared with the effect of the existing segmentation mask method. Figure 4 The effect of the image target restoration method provided by this embodiment is compared with the effect of the existing bounding box mask method as shown in FIG. Figure 5 shown.

[0100] The image target restoration method provided in this embodiment is based on the paradigm of two-stage text-guided image restoration, which divides the restoration process into semantic pre-filling and object generation, and proves its ability of high-fidelity object restoration. The image target restoration method provided in this embodiment makes full use of the text-image alignment capability of the cross-modal semantic space to avoid low-quality alignment in the noisy sampling space. The reference adapter layer of the diffusion model can well control the object restoration by adjusting the visual cues (i.e., the semantic features after restoration). The image target restoration method provided in this embodiment has the most advanced performance in the text-guided image restoration task.

[0101] The present embodiment provides an image target restoration method, which obtains an image including a restoration area, a text prompt and a binary mask code; the text prompt is used to describe the target object in the restoration area; the binary mask code is used to indicate the area to be restored; the mask image of the image is determined according to the binary mask code; the restored features of the target object are inferred according to the text prompt and the mask image; the restored features are used as visual prompts to guide the diffusion model to generate the target object. The method provided in the present embodiment infers the restored features of the target object based on the image, text prompt and binary mask code of the restoration area, and then uses the restored features as visual prompts to guide the diffusion model to generate the target object, so that the generated target object is consistent with the text prompt, and there are no obvious artifacts, with high fidelity.

[0102] Based on the same inventive concept of the image object restoration method, this embodiment provides an image object restoration device, such as Figure 5 As shown, the device comprises:

[0103] The acquisition module 501 is used to acquire an image including a repair area, a text prompt and a binary mask code. The text prompt is used to describe the target object in the repair area. The binary mask code is used to indicate the area to be repaired.

[0104] The determination module 502 is used to determine the mask image of the image according to the binary mask code obtained by the acquisition module 501 .

[0105] The inference module 503 is used to infer the restored features of the target object according to the text prompt obtained by the acquisition module 501 and the mask image determined by the determination module 502.

[0106] The generation module 504 is used to use the restored features obtained by the inference module 503 as visual cues to guide the diffusion model to generate the target object.

[0107] The determination module 502 is used to determine the mask image x'=x⊙(1-m) of the image.

[0108] Among them, x' is the mask image, x is the image, m is the binary mask code, and ⊙ is the multiplication operator.

[0109] The inference module 503 is used to extract visual features from the mask image. in, is the feature vector, i is the feature identifier, N is the feature length, d is the feature dimension. Reconstructing semantic features Based on semantic features Determine the repaired features of the target object.

[0110] Among them, the inference module 503 is used to Connect with the text prompt to get the connection feature, where PE is the position code, PE∈R N×d , ME is the mask code, ME∈R N×d . Reconstructing semantic features from concatenated features

[0111] The inference module 503 is used to determine the restored features of the target object.

[0112] Among them, m is a binary mask code, ⊙ is a multiplication operator, and DownSample(·) is downsampling.

[0113] The diffusion model includes a U-Net module, which includes multiple attention blocks.

[0114] Each attention block consists of a self-attention layer, a reference adapter layer, and a cross-attention layer.

[0115] In each attention block, the output of the self-attention layer is the input of the reference adapter layer. The output of the reference adapter layer is the input of the criss-cross attention layer. The output of the criss-cross attention layer is the input of the next attention block.

[0116] The reference adapter layer is a multi-head cross-attention layer.

[0117] Among them, the output H of the self-attention layer s =SelfAttn(H,C,C)+H. H is the input of the attention block where the self-attention layer is located, C is the text feature extracted from the text prompt, and SelfAttn(·) is the self-attention layer.

[0118] Output of the reference adapter layer in, is the repaired feature, and RefAdapter(·) is the reference adapter layer.

[0119] The output H of the cross attention layer x =CrossAttn(H r ,C,C)+H r .

[0120] The device provided in this embodiment obtains an image, a text prompt and a binary mask code including a repair area; the text prompt is used to describe the target object in the repair area; the binary mask code is used to indicate the area to be repaired; the mask image of the image is determined according to the binary mask code; the repaired features of the target object are inferred according to the text prompt and the mask image; the repaired features are used as visual cues to guide the diffusion model to generate the target object. The method provided in this application infers the repaired features of the target object based on the image, text prompt and binary mask code of the repair area, and then uses the repaired features as visual cues to guide the diffusion model to generate the target object, so that the generated target object is consistent with the text prompt, and there are no obvious artifacts, with high fidelity.

[0121] Based on the same inventive concept of the image object restoration method, this embodiment provides an electronic device, the electronic device such as Figure 6 As shown, it includes: a memory 601, a processor 602, and a computer program.

[0122] The computer program is stored in the memory 601 and is configured to be executed by the processor 602 to implement the above-mentioned image object restoration method.

[0123] Specifically,

[0124] Get an image including the repair area, a text hint, and a binary mask code. The text hint is used to describe the target object in the repair area. The binary mask code is used to indicate the area to be repaired.

[0125] Determine a mask image for the image based on the binary mask code.

[0126] Based on the textual hint and the mask image, the inpainted features of the target object are inferred.

[0127] The restored features are used as visual cues to guide the diffusion model to generate the target object.

[0128] Optionally, determining a mask image of the image according to the binary mask code includes:

[0129] The mask image x'=x⊙(1-m) of the image is determined.

[0130] Among them, x' is the mask image, x is the image, m is the binary mask code, and ⊙ is the multiplication operator.

[0131] Optionally, based on the textual hint and the mask image, infer the restored features of the target object, including:

[0132] Extracting visual features from mask images in, is the feature vector, i is the feature identifier, N is the feature length, d is the feature dimension.

[0133] According to the text prompts and Reconstructing semantic features

[0134] Based on semantic features Determine the repaired features of the target object.

[0135] Optionally, based on text prompts and Reconstructing semantic features include:

[0136] Will Connect with the text prompt to get the connection feature, where PE is the position code, PE∈R N×d , ME is the mask code, ME∈R N×d .

[0137] Reconstructing semantic features from concatenated features

[0138] Optionally, based on semantic features Determine the post-repair features of the target object, including:

[0139] Determine the restored features of the target object

[0140] Among them, m is a binary mask code, ⊙ is a multiplication operator, and DownSample(·) is downsampling.

[0141] Optionally, the diffusion model includes a U-Net module. The U-Net module includes multiple attention blocks.

[0142] Each attention block consists of a self-attention layer, a reference adapter layer, and a cross-attention layer.

[0143] In each attention block, the output of the self-attention layer is the input of the reference adapter layer. The output of the reference adapter layer is the input of the criss-cross attention layer. The output of the criss-cross attention layer is the input of the next attention block.

[0144] The reference adapter layer is a multi-head cross-attention layer.

[0145] Optionally, the output H of the self-attention layer s=SelfAttn(H,C,C)+H. H is the input of the attention block where the self-attention layer is located, C is the text feature extracted from the text prompt, and SelfAttn(·) is the self-attention layer.

[0146] Output of the reference adapter layer in, is the repaired feature, and RefAdapter(·) is the reference adapter layer.

[0147] The output H of the cross attention layer x =CrossAttn(H r ,C,C)+H r .

[0148] The electronic device provided by this embodiment has a computer program executed by a processor to obtain an image including a repair area, a text prompt and a binary mask code; the text prompt is used to describe the target object in the repair area; the binary mask code is used to indicate the area to be repaired; the mask image of the image is determined according to the binary mask code; the repaired features of the target object are inferred according to the text prompt and the mask image; the repaired features are used as visual cues to guide the diffusion model to generate the target object. The method provided by this application infers the repaired features of the target object based on the image, text prompt and binary mask code of the repair area, and then uses the repaired features as visual cues to guide the diffusion model to generate the target object, so that the generated target object is consistent with the text prompt, and there are no obvious artifacts, with high fidelity.

[0149] Based on the same inventive concept of the image object restoration method, this embodiment provides a computer-readable storage medium, and a computer program is stored thereon. The computer program is executed by a processor to implement the above-mentioned image object restoration method.

[0150] Specifically,

[0151] Get an image including the repair area, a text hint, and a binary mask code. The text hint is used to describe the target object in the repair area. The binary mask code is used to indicate the area to be repaired.

[0152] Determine a mask image for the image based on the binary mask code.

[0153] Based on the textual hint and the mask image, the inpainted features of the target object are inferred.

[0154] The restored features are used as visual cues to guide the diffusion model to generate the target object.

[0155] Optionally, determining a mask image of the image according to the binary mask code includes:

[0156] The mask image x'=x⊙(1-m) of the image is determined.

[0157] Among them, x' is the mask image, x is the image, m is the binary mask code, and ⊙ is the multiplication operator.

[0158] Optionally, based on the textual hint and the mask image, infer the restored features of the target object, including:

[0159] Extracting visual features from mask images in, is the feature vector, i is the feature identifier, N is the feature length, d is the feature dimension.

[0160] According to the text prompts and Reconstructing semantic features

[0161] Based on semantic features Determine the repaired features of the target object.

[0162] Optionally, based on text prompts and Reconstructing semantic features include:

[0163] Will Connect with the text prompt to get the connection feature, where PE is the position code, PE∈R N×d , ME is the mask code, ME∈R N×d .

[0164] Reconstructing semantic features from concatenated features

[0165] Optionally, based on semantic features Determine the post-repair features of the target object, including:

[0166] Determine the restored features of the target object

[0167] Among them, m is a binary mask code, ⊙ is a multiplication operator, and DownSample(·) is downsampling.

[0168] Optionally, the diffusion model includes a U-Net module. The U-Net module includes multiple attention blocks.

[0169] Each attention block consists of a self-attention layer, a reference adapter layer, and a cross-attention layer.

[0170] In each attention block, the output of the self-attention layer is the input of the reference adapter layer. The output of the reference adapter layer is the input of the criss-cross attention layer. The output of the criss-cross attention layer is the input of the next attention block.

[0171] The reference adapter layer is a multi-head cross-attention layer.

[0172] Optionally, the output H of the self-attention layer s =SelfAttn(H,C,C)+H. H is the input of the attention block where the self-attention layer is located, C is the text feature extracted from the text prompt, and SelfAttn(·) is the self-attention layer.

[0173] Output of the reference adapter layer in, is the repaired feature, and RefAdapter(·) is the reference adapter layer.

[0174] The output H of the cross attention layer x =CrossAttn(H r ,C,C)+H r .

[0175] The computer-readable storage medium provided in this embodiment has a computer program executed by a processor to obtain an image including a repaired area, a text prompt, and a binary mask code; the text prompt is used to describe the target object in the repaired area; the binary mask code is used to indicate the area to be repaired; the mask image of the image is determined according to the binary mask code; the repaired features of the target object are inferred according to the text prompt and the mask image; the repaired features are used as visual prompts to guide the diffusion model to generate the target object. The method provided in this application infers the repaired features of the target object based on the image, text prompt, and binary mask code of the repaired area, and then uses the repaired features as visual prompts to guide the diffusion model to generate the target object, so that the generated target object is consistent with the text prompt, and there are no obvious artifacts, with high fidelity.

[0176] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.

[0177] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0178] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0180] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0181] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A method for image object restoration, characterized in that: The method comprises: Acquire an image including a repair area, a text prompt and a binary mask code; the text prompt is used to describe the target object in the repair area; the binary mask code is used to indicate the area to be repaired; Determine a mask image of the image according to the binary mask code; Inferring restored features of the target object according to the text prompt and the mask image; Using the restored features as visual cues to guide the diffusion model to generate the target object; The inferring the restored features of the target object according to the text prompt and the mask image includes: Extract visual features from the mask image ,in, is the feature vector, is the feature identifier, is the characteristic length, , is the feature dimension; According to the text prompt and , reconstruct semantic features ; Based on the semantic features , determine the restored features of the target object.

2. The method according to claim 1, characterized in that The step of determining a mask image of the image according to the binary mask code comprises: Determines the mask image for the image ; in, is the mask image, For the image, is the binary mask code, is the multiplication operator.

3. The method according to claim 1, characterized in that: The text prompt and , reconstruct semantic features ,include: Will Connect with the text prompt to obtain a connection feature, wherein, is the position code, , is the mask encoding, ; Reconstruct semantic features from the concatenated features .

4. The method according to claim 1, characterized in that Based on the semantic features , determining the restored features of the target object, including: Determine the restored features of the target object ; in, is the binary mask code, is the multiplication operator, is down sampling.

5. The method according to claim 1, characterized in that The diffusion model includes a U-Net module; the U-Net module includes multiple attention blocks; Each attention block consists of a self-attention layer, a reference adapter layer, and a cross-attention layer; In each attention block, the output of the self-attention layer is the input of the reference adapter layer; the output of the reference adapter layer is the input of the cross-attention layer; the output of the cross-attention layer is the input of the next attention block; The reference adapter layer is a multi-head cross-attention layer.

6. The method according to claim 5, characterized in that The output of the self-attention layer ; is the input of the attention block where the self-attention layer is located, is the text feature extracted from the text prompt, is the self-attention layer; The output of the reference adapter layer ,in, is the repaired feature, is the reference adapter layer; The output of the crisscross attention layer .

7. An image object restoration device, characterized in that: The device comprises: An acquisition module, used to acquire an image including a repair area, a text prompt and a binary mask code; the text prompt is used to describe a target object in the repair area; the binary mask code is used to indicate the area to be repaired; A determination module, used for determining a mask image of the image according to the binary mask code acquired by the acquisition module; An inference module, used to infer the restored features of the target object according to the text prompt acquired by the acquisition module and the mask image determined by the determination module; A generation module, used for using the restored features obtained by the inference module as visual cues to guide the diffusion model to generate the target object; The inference module is used to extract visual features from the mask image. ,in, is the feature vector, is the feature identifier, is the characteristic length, , is the feature dimension; according to the text prompt and , reconstruct semantic features Based on the semantic features , determine the restored features of the target object.

8. An electronic device, characterized in that: include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon; the computer program is executed by a processor to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image synthesis method and device, electronic equipment and storage medium

    CN117036184A