Text-guided zero-sample transparent layer and layered image generation method

Through text-guided zero-sample transparent layer and layered image generation method, the problem of poor control of image layout and subject object position in the prior art is solved, efficient and controllable image generation is achieved, and dependence on training data and computing resources is reduced.

CN120070638AActive Publication Date: 2025-05-30UNIV OF SCI & TECH OF CHINA

Patent Information

Application Number
CN202510202270.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-30
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing image generation method based on diffusion model lacks precise control of image layout and subject object position when processing complex scenes of multi-subject objects, and requires a large amount of training data and computing resources.

Method used

Using text-guided zero-sample transparent layer and layered image generation method, images with transparent background are generated through text encoder and transparent image encoder, and the position and layout of the subject object are accurately controlled using the cross attention matrix and denoising process.

Benefits of technology

Accurate control of image layout and subject object position is realized, reducing dependence on training data and computing resources, and improving the controllability and efficiency of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070638A_ABST
    Figure CN120070638A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image processing, in particular to a text-guided zero-sample transparent image layer and layered image generation method, and the layered image generation method comprises the steps: inputting a global image text prompt, a target image size and a layer text prompt to a foreground position information generation model, and obtaining foreground position information; generating a first target image for each layer of text prompt; generating a soft segmentation mask according to the transparent channels of all the first target images; superposing all the first target images, and encoding the first target images to a potential space to obtain foreground superposition potential features; gaussian noise is randomly sampled as an initial background potential feature; and according to the soft segmentation mask, mixing the foreground superposition potential feature and the initial background potential feature in an iterative denoising process to obtain a global image potential feature, and decoding the global image potential feature into a second target image. According to the method, the position of each main body object is accurately controlled, and the image layout capability of the model is enhanced; the step of model training is omitted, and computing resources are greatly saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a text-guided zero-shot transparent layer and hierarchical image generation method. Background Art

[0002] In computer image processing, an image usually consists of various channels, and the most common one is the RGBA channel, where R, G, and B represent the three colors red, green, and blue respectively, and A represents the transparency channel. An image can usually be divided into a background and a main object. For example, in a photo of an apple on a dining table, the apple is the main object, and the dining table and possibly other objects constitute the background. In practical applications, generating an image from a text description, known as image generation, is an important task in the cross-field of computer vision and natural language processing. The development of this technology enables people to generate images with specific content and layout from text prompts, which has far-reaching significance for various fields, including assisted design, data augmentation, and content creation.

[0003] In existing image generation methods, the method based on diffusion models is a mainstream solution. This type of image generation method generates an image that conforms to the text description through an iterative denoising process within a certain number of time steps. However, this method lacks precise control over the image layout and the position of the main object when dealing with complex scenes with multiple main objects; in addition, this method usually requires a large amount of training data and computing resources, which is an important limitation in many practical applications. Summary of the Invention

[0004] To solve the above problems, the present invention provides a text-guided zero-shot transparent layer and hierarchical image generation method.

[0005] The first aspect of the present invention discloses a text-guided zero-shot transparent layer generation method, which is used to generate a first target image with a size of the target image size according to the input layer text prompt, target image size, and foreground position information. The layer text prompt describes the content of the first target image. The first target image includes a main object and a transparent background. The foreground position information represents the position of the main object in the first target image. The method includes: Receiving the layer text prompt, target image size, and foreground position information; Inputting the layer text prompt into a text encoder to obtain a text prompt embedding; Generating a binary mask according to the foreground position information; Create a transparent image according to the target image size, and encode the transparent image into the latent space through a transparent image encoder to obtain the latent features of the transparent image; wherein, the sizes of the first target image, the binary mask, and the transparent image are the same; In the latent space, randomly sample Gaussian noise as the initial noise latent features; Iterate a preset first number of time steps. At each time step, according to the latent features of the transparent image and the binary mask, modify the noise latent features output in the previous time step to obtain the corrected noise latent features; set the cross-attention matrix according to the binary mask, and perform denoising on the corrected noise latent features by adjusting the cross-attention correlation between the text prompt embedding and the corrected noise latent features to obtain the noise latent features output in the current time step; wherein, in the first time step, the initial noise latent features are used as the noise latent features output in the previous time step; Input the noise latent features output in the last time step into the transparent image decoder to obtain the first target image.

[0006] Further, the step of modifying the noise latent features output in the previous time step according to the latent features of the transparent image and the binary mask to obtain the corrected noise latent features includes: Add the first image noise to the latent features of the transparent image to obtain the latent features of the transparent noisy background, and modify the noise latent features output in the previous time step according to the latent features of the transparent noisy background and the binary mask to obtain the corrected noise latent features; wherein, the first image noise is Gaussian noise, and the intensity of the first image noise added at each time step is less than the intensity of the first image noise added in the previous time step.

[0007] Further, the step of modifying the noise latent features output in the previous time step according to the latent features of the transparent noisy background and the binary mask to obtain the corrected noise latent features includes: Judge whether the number of iterations in the current time step exceeds a preset second number; wherein, the second number is less than the first number; When it does not exceed, based on the binary mask, determine the background region features in the noise latent features output in the previous time step, and replace the background region features with the latent features of the transparent noisy background to obtain the corrected noise latent features; When it exceeds, determine the noise latent features output in the previous time step as the corrected noise latent features.

[0008] Further, the step of setting the cross-attention matrix according to the binary mask and performing denoising on the corrected noise latent features by adjusting the cross-attention correlation between the text prompt embedding and the corrected noise latent features to obtain the noise latent features output in the current time step includes: Input the text prompt embedding and the corrected noise latent features into the U-shaped network; Embed the text prompt as the keys and values of the cross-attention layer of the U-Net, use the corrected noise latent feature as the query of the cross-attention layer of the U-Net, set the cross-attention matrix according to the binary mask, and perform denoising on the corrected noise latent feature to obtain the noise latent feature output at the current time step.

[0009] Further, the step of setting the cross-attention matrix according to the binary mask includes: Based on the binary mask, divide the pixel region corresponding to the corrected noise latent feature into a foreground region and a background region; Set the cross-attention matrix corresponding to the foreground region as: ; Set the cross-attention matrix corresponding to the background region as: ; where, represents the original cross-attention matrix corresponding to the foreground region, represents the original cross-attention matrix corresponding to the background region, represents the correction coefficient corresponding to the foreground region, represents the correction coefficient corresponding to the background region, represents the th text prompt embedding, represents the number of text prompt embeddings, represents the text prompt embedding corresponding to the preset text start token, represents the text prompt embedding corresponding to the preset text end token.

[0010] The second aspect of the present invention discloses a text-guided zero-shot hierarchical image generation method, which is used to generate a second target image according to a global image text prompt, a target image size, and at least one layer text prompt. The global image text prompt describes the content of the second target image, and each layer text prompt describes a main object in the second target image. The method includes: Input the global image text prompt, the target image size, and at least one layer text prompt into the foreground position information generation model to obtain the foreground position information corresponding to each layer text prompt; wherein, the foreground position information generation model is a pre-trained large language model; According to the target image size, for each layer text prompt and its corresponding foreground position information, use the text-guided zero-shot transparent layer generation method described in any item of the first aspect of the present invention to generate the first target image corresponding to this layer text prompt; Generate a soft segmentation mask according to the transparent channels in all the first target images; Overlay all the first target images to obtain a foreground-overlaid image; Encode the foreground-overlaid image into the latent space through an image encoder to obtain foreground-overlaid latent features; In the latent space, randomly sample Gaussian noise as the initial background latent features; According to the soft segmentation mask, mix the foreground-overlaid latent features with the initial background latent features to obtain global image latent features; Input the global image latent features into an image decoder to obtain the second target image.

[0011] Further, the step of mixing the foreground-overlaid latent features with the initial background latent features according to the soft segmentation mask to obtain global image latent features includes: Iterate a preset third number of time steps. At each time step, generate foreground-noisy latent features according to the foreground-overlaid latent features and Gaussian noise; according to the soft segmentation mask, mix the foreground-noisy latent features with the mixed image latent features output in the previous time step, and output the mixed image latent features at the current time step; wherein, in the first time step, the initial background latent features are used as the mixed image latent features output in the previous time step; Determine the mixed image latent features output in the last time step as the global image latent features.

[0012] Further, the step of mixing the foreground-noisy latent features with the mixed image latent features output in the previous time step according to the soft segmentation mask includes: The th time step is mixed according to the following formula: ; wherein, represents the mixed image latent features output in the th time step, represents the soft segmentation mask, represents the th time step of the foreground-noisy latent features.

[0013] Further, at each time step, the step of generating foreground-noisy latent features according to the foreground-overlaid latent features and Gaussian noise includes: At each time step, add second image noise to the foreground-overlaid latent features to obtain foreground-noisy latent features; wherein, the second image noise is Gaussian noise, and the intensity of the second image noise added at each time step is less than the intensity of the second image noise added in the previous time step.

[0014] Further, the step of generating a soft segmentation mask according to the transparency channels in all the first target images includes: Generate a soft segmentation mask according to the following formula : ; wherein, is the alpha channel of the th first target image, and is the number of first target images.

[0015] By introducing an alpha layer and foreground position information, the present invention can precisely control the position of each subject object, thereby enhancing the image layout ability of the model. The generated images not only conform to the semantic content of the text prompt but also meet the layout requirements of users. Moreover, the present invention adopts the idea of zero-shot learning, eliminating the model training step and greatly saving computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 is a schematic flowchart of a text-guided zero-shot alpha layer generation method disclosed in an embodiment of the present invention; Figure 2 is a schematic diagram of an alpha layer disclosed in an embodiment of the present invention; Figure 3 is another schematic diagram of an alpha layer disclosed in an embodiment of the present invention; Figure 4 is yet another schematic diagram of an alpha layer disclosed in an embodiment of the present invention; Figure 5 is a schematic flowchart of a text-guided zero-shot hierarchical image generation method disclosed in an embodiment of the present invention; Figure 6 is a schematic diagram of a hierarchical image disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] In the specification, claims, and above-mentioned drawings of the present invention, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, or product that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include unlisted steps or units, or may optionally further include other steps or units inherent to these processes, methods, devices, or products.

[0020] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0021] Please refer to Figure 1 as shown Figure 1 is a schematic flowchart of a text-guided zero-shot transparent layer generation method disclosed in an embodiment of the present invention. The zero-shot transparent layer generation method is used to generate a first target image with a size of the target image size according to the input layer text prompt, target image size, and foreground position information. The layer text prompt describes the content of the first target image, and the first target image includes a main object and a transparent background. Figure 2 , Figure 3 , Figure 4 respectively show different first target images, and their corresponding main objects are different zebras. The foreground position information represents the position of the main object in the first target image. As Figure 1 shown, the text-guided zero-shot transparent layer generation method may include the following operations: S101. Receive the layer text prompt, target image size, and foreground position information; In this optional embodiment, the layer text prompt refers to a text description of the main object in the first target image, describing the visual attributes of the main object, such as category, color, pose, state, etc. Taking Figure 2 as an example, its layer text prompt may include, but is not limited to, the following expression methods: A standing zebra; A realistic-style zebra with black and white stripes, facing right; A side view of a zebra.

[0022] The foreground position information refers to the spatial position and range of the subject object in the first target image, which is represented by a bounding box in the embodiments of the present invention. The bounding box is a rectangular box, and is described by the abscissa and ordinate of its upper left corner, the width of the bounding box, and the height of the bounding box. Taking Figure 2 as an example, assume Figure 2 the size of the corresponding first target image is , then the bounding box of its subject object can be represented as [60, 90, 380, 410]. This quadruple represents that the upper left corner coordinates of the bounding box of the subject object are (60, 90), the width is 380 pixels, and the height is 410 pixels. The numbers in this quadruple are only examples and do not represent real data.

[0023] Those of ordinary skill in the art can deduce that the width and height in the foreground position information are respectively smaller than the width and height in the target image size.

[0024] S102: Input the layer text prompt into the text encoder to obtain a text prompt embedding; In this optional embodiment, the text encoder is a model that converts text into a numerical vector representation, which can map an indefinite-length text sequence to a dense vector space of a fixed dimension. The text encoder in the embodiments of the present invention can select a contrastive language-image pre-training model (CLIP, Contrastive Language-Image Pre-training), a text-to-text transfer transformer (T5 model, Text-to-Text Transfer Transformer), etc., and the embodiments of the present invention do not make limitations.

[0025] S103: Generate a binary mask according to the foreground position information; In this optional embodiment, the binary mask is a two-dimensional matrix with the same size as the first target image, where each element corresponds to a pixel, and each element can only take two values of 0 or 1. The binary mask is used to represent the position and range of the subject object in the first target image. Specifically, the value of the binary mask being 1 indicates that the corresponding pixel belongs to the subject object, and the value being 0 indicates that the corresponding pixel belongs to the background.

[0026] S104: Create a transparent image according to the target image size, and encode the transparent image into the latent space through a transparent image encoder to obtain a transparent image latent feature; wherein, the first target image, the binary mask, and the transparent image have the same size; In this optional embodiment, a transparent image refers to an image with all transparent channel values being 0. A transparent image encoder is a neural network model that converts a transparent image into a low-dimensional latent representation. The sizes of the first target image, the binary mask, and the transparent image are all the size of the target image. In the embodiments of the present invention, the type of the transparent image encoder is a variational auto-encoder (VAE, Variational auto-encoder).

[0027] The latent space in this embodiment is a continuous low-dimensional vector space, and the space of the transparent image before encoding is the image space. For example, in the image space, the size of the transparent image is and the latent feature of the transparent image in the latent space is where 4 is the number of image channels, the resolution of the latent space is and the resolution of the image space is The resolution of the latent space is 1 / 8 of that of the image space.

[0028] S105. In the latent space, randomly sample Gaussian noise as the initial noise latent feature; In this optional embodiment, random sampling refers to the process of randomly drawing samples from a probability distribution. Random sampling can be implemented by generating pseudo-random numbers by a computer, such as inverse transform sampling, rejection sampling, importance sampling, etc.

[0029] In the present invention, sampling Gaussian noise in the latent space means that the size of the sampled Gaussian noise is the same as the resolution of the latent space. For example, when the resolution of the latent space is the size of the sampled Gaussian noise is where 4 is the number of image channels.

[0030] S106. Iterate a preset first number of time steps. At each time step, modify the noise latent feature output in the previous time step according to the latent feature of the transparent image and the binary mask to obtain a corrected noise latent feature; set a cross-attention matrix according to the binary mask, and perform denoising on the corrected noise latent feature by adjusting the cross-attention correlation between the text prompt embedding and the corrected noise latent feature to obtain the noise latent feature output at the current time step; wherein, in the first time step, the initial noise latent feature is used as the noise latent feature output in the previous time step; The time step in the embodiments of the present invention refers to a time interval, which is sorted according to the execution time.

[0031] The denoising in the embodiments of the present invention is performed by a sampler. In the embodiments of the present invention, a denoising diffusion implicit models sampler (DDIM, Denoising Diffusion Implicit Models Sampler) is used.

[0032] In an optional embodiment, the step of modifying the noise latent feature output in the previous time step according to the transparent image latent feature and the binary mask to obtain the corrected noise latent feature includes: Adding first image noise to the transparent image latent feature to obtain a transparent noise-added background latent feature, and modifying the noise latent feature output in the previous time step according to the transparent noise-added background latent feature and the binary mask to obtain the corrected noise latent feature; wherein, the first image noise is Gaussian noise, and the intensity of the first image noise added in each time step is less than the intensity of the first image noise added in the previous time step.

[0033] It can be seen that in this optional embodiment, by using first image noise with different intensities at each time step and gradually reducing the noise intensity as the time step progresses, a progressive denoising and optimization process can be achieved. In the early time steps, denoising the corrected noise latent feature with strong Gaussian noise is mainly responsible for creating structural features such as the outline of the main object; while in the later time steps, denoising the corrected noise latent feature with weak Gaussian noise is mainly responsible for refining the details of the main object, which can achieve local fine-tuning and optimization, making the finally generated transparent layer more delicate and accurate.

[0034] In another optional embodiment, the step of modifying the noise latent feature output in the previous time step according to the transparent noise-added background latent feature and the binary mask to obtain the corrected noise latent feature includes: Determining whether the number of iterations in the current time step exceeds a preset second quantity; wherein, the second quantity is less than the first quantity; When it does not exceed, based on the binary mask, determining the background region feature in the noise latent feature output in the previous time step, and replacing the background region feature with the transparent noise-added background latent feature to obtain the corrected noise latent feature; When it exceeds, determining the noise latent feature output in the previous time step as the corrected noise latent feature.

[0035] In an optional embodiment, the first quantity can be set to 800, and the second quantity can be set to 250.

[0036] It can be seen that this optional embodiment can effectively control the modification process of the noise latent feature. By replacing the background region feature in the early time steps, it can ensure that in the early time steps, the background region determined by the binary mask can generate a transparent background, thereby improving the quality of the generated image. In the later time steps, not modifying the background region feature before denoising the noise latent feature helps to maintain the stability of the generated image and reduce visual incoherence caused by unnecessary minor changes.

[0037] In yet another alternative embodiment, the steps of setting the cross-attention matrix according to the binary mask and performing denoising on the corrected noise latent feature by adjusting the cross-attention correlation between the text prompt embedding and the corrected noise latent feature to obtain the noise latent feature output at the current time step include: Input the text prompt embedding and the corrected noise latent feature into the U-shaped network; Use the text prompt embedding as the key and value of the cross-attention layer of the U-shaped network, use the corrected noise latent feature as the query of the cross-attention layer of the U-shaped network, set the cross-attention matrix according to the binary mask, and perform denoising on the corrected noise latent feature to obtain the noise latent feature output at the current time step.

[0038] In this alternative embodiment, the U-shaped network refers to the U-Net network.

[0039] It can be seen that this alternative embodiment effectively improves the accuracy of the generated image, maintains the balance between the global layout and local details, enhances the semantic consistency and visual quality of the generated image by introducing the U-shaped network and adjusting the cross-attention correlation between the text prompt embedding and the corrected noise latent feature, thereby significantly improving the performance and application value of the text-guided zero-shot transparent layer generation method.

[0040] In yet another alternative embodiment, the steps of setting the cross-attention matrix according to the binary mask include: Based on the binary mask, divide the pixel region corresponding to the corrected noise latent feature into a foreground region and a background region; Set the cross-attention matrix corresponding to the foreground region as: ; Set the cross-attention matrix corresponding to the background region as: ; where, represents the original cross-attention matrix corresponding to the foreground region, represents the original cross-attention matrix corresponding to the background region, represents the correction coefficient corresponding to the foreground region, represents the correction coefficient corresponding to the background region, represents the th text prompt embedding, represents the number of text prompt embeddings, represents the text prompt embedding corresponding to the preset text start token, represents the text prompt embedding corresponding to the preset text end token.

[0041] In this optional embodiment, based on the binary mask, the pixel region corresponding to the corrected noise latent feature is divided into a foreground region and a background region, which means that the pixel region with a binary mask value of 1 is divided into the foreground region, and the pixel region with a binary mask value of 0 is divided into the background region.

[0042] The text start marker is a preset marker, and the text prompt embedding corresponding to this marker is located at the starting position of all text prompt embeddings. In the embodiments of the present invention, [SoT] is used as the text start marker. The text end marker is a preset marker, and the text prompt embedding corresponding to this marker is located at the ending position of all text prompt embeddings. In the embodiments of the present invention, [EoT] is used as the text end marker.

[0043] The original cross-attention matrix is the initial attention distribution state of the U-shaped network before any correction or adjustment.

[0044] It can be seen that in this optional embodiment, based on the binary mask, the pixel region corresponding to the corrected noise latent feature is divided into a foreground region and a background region, which can more precisely enhance the correlation between the text prompt embedding and different regions of the image, thereby improving the quality of image generation; moreover, by using the newly set cross-attention matrix to denoise the corrected noise latent feature, the consistency between the image and the text can be effectively enhanced, important features can be strengthened, and thus an image that conforms to the text prompt and has richer details can be generated.

[0045] S107. Input the noise latent feature output at the last time step into the transparent image decoder to obtain the first target image.

[0046] In this optional embodiment, the transparent image decoder refers to a model that decodes the noise latent feature in the latent space into the first target image in the image space. The type of the transparent image decoder in the embodiments of the present invention is a variational auto-encoder (VAE).

[0047] It can be seen that in this optional embodiment, through the text guidance and zero-shot learning mechanism, combined with the foreground position information and the cross-attention adjustment technology, the ability to generate a high-precision transparent background image without training data is achieved, which not only ensures a high degree of consistency between the generated subject and the text description, but also precisely controls the object layout through the binary mask. At the same time, the transparent image encoder and the transparent image decoder effectively retain the background transparency information, significantly improving the controllability of image synthesis and the convenience of post-editing.

[0048] Please refer to Figure 5 as shown Figure 5It is a schematic flowchart of a text-guided zero-shot hierarchical image generation method disclosed in an embodiment of the present invention. The text-guided zero-shot hierarchical image generation method is used to generate a second target image according to a global image text prompt, a target image size, and at least one layer text prompt. The global image text prompt describes the content of the second target image, and each layer text prompt describes a main object in the second target image. Figure 6 It shows a second target image, in which there are three main objects, namely three zebras. As Figure 5 shown, the text-guided zero-shot hierarchical image generation method may include the following operations: S501. Input the global image text prompt, the target image size, and at least one layer text prompt into a foreground position information generation model to obtain the foreground position information corresponding to each layer text prompt; wherein, the foreground position information generation model is a pre-trained large language model; In an optional embodiment, before the step of inputting the global image text prompt, the target image size, and at least one layer text prompt into the foreground position information generation model, the method further includes: Input a preset number of foreground position information generation examples into the large language model. Each foreground position information generation example includes an example global image text prompt, an example target image size, at least one example layer text prompt, and example foreground position information corresponding to each example layer text prompt. The example global image text prompt describes the image content in an example image, the example target image size represents the size of the example image, the example image includes at least one main object, each example layer text prompt describes a main object, and each example foreground position information describes the position information of a main object in the example image.

[0049] In an optional embodiment, a possible foreground position information generation example is as follows: Example global image text prompt: A realistic image of a brown tabby cat on a white staircase; Example target image size: ; The first example layer text prompt: A brown tabby cat; The second example layer text prompt: A glass bottle; The first example foreground position information: [88, 28, 355, 255]; The second example foreground position information: [10, 317, 102, 152].

[0050] Among them, [88, 28, 355, 255] indicates that the abscissa of the upper left corner of the bounding box of the brown tabby cat is 88, the ordinate of the upper left corner is 28, the width is 355, and the height is 255; [10, 317, 102, 152] indicates that the abscissa of the upper left corner of the bounding box of the glass bottle is 10, the ordinate of the upper left corner is 317, the width is 102, and the height is 152.

[0051] It can be seen that in this alternative embodiment, by using a preset number of foreground position information to generate example pairs for context learning of the large language model, the large language model can activate the pre-trained capabilities through designed prompts during the inference stage, dynamically establish the mapping relationship between text prompts and foreground position information, without manually specifying the position information of each subject object, greatly improving the efficiency of generating foreground position information and reducing manual intervention and labor costs. At the same time, the pre-trained large language model has good generalization ability and can automatically generate reasonable foreground position information according to different text prompts, enhancing the practicality of the present invention.

[0052] S502. According to the target image size, for each layer of text prompt and its corresponding foreground position information, use the text-guided zero-shot transparent layer generation method described in any of the foregoing embodiments to generate the first target image corresponding to this layer of text prompt; In this alternative embodiment, using the target image size, each layer of text prompt and its corresponding foreground position information as the input of the text-guided zero-shot transparent layer generation method, the first target image corresponding to this layer of text prompt is obtained, where the size of the first target image is the target image size.

[0053] S503. Generate a soft segmentation mask according to the transparency channels in all the first target images; In an alternative embodiment, the step of generating a soft segmentation mask according to the transparency channels in all the first target images includes: Generate a soft segmentation mask according to the following formula : ; where, is the transparency channel of the th first target image, and is the number of first target images.

[0054] It can be seen that this alternative embodiment uses a linear weighted formula to explicitly define the mixing process, precisely controls the contribution weights of the foreground and background through the soft segmentation mask, ensures that the mixing process is interpretable and computationally efficient. The soft segmentation mask in the linear weighted formula is directly related to the transparency channel information, enabling position-sensitive mixing with pixel-level accuracy, avoiding the edge jaggedness problem caused by traditional hard masks, and at the same time supporting the natural transition of multi-level transparency, improving the quality of image generation.

[0055] S504. Superimpose all the first target images to obtain a foreground superimposed image; In this alternative embodiment, all the first target images are superimposed according to the pixel positions to obtain a foreground superimposed image, and the size of the foreground superimposed image is the size of the target image.

[0056] S505. Encode the foreground superimposed image into the latent space through an image encoder to obtain foreground superimposed latent features; In this alternative embodiment, the image encoder refers to a neural network model that converts an image into a low-dimensional latent representation. It can extract high-level semantic features of the image and map them to the latent space. The type of image encoder in the embodiments of the present invention is a variational auto-encoder (VAE, Variational auto-encoder). The latent space in this embodiment has the same resolution as the latent space in any of the embodiments of the foregoing text-guided zero-shot transparent layer generation method.

[0057] S506. Randomly sample Gaussian noise as the initial background latent features in the latent space; S507. Mix the foreground superimposed latent features and the initial background latent features according to the soft segmentation mask to obtain global image latent features; In another alternative embodiment, the step of mixing the foreground superimposed latent features and the initial background latent features according to the soft segmentation mask to obtain global image latent features includes: Iterate a preset third number of time steps. At each time step, generate foreground-noisy latent features according to the foreground superimposed latent features and Gaussian noise; mix the foreground-noisy latent features with the mixed image latent features output in the previous time step according to the soft segmentation mask, and output the mixed image latent features of the current time step; wherein, in the first time step, the initial background latent features are used as the mixed image latent features output in the previous time step; Determine the mixed image latent features output in the last time step as the global image latent features.

[0058] It can be seen that in this alternative embodiment, through the progressive mixing strategy of iterating time steps, the foreground and background are gradually fused in the latent space, effectively alleviating the problems of blurred boundaries of the main object or residual noise caused by direct superimposition, making the generation process more adaptable to the denoising characteristics of the diffusion model, improving the physical rationality and visual coherence of the generated image, retaining the integrity of the foreground structure and details, and further improving the quality of image generation.

[0059] In yet another alternative embodiment, the step of mixing the foreground-noisy latent features with the mixed image latent features output in the previous time step according to the soft segmentation mask includes: The At each time step, mixing is performed according to the following formula: ; wherein, represents the latent features of the mixed image output at the -th time step, represents the soft segmentation mask, represents the foreground noise-added latent features at the -th time step.

[0060] It can be seen that in this alternative embodiment, the soft segmentation mask is directly generated by accumulating the transparency channels, converting the multi-layer transparency information into a continuous spatial weight distribution, which not only preserves the characteristics of each independent subject object, but also quantifies the cumulative influence of the overlapping regions through superposition; on the other hand, this embodiment does not require additional training of a segmentation model, and uses the inherent attributes of the transparency channels to implement mask generation, improving the efficiency of generating pictures.

[0061] In yet another alternative embodiment, at each time step, the steps of generating the foreground noise-added latent features according to the foreground superimposed latent features and Gaussian noise include: At each time step, adding second image noise to the foreground superimposed latent features to obtain the foreground noise-added latent features; wherein, the second image noise is Gaussian noise, and the intensity of the second image noise added at each time step is less than the intensity of the second image noise added at the previous time step.

[0062] It can be seen that in this alternative embodiment, through the design of the noise intensity decreasing with time steps, the standard denoising process of the diffusion model is simulated during the foreground noise-adding stage. The strong noise in the early stage retains the structural freedom, and the weak noise in the later stage focuses on detail optimization, making the generated foreground features have high authenticity and fidelity.

[0063] S508. Input the global image latent features into an image decoder to obtain a second target image.

[0064] In this alternative embodiment, the image decoder refers to a neural network model that converts a low-dimensional latent representation into high-dimensional image data. The type of the image decoder in this embodiment is a variational auto-encoder (VAE, Variational auto-encoder).

[0065] In summary, a text-guided zero-shot transparent layer and hierarchical image generation method disclosed by the present invention adopts a hierarchical image generation paradigm. First, an RGBA image with transparency information is generated as a separate instance of the main object to enhance the fine-grained control of the attributes of the main object. To generate a complex hierarchical image corresponding to the text prompt, the generated main object layers are combined and repaired to obtain a complete image containing each main object. At the same time, to achieve the editability of the layout of the main object in the image, in the RGBA image generation stage, the present invention uses a layout-aware cross-attention technique to introduce layout conditions by a training-free method, guiding the RGBA main object to follow the layout information for generation, realizing the automatic understanding of complex layouts in the image generation process and enhancing the controllability of image generation. Therefore, the present invention effectively overcomes various drawbacks in the prior art and has high industrial utilization value.

[0066] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A text-guided zero-sample transparent layer generation method, characterized in that: The method is used to generate a first target image with a size of the target image according to an input layer text prompt, a target image size and foreground position information, wherein the layer text prompt describes the content of the first target image, the first target image includes a main object and a transparent background, and the foreground position information indicates the position of the main object in the first target image, and the method comprises: Receive layer text prompt, target image size and foreground position information; Input the layer text prompt into the text encoder to obtain the text prompt embedding; Generate a binary mask based on the foreground position information; Creating a transparent image according to the size of the target image, and encoding the transparent image into a latent space through a transparent image encoder to obtain latent features of the transparent image; wherein the first target image, the binary mask, and the transparent image have the same size; In the latent space, randomly sampled Gaussian noise as the initial noise latent features; Iterate a preset first number of time steps, and at each time step, modify the noise latent feature outputted at the previous time step according to the transparent image latent feature and the binary mask to obtain the modified noise latent feature; set the cross attention matrix according to the binary mask, perform denoising on the modified noise latent feature by adjusting the cross attention correlation between the text prompt embedding and the modified noise latent feature, and obtain the noise latent feature outputted at the current time step; wherein, in the first time step, the initial noise latent feature is used as the noise latent feature outputted at the previous time step; The noisy latent features output at the last time step are input into the transparent image decoder to obtain the first target image.

2. The text-guided zero-sample transparent layer generation method according to claim 1, characterized in that: According to the transparent image potential features and the binary mask, the noise potential features outputted in the previous time step are modified to obtain the modified noise potential features, and the steps include: The first image noise is added to the transparent image latent feature to obtain the transparent noisy background latent feature, and the noise latent feature output at the previous time step is modified according to the transparent noisy background latent feature and the binary mask to obtain the modified noise latent feature; wherein the first image noise is Gaussian noise, and the intensity of the first image noise added at each time step is less than the intensity of the first image noise added at the previous time step.

3. The text-guided zero-sample transparent layer generation method according to claim 2, characterized in that: According to the transparent noise background potential features and binary mask, the noise potential features output in the previous time step are modified to obtain The steps to correct for the underlying characteristics of the noise include: Determine whether the number of iterations of the current time step exceeds a preset second number; wherein the second number is less than the first number; When it does not exceed, the background area features in the noise potential features output in the previous time step are determined based on the binary mask, and the background area features are replaced with the transparent noise background potential features to obtain the corrected noise potential features; When it exceeds, the noise latent feature outputted at the previous time step is determined to be the corrected noise latent feature.

4. The text-guided zero-sample transparent layer generation method according to claim 1, characterized in that: The steps of setting a cross attention matrix according to a binary mask, performing denoising on the modified noise latent feature by adjusting the cross attention correlation between the text prompt embedding and the modified noise latent feature, and obtaining the noise latent feature outputted at the current time step include: The text prompt embedding and corrected noise latent features are fed into the U-shaped network; The text prompt embedding is used as the key and value of the cross attention layer of the U-type network, the corrected noisy latent feature is used as the query of the cross attention layer of the U-type network, the cross attention matrix is ​​set according to the binary mask, the corrected noisy latent feature is denoised, and the noisy latent feature output at the current time step is obtained.

5. The text-guided zero-sample transparent layer generation method according to claim 4, characterized in that: The steps of setting the criss-cross attention matrix according to the binary mask include: Based on the binary mask, the pixel area corresponding to the latent feature of the corrected noise is divided into the foreground area and the background area; Set the cross attention matrix corresponding to the foreground area for: ; Set the cross attention matrix corresponding to the background area for: ; in, represents the original cross attention matrix corresponding to the foreground area, represents the original cross attention matrix corresponding to the background area, represents the correction coefficient corresponding to the foreground area, represents the correction coefficient corresponding to the background area, Indicates A text hint is embedded, represents the number of text hint embeddings, Indicates the text hint embedding corresponding to the preset text start tag. Indicates the text hint embedding corresponding to the preset text end tag.

6. A text-guided zero-shot hierarchical image generation method, characterized in that: The method is used to generate a second target image according to a global image text prompt, a target image size and at least one layer text prompt, wherein the global image text prompt describes the content of the second target image, and each layer text prompt describes a main object in the second target image, and the method comprises: Inputting the global image text prompt, the target image size and at least one layer text prompt into a foreground position information generation model to obtain foreground position information corresponding to each layer text prompt; wherein the foreground position information generation model is a pre-trained large language model; According to the target image size, for each layer text prompt and its corresponding foreground position information, using the text-guided zero-sample transparent layer generation method as described in any one of claims 1 to 5, a first target image corresponding to the layer text prompt is generated; Generate a soft segmentation mask based on all transparent channels in the first target image; Superimposing all first target images to obtain a foreground superimposed image; The foreground overlay image is encoded into the latent space through the image encoder to obtain the foreground overlay latent features; In the latent space, randomly sampled Gaussian noise as the initial background latent features; According to the soft segmentation mask, the foreground superposition latent features are mixed with the initial background latent features to obtain the global image latent features; The global image latent features are input into the image decoder to obtain a second target image.

7. The text-guided zero-shot hierarchical image generation method according to claim 6, characterized in that: The steps of mixing the foreground superposition latent features and the initial background latent features according to the soft segmentation mask to obtain the global image latent features include: Iterate a preset third number of time steps, and at each time step, generate a foreground noise-added latent feature according to the foreground superimposed latent feature and Gaussian noise; according to the soft segmentation mask, mix the foreground noise-added latent feature with the mixed image latent feature output at the previous time step, and output the mixed image latent feature at the current time step; wherein, in the first time step, the initial background latent feature is used as the mixed image latent feature output at the previous time step; The mixed image latent features output at the last time step are determined as the global image latent features.

8. The text-guided zero-sample layered image generation method according to claim 7, characterized in that: The steps of mixing the foreground noise latent features with the mixed image latent features output at the previous time step according to the soft segmentation mask include: No. The time steps are mixed according to the following formula: ; in, Representative The mixed image potential features output at time steps are represents the soft segmentation mask, Representative The foreground noisy latent features at time steps.

9. The text-guided zero-shot hierarchical image generation method according to claim 7, characterized in that: At each time step, the steps of generating the foreground plus noise latent features according to the foreground superimposed latent features and Gaussian noise include: At each time step, the second image noise is added to the foreground superimposed latent feature to obtain the foreground noisy latent feature; wherein the second image noise is Gaussian noise, and the intensity of the second image noise added at each time step is less than the intensity of the second image noise added at the previous time step.

10. The text-guided zero-shot hierarchical image generation method according to claim 6, characterized in that: The step of generating a soft segmentation mask according to the transparent channels in all the first target images comprises: The soft segmentation mask is generated according to the following formula : ; in, For the The transparent channel of the first target image, is the number of the first target image.

Citation Information

Patent Citations

  • Context-aware reference image segmentation method, system and device and storage medium

    CN117078942A

  • Image fine-grained editing method and system based on text graph large model

    CN117808926A

  • Method and device for converting text data into image data

    CN118071867A

  • Video generation method and device, electronic equipment and storage medium

    CN118608907A

  • Multi-instance controllable image generation method based on cross attention redistribution

    CN118628611A

Cited By

  • Image enhancement method and device, equipment and storage medium

    CN120876316A

  • Image enhancement method, device, apparatus, and storage medium

    CN120876316B

  • Cross attention modulated digital printing pattern color semantic consistency generation method

    CN121685747A

  • Foreground-background soft separation zero sample anomaly detection method based on CLIP model

    CN121765610A