A training-free layout-controllable pattern generation method

CN122820908APending Publication Date: 2026-09-25ZHEJIANG HAIYIN DIGITAL TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611313886.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,单纯依靠文本提示词难以精确控制图像中元素的空间分布、疏密程度与前景背景的区域划分等关键布局信息,导致生成结果在布局上不稳定,难以满足图案设计对精细布局控制的需求

Benefits of technology

[0044]本方案无需对基础扩散模型进行额外微调,无需新增控制网络与额外参数,仅在推理阶段通过流程改造即可实现布局控制,无需大规模数据集与训练算力,部署灵活、适配性强,可直接对接各类主流预训练扩散模型;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820908A_ABST
    Figure CN122820908A_ABST
Patent Text Reader

Abstract

The application discloses a kind of training-free layout controllable pattern generation method, is realized based on pre-training text image diffusion model, whole process does not need model fine-tuning, complete process includes five key steps of S1 prompt word decomposition, S2 gray layout mask generation, S3 background prior injection purification, S4 early multi-branch attention fusion denoising, S5 late global fine-tuning decoding.The present application divides the image space into foreground region and background region by gray layout mask or point prompt and other light space prompts, applies different text conditions and denoising strategies to different regions during diffusion reasoning process, thereby realizing controllable generation of foreground element position and density, and obtaining a cleaner background through background purification mechanism, ensuring that the background purity and global style consistency of the generated pattern are consistent, without model fine-tuning, reducing interaction and computing cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image generation technology, and more specifically, to a training-free, layout-controllable pattern generation method. Background Technology

[0002] With the development of diffusion models, the quality of text-to-image generation has significantly improved, and these models are widely used in image content generation scenarios with repetitive or sparse elements, such as pattern design, clothing fabric patterns, decorative textures, and sky clouds. However, relying solely on text prompts makes it difficult to accurately control key layout information such as the spatial distribution, density, and foreground / background region division of elements in an image. This results in unstable layouts in the generated results, failing to meet the fine-grained layout control requirements of pattern design. Furthermore, in pattern generation tasks, the background is usually expected to remain a solid color or have a simple texture, but diffusion models tend to generate unnecessary subjects or complex textures in blank or simple background areas, affecting the overall visual quality.

[0003] One type of existing method relies on additional training or additional networks to achieve spatial control, but this often requires the introduction of additional parameters, relies on large-scale paired datasets and high computational costs for model fine-tuning, thus limiting its practicality. Another type of method attempts to achieve layout guidance through cue word engineering or post-processing constraints, but it is still insufficient to ensure the stability of complex spatial layouts, density control and background cleanliness, and it is easy to destroy the semantic shape of the object itself and reduce its recognizability.

[0004] Therefore, there is an urgent need for a pattern and texture generation scheme that requires no additional training, has low interaction costs, and can stably control the layout of foreground elements and the purity of the background. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a training-free, layout-controllable pattern generation method. By using lightweight spatial cues such as grayscale layout masks or dot hints, the image space is divided into foreground and background regions. During the diffusion inference process, different text conditions and denoising strategies are applied to different regions, thereby achieving controllable generation of the position and density of foreground elements. A cleaner background is obtained through a background purification mechanism, ensuring the background purity and global style consistency of the generated pattern. No model fine-tuning is required, reducing interaction and computation costs.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A training-free, layout-controllable pattern generation method includes the following steps:

[0008] S1: Get the overall prompt words , overall prompt words Decomposed into foreground cue words With background cue words ;

[0009] S2: Obtain spatial prompt information containing spatial layout information and generate a grayscale layout mask M;

[0010] S3: Generate background image The low-frequency prior is injected into the initial diffuse noise to suppress the appearance of redundant subjects in the background region;

[0011] S4: In the early stages of diffusion reasoning, multiple foreground cues are used respectively. With background cue words Generate key-value pairs, calculate multi-path attention outputs and perform spatial weighted fusion according to the grayscale layout mask to achieve semantic decoupling of regions, and complete the corresponding step of denoising update based on the fusion result;

[0012] S5: In the later stages of diffusion reasoning, switch back to the overall clue. A unified revision will be made.

[0013] Further, in step S2, the grayscale layout mask ,in, The target resolution for spatial layout information is defined as follows: a region with a pixel value of 1 corresponds to the distribution area of ​​foreground elements, a region with a pixel value of 0 corresponds to the background area, and a region with a pixel value between 0 and 1 is a grayscale transition region. The grayscale transition region is used to provide continuous blending weights during fusion to smooth the region boundaries.

[0014] Furthermore, in step S2, the grayscale layout mask M has multiple foreground branches, forming a corresponding mask set. ,in, Represents the i-th foreground region. This indicates that it does not belong to the i-th foreground branch;

[0015] The background area consists of residual weights. Implicit representation, the calculation formula is:

[0016]

[0017] Where i represents the index of the foreground branch, and N represents the total number of foreground branches.

[0018] Further, in step S3, the background image The generation method is as follows: based on background prompt words Parse the background color parameter or the user-specified background color parameter to generate a solid color background image with the same resolution as the target resolution. ;

[0019] Background image The method for low-frequency prior and injection of diffuse initial noise is as follows:

[0020] Will The input is fed into a variational autoencoder to obtain solid color latent variables, and then averaged across the spatial dimensions to obtain the channel-level background vector. Preserve low-frequency color priors and remove high-frequency spatial variations;

[0021] Sampling yields diffused noise And perform additive injection of background vectors in the background region.

[0022] Furthermore, the calculation formula for the additive injection is as follows:

[0023]

[0024] in, For injection intensity hyperparameters, This is an element-wise multiplication operation.

[0025] Furthermore, the early stage of the diffusion inference is a denoising time step that satisfies... In the later stage of diffusion inference, the denoising time step satisfies... The stage in which, For the noise reduction time step, This is the preset cutoff time threshold for multi-branch denoising.

[0026] Furthermore, in step S4, the steps for fusing the multi-path attention output with spatial weighting are as follows:

[0027] In the cross-attention layer of the U-Net denoising loop, spatial query features Q, with dimension d, are generated from the current latent variable features; the current background branch prompts are also included. Text embeddings are obtained by text encoders, and key-value pairs are obtained by projection. ;

[0028] Calculate multi-path attention output :

[0029]

[0030]

[0031] The standard formula for calculating attention is:

[0032]

[0033] The grayscale layout mask M is downsampled to match the feature resolution of the current attention layer, and then spatially weighted and fused to obtain the hybrid output. :

[0034]

[0035] The fused attention output is used to complete this step of U-Net noise prediction and denoising update.

[0036] Furthermore, in step S5, the denoising process ends to obtain the final latent variables. Afterwards, The input variational autoencoder decoder obtains the final output image.

[0037] Furthermore, in step S2, the method for generating the grayscale layout mask M includes at least three methods: direct drawing, coordinate-guided generation, and density-guided generation.

[0038] Furthermore, the density-guided generation process is as follows:

[0039] User-specified layout constraint parameters include foreground coverage. Element radius range, minimum interval With the maximum number of points K, on ​​the canvas Initial points are generated by random sampling. ;

[0040] The interval between all existing points is greater than [a certain value]. Under the given conditions, candidate points are randomly sampled from the canvas;

[0041] If the current number of points is less than K and the current foreground coverage has not been reached. If the candidate point is not found, it is accepted and added to the point set; this process is repeated until the coverage rate is met or the maximum number of points is reached.

[0042] Each point is rendered as a circular region with its own radius; and normalized to the grayscale range of [0,1] to obtain a grayscale layout mask.

[0043] In summary, the present invention has the following beneficial effects:

[0044] This solution requires no additional fine-tuning of the basic diffusion model, no addition of control networks or extra parameters. Layout control can be achieved simply by modifying the process during the inference stage. It does not require large-scale datasets or training computing power, and is flexible in deployment and highly adaptable. It can directly connect to various mainstream pre-trained diffusion models.

[0045] In this solution, users only need to draw, input coordinates or specify density parameters to generate control masks, resulting in low interaction costs. It also supports multi-dimensional layout control such as point position, element density, and multiple foreground partitions, which can meet diverse pattern design needs.

[0046] By using spatial weighted fusion of cross-attention layers to achieve semantic decoupling of regions, under the premise of strictly constraining the element layout, the semantic shape and visual recognizability of the foreground object itself will not be destroyed, and the generated elements will have a natural and reasonable shape without obvious deformation or semantic distortion.

[0047] In this scheme, the transition area of ​​the grayscale mask achieves a smooth blending of foreground and background semantics, avoiding boundary seam artifacts; in the later global refinement stage, the overall tone and texture style are further unified, regional segmentation traces are eliminated, and the visual integrity of the generated image is ensured. Attached Figure Description

[0048] Figure 1 This is a flowchart illustrating a training-free, layout-controllable pattern generation method according to this embodiment.

[0049] Figure 2 This is a flowchart illustrating steps S1 to S3 of this embodiment;

[0050] Figure 3 This is a flowchart illustrating step S1 of this embodiment;

[0051] Figure 4 This is a flowchart illustrating step S2 of this embodiment;

[0052] Figure 5 This is a flowchart illustrating step S3 of this embodiment;

[0053] Figure 6 This is a flowchart illustrating steps S4 and S5 of this embodiment;

[0054] Figure 7 This is a flowchart illustrating step S4 of this embodiment;

[0055] Figure 8 This is a flowchart illustrating step S5 of this embodiment. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] This embodiment discloses a training-free, layout-controllable pattern generation method, referring to... Figures 1-8As shown, the implementation is based on a pre-trained text-based image diffusion model, requiring no model fine-tuning throughout the process. The complete workflow includes five core steps: S1 cue word decomposition, S2 grayscale layout mask generation, S3 background prior injection and purification, S4 early multi-branch attention fusion denoising, and S5 late-stage global fine-tuning decoding. The specific implementation process is as follows:

[0058] Step S1: Semantic decomposition of prompt words;

[0059] Reference Figures 1-3 As shown, the overall prompt words input by the user are obtained. For example, "A clean white background dotted with red flowers and green leaves." The overall prompt is analyzed using pre-defined semantic parsing rules or a large language model. Break it down into foreground cue words With background cue words Foreground prompts The main elements in the corresponding image, in this embodiment, are "red flowers and green leaves"; background prompts The background area of ​​the corresponding image is, in this embodiment, a "clean white background".

[0060] For scenarios containing multiple types of foreground elements and requiring independent layout control, they can be broken down into multiple sets of foreground cue words. For example, it can be split into two sets of foreground prompts, namely "red flowers" and "green leaves", corresponding to two foreground branches, to achieve partitioned layout control of different elements.

[0061] Step S2: Generate grayscale layout mask;

[0062] Reference Figure 1 , Figure 2 , Figure 4 As shown, the system obtains the spatial prompts provided by the user and generates a grayscale layout mask M with the same size as the target generated image. ,in, H represents the target resolution for spatial layout information, where H is the image height and W is the image width.

[0063] Regions with a pixel value of 1 correspond to foreground element distribution areas. The closer the pixel value is to 1, the higher the weight of the region as a foreground element distribution area. Regions with a pixel value of 0 correspond to background areas. The closer the pixel value is to 0, the higher the weight of the region as a background area. Pixel values ​​between 0 and 1 are grayscale transition areas. Grayscale transition areas are used to provide continuous blending weights during fusion and to smooth the region boundaries.

[0064] Reference Figure 4 As shown, this embodiment supports three mask generation methods, which users can choose as needed:

[0065] Directly drawing the mask: Users draw black dots or blocks on a white canvas using freehand tools. After drawing, the pixel values ​​of the image are inverted and normalized from [0,255] to the [0,1] range to obtain a grayscale layout mask. This method is intuitive and suitable for custom irregular layout requirements.

[0066] Coordinate-guided generation: The user inputs a set of coordinates for a group of points, along with the rendering radius parameter for each point; the algorithm renders a corresponding circular grayscale area on the canvas based on the coordinates and radius, feathering the edges to form a transition band, and finally generating a grayscale layout mask. This method is suitable for precise layout control of points.

[0067] Density-guided generation: Users do not need to manually draw points; they only need to specify layout constraint parameters. The algorithm automatically generates a set of points that meet the constraints through random sampling, and then renders the mask.

[0068] The specific execution process of density-guided generation is as follows:

[0069] Preset foreground coverage Element radius range, minimum interval Layout constraint parameters such as maximum number of points K;

[0070] On the canvas Random sampling within the range generates initial points Add to the set of points;

[0071] Randomly sample candidate points from the canvas and determine whether the interval between each candidate point and all existing points is greater than the minimum interval. ;

[0072] If the interval condition is met, and the current number of points is less than K, and the current foreground coverage has not reached the target, then... If so, then the candidate point is accepted and added to the point set;

[0073] Repeat the above sampling and judgment process until the foreground coverage reaches [a certain level]. Sampling stops when the number of points reaches K.

[0074] Each point in the point set is rendered as a circular area with the corresponding radius, and the edges are Gaussian blurred to form a grayscale transition band. The grayscale layout mask is obtained by normalizing the pixel values.

[0075] For scenarios with multiple foreground branches, each foreground branch generates an independent sub-mask. This forms a corresponding mask set. , Where i represents the index of the foreground branch and N represents the total number of foreground branches;

[0076] Mask weight of background area Calculated using the following formula:

[0077]

[0078] The weights of all regions in the entire image are guaranteed to sum to 1, thus achieving an implicit representation of the background weights.

[0079] Step S3: Background low-frequency prior injection purification;

[0080] Reference Figure 5 As shown, to suppress the generation of redundant main elements or cluttered textures in the background area, this step injects low-frequency background color priors into the initial diffusion noise, constraining the content of the background area from the generation source. The specific process is as follows:

[0081] The background primary color is parsed from the background prompt words obtained in step S1. In this embodiment, the primary color is obtained as white from "clean white background"; it also supports users to directly specify RGB color values ​​as background colors.

[0082] Based on the obtained background color, generate a solid color background image with the same resolution as the target output image. .

[0083] solid color background image Input the variational autoencoder (VAE) encoder corresponding to the pre-trained diffusion model to obtain the corresponding solid color latent variables.

[0084] Perform global average pooling on the solid color latent variables in the spatial dimension to obtain the channel-level background vector in the channel dimension. This operation removes high-frequency spatial details of latent variables, retaining only the low-frequency prior of global color, thus avoiding disruption of background smoothness after injection.

[0085] The initial noise obtained by sampling conforms to a standard normal distribution. Perform additive injection to update the initial noise:

[0086]

[0087] in, This is used to inject intensity hyperparameters; for example, in this embodiment, the value is 0.05, which can be adjusted in the range of 0 to 1 as needed: the larger the value, the stronger the background purity, but the easier it is to lose the style performance of the original model; the smaller the value, the better the model style is preserved, but the improvement in background purity is limited. This indicates an element-wise multiplication operation, meaning that only the background prior is injected into the background region, while the foreground region remains unaffected.

[0088] Step S4: Multi-branch attention fusion and denoising in the early stage of diffusion inference;

[0089] Reference Figure 6, Figure 7 As shown, this step is performed in the early stages of diffusion inference, i.e., the denoising time step satisfies... The stage in which In this embodiment, the preset multi-branch denoising cutoff time threshold is used. The threshold is set to 60% of the total denoising steps. Users can adjust this threshold according to their layout constraint requirements: the larger the threshold, the stronger the layout constraint, but the weaker the global consistency; the smaller the threshold, the more uniform the global style, but the weaker the layout constraint.

[0090] The core of this step is to achieve semantic decoupling of regions in the cross-attention layer of U-Net, and to achieve layout constraints through mask weighted fusion. The specific process is as follows:

[0091] In the cross-attention layer of each U-Net denoising loop, the spatial query feature Q is generated by projecting the latent variable features of the current input. The feature dimension is d, and this feature carries spatial location information and remains unchanged in this step.

[0092] Each set of foreground and background cue words is input into a pre-trained text encoder to obtain the corresponding text embedding vectors, which are then mapped through a projection layer to obtain key-value pairs. .

[0093] Based on the standard attention formula, the attention output for each foreground branch and background branch is calculated separately: Attention output for the i-th foreground branch:

[0094]

[0095] Attention output for the background branch:

[0096]

[0097] The standard formula for calculating attention is:

[0098]

[0099] The grayscale layout mask M is downsampled to a resolution consistent with the current attention layer features using bilinear interpolation, thus aligning the mask with the attention feature space.

[0100] Using the mask pixel values ​​as weights, perform pixel-wise weighted fusion on each attention output to obtain a hybrid attention output:

[0101]

[0102] In a pure foreground region with a mask value of 1, the attention output corresponds exactly to the semantics of the foreground cue words; in a pure background region with a mask value of 0, the attention output corresponds exactly to the semantics of the background cue words; in grayscale transition regions, the two semantics are smoothly mixed according to weights to avoid abrupt seams and artifacts at the boundaries between the foreground and background.

[0103] Use the fused hybrid attention output Complete the subsequent U-Net computation, obtain the noise prediction result of the current step, and perform denoising update of the latent variables.

[0104] Repeat the above denoising process until the denoising time step reaches... This concludes the early stage of diffusion reasoning.

[0105] Step S5: Global refinement and decoding output in the later stage of diffusion inference;

[0106] Reference Figure 6 , Figure 8 As shown, when the denoising time step satisfies At this point, the process enters the later stage of diffusion reasoning. In this stage, multi-branch attention fusion ceases, and the use of holistic cue words is switched. As a condition, the standard diffusion denoising process is executed until all denoising steps are completed.

[0107] The purpose of this stage is to unify the global tone, lighting, and texture style of foreground elements and background areas, eliminate problems such as boundary breaks, inconsistent tones, and unnatural shapes that may be caused by early area decoupling, and improve the overall visual consistency and detail quality of the image without significantly disrupting layout adherence.

[0108] After the denoising process is completed, the final latent variables are obtained. The image is then input into a variational autoencoder (VAE) decoder for upsampling and decoding to obtain the final RGB format generated image.

[0109] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A training-free, layout-controllable pattern generation method, characterized in that, Includes the following steps: S1: Get the overall prompt words , overall prompt words Decomposed into foreground cue words With background cue words ; S2: Obtain spatial prompt information containing spatial layout information and generate a grayscale layout mask M; S3: Generate background image The low-frequency prior is injected into the initial diffuse noise to suppress the appearance of redundant subjects in the background region; S4: In the early stages of diffusion reasoning, multiple foreground cues are used respectively. With background cue words Generate key-value pairs, calculate multi-path attention outputs and perform spatial weighted fusion according to the grayscale layout mask to achieve semantic decoupling of regions, and complete the corresponding step of denoising update based on the fusion result; S5: In the later stages of diffusion reasoning, switch back to the overall clue. A unified revision will be made.

2. The training-free, layout-controllable pattern generation method according to claim 1, characterized in that, In step S2, the grayscale layout mask ,in, The target resolution for spatial layout information is defined as follows: a region with a pixel value of 1 corresponds to the distribution area of ​​foreground elements, a region with a pixel value of 0 corresponds to the background area, and a region with a pixel value between 0 and 1 is a grayscale transition region. The grayscale transition region is used to provide continuous blending weights during fusion to smooth the region boundaries.

3. The training-free, layout-controllable pattern generation method according to claim 2, characterized in that, In step S2, the grayscale layout mask M has multiple foreground branches, forming a corresponding mask set. ,in, Represents the i-th foreground region. This indicates that it does not belong to the i-th foreground branch; The background area consists of residual weights. Implicit representation, the calculation formula is: Where i represents the index of the foreground branch, and N represents the total number of foreground branches.

4. The training-free, layout-controllable pattern generation method according to claim 1, characterized in that, In step S3, the background image The generation method is as follows: based on background prompt words Parse the background color parameter or the user-specified background color parameter to generate a solid color background image with the same resolution as the target resolution. ; Background image The method for low-frequency prior and injection of diffuse initial noise is as follows: Will The input is fed into a variational autoencoder to obtain solid color latent variables, and then averaged across the spatial dimensions to obtain the channel-level background vector. Preserve low-frequency color priors and remove high-frequency spatial variations; Sampling yields diffused noise And perform additive injection of background vectors in the background region.

5. The training-free, layout-controllable pattern generation method according to claim 4, characterized in that, The calculation formula for additive injection is: in, For injection intensity hyperparameters, This is an element-wise multiplication operation.

6. The training-free, layout-controllable pattern generation method according to claim 1, characterized in that, The early stage of the diffusion inference is that the denoising time step satisfies... In the later stage of diffusion inference, the denoising time step satisfies... The stage in which, For the noise reduction time step, This is the preset cutoff time threshold for multi-branch denoising.

7. The training-free, layout-controllable pattern generation method according to claim 1, characterized in that, In step S4, the steps for multi-path attention output and spatial weighted fusion are as follows: In the cross-attention layer of the U-Net denoising loop, spatial query features Q, with dimension d, are generated from the current latent variable features; the current background branch prompts are also included. Text embeddings are obtained by text encoders, and key-value pairs are obtained by projection. ; Calculate multi-path attention output : The standard formula for calculating attention is: The grayscale layout mask M is downsampled to match the feature resolution of the current attention layer, and then spatially weighted and fused to obtain the hybrid output. : The fused attention output is used to complete this step of U-Net noise prediction and denoising update.

8. The training-free, layout-controllable pattern generation method according to claim 1, characterized in that, In step S5, the denoising process is completed, and the final latent variables are obtained. Afterwards, The input variational autoencoder decoder obtains the final output image.

9. The training-free, layout-controllable pattern generation method according to claim 1, characterized in that, In step S2, the grayscale layout mask M is generated by at least three methods: direct drawing, coordinate-guided generation, and density-guided generation.

10. The training-free, layout-controllable pattern generation method according to claim 9, characterized in that, The density-guided generation process is as follows: User-specified layout constraint parameters include foreground coverage. Element radius range, minimum interval With the maximum number of points K, on ​​the canvas Initial points are generated by random sampling. ; The interval between all existing points is greater than [a certain value]. Under the given conditions, candidate points are randomly sampled from the canvas; If the current number of points is less than K and the current foreground coverage has not been reached. If the candidate point is not found, it is accepted and added to the point set; this process is repeated until the coverage rate is met or the maximum number of points is reached. Each point is rendered as a circular region with its own radius; and normalized to the grayscale range of [0,1] to obtain a grayscale layout mask.