An image generation method based on multi-reference image dynamic adaptive fusion

CN122335606BActive Publication Date: 2026-09-29ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610813037.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-09-29
Estimated Expiration
2046-06-08

AI Technical Summary

Technical Problem

首先,现有方法以整幅图像为引导对象,缺少对目标区域的显式空间约束,在服装设计等需要对特定区域进行颜色迁移的应用场景中,容易导致颜色信息扩散至背景区域,破坏原始图像的背景一致性与视觉完整性

Benefits of technology

本发明通过构建面向多参考图像的动态自适应融合机制,使综合色彩目标不再表现为预先给定的单一参考颜色分布,也不再表现为多张参考图像颜色信息的固定比例叠加结果,而是表现为与当前生成状态相匹配的综合色彩约束目标。具体地,系统依据当前生成结果对不同参考图像色彩特征的补偿需求,动态分配各参考图像在当前时间步下的贡献强度,并通过时间步间连续平滑机制保持权重演化的稳定性。由此,生成过程能够有选择地吸收不同参考图像中的主色调信息、局部点缀色信息及层次化配色关系,降低颜色差异较大的多参考图像直接融合时出现色彩脏污、浑浊或局部色彩冲突的风险,从而获得具有实际设计意义的综合色彩表达效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122335606B_ABST
    Figure CN122335606B_ABST
Patent Text Reader

Abstract

The application discloses an image generation method based on multi-reference image dynamic self-adaptive fusion. The application dynamically perceives the missing degree of different reference image color features in the current generation state, adaptively constructs a mixed reference color distribution matched with a time-varying generation track, and continuously guides and updates the reverse denoising process of the diffusion model with the distribution as the target, and finally generates a target image fused with elements of multiple reference images. The application effectively solves the color contamination and turbidity problem caused by simple averaging when the color difference of multiple reference images is large, and can generate a high-quality image with natural color and design performance while maintaining the stability of the content image structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence-generated content, computer vision, and digital image processing, and in particular to an image generation method and system based on mask region constraints and dynamic adaptive fusion of multiple reference images. Background Technology

[0002] Image generation techniques based on diffusion models have made significant progress in recent years. In conditional generation tasks, by introducing guiding signals during the inverse denoising process, the generated image can be controlled to meet specific structural, semantic, or stylistic constraints. Among these, using the color distribution of a reference image to guide the generation process is a crucial technical approach for achieving color-conditional generation. For example, the SW-Guidance method proposed by Lobashev et al. introduces slice Wasserstein distance during the sampling process of the diffusion model to calculate the color distribution difference between the generated image and the reference image. It then guides the denoising trajectory with the goal of minimizing this difference, thereby generating an image with a color distribution consistent with the reference image. This method achieves flexible control over the color of the generated image without retraining the model, providing an effective technical solution for color-conditional generation.

[0003] However, existing methods such as SW-Guidance still have the following limitations in practical applications. First, existing methods use the entire image as the guiding object, lacking explicit spatial constraints on the target area. In application scenarios such as clothing design, which require color transfer of specific areas, this can easily lead to color information spreading to the background area, disrupting the background consistency and visual integrity of the original image. Second, existing methods typically rely on a single reference image for color guidance. When users provide multiple reference images with different color styles, they cannot integrate the complementary color features of each reference image. When there are significant color differences between reference images, the generated result can easily result in dirty, muddy, or localized color conflicts. Therefore, how to achieve accurate color guidance for the target area under multiple reference image conditions, while maintaining the structural stability and visual naturalness of the generated result, is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0004] The purpose of this invention is to construct a mixed reference color distribution that matches the current time step by dynamically comparing the color feature differences between the current generated state and each reference image, thereby achieving stable adaptive fusion of multiple reference images. Another purpose of this invention is to improve the visual naturalness of color transfer by uniformly mapping color features to a perceptually uniform color space. Yet another purpose of this invention is to implement hierarchical constraints on the color target from the whole to the details by using a multi-scale slice Wasserstein distance loss algorithm. A further purpose of this invention is to achieve accurate color transfer constraints on the target region by simultaneously applying a binary mask to both the image space and the latent space to guide the update process.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an image generation method based on dynamic adaptive fusion of multiple reference images, comprising the following steps: S1. Obtain the content image, reference image, and mask image; S2. Extract the structural features of the content image and inject them into the initialized diffusion model; S3. Based on the content image and the mask image, generate the target image through the reverse denoising process using the diffusion model; An adaptive fusion operation is performed at any time step of the reverse denoising process: S31. At the current time step, based on the target area indicated by the mask image, calculate the color feature difference between the current generation state and each reference image, thereby obtaining the color compensation requirement intensity corresponding to each reference image. S32. Generate corresponding dynamic weights for each reference image based on the intensity of the color compensation requirement; S33. The color distribution of each reference image is weighted and fused according to the dynamic weight to construct a hybrid reference color distribution; S34. Calculate the distribution difference loss between the current generated state and the mixed reference color distribution, update the latent variables of the inverse denoising process, and obtain the generated state of the next time step.

[0006] Preferably, the calculation of color feature differences in step S31 is performed in the CIELAB perceptual uniform color space.

[0007] Preferably, the greater the difference in color features between the reference image and the current generated state, the higher the intensity of the corresponding color compensation requirement, and the higher the dynamic weight obtained by the reference image.

[0008] Preferably, the dynamic weights are calculated by using a normalized exponential function to assess the intensity of the color compensation requirement. The normalized exponential function includes a temperature coefficient that controls the smoothness of the weight distribution.

[0009] Preferably, step S32 further includes smoothly updating the dynamic weights using an exponential moving average algorithm, and the dynamic weights of the current time step and the previous time step are weighted and fused to obtain the final dynamic weights of the current time step.

[0010] Preferably, the calculation process of the distribution difference loss in step S34 includes: The set of color pixels within the target area is extracted using the mask image; The target region is represented in layers using different spatial scales; At each spatial scale, the hierarchical representation is projected onto multiple random sampling direction vectors to obtain a one-dimensional projection distribution; Calculate the one-dimensional Wasserstein distance between the current generated state and the mixed reference color distribution; The distribution difference loss is obtained by weighted aggregation of the Wasserstein distances at all scales and in all directions.

[0011] Preferably, step S34 also includes a color adjustment tensor, which has the same size as the current latent variable. The gradient of the color adjustment tensor is calculated based on the distribution difference loss, and the color adjustment tensor is updated using the gradient. The updated color adjustment tensor is superimposed on the current latent variable to obtain the latent variable for the next time step.

[0012] Preferably, the method further includes: gating the updated color adjustment tensor with a mask corresponding to the latent space so that it only applies to the target region.

[0013] Preferably, the structural features described in step S2 are extracted using a pre-trained structural control network.

[0014] Preferably, the number of reference images is at least two; the total number of time steps in the reverse denoising process is 20 to 50 steps.

[0015] Compared with the prior art, the present invention has the following beneficial effects: This invention constructs a dynamic adaptive fusion mechanism for multiple reference images, so that the comprehensive color target is no longer represented by a pre-given single reference color distribution, nor by a fixed-ratio superposition of color information from multiple reference images, but rather by a comprehensive color constraint target that matches the current generation state. Specifically, the system dynamically allocates the contribution intensity of each reference image at the current time step based on the compensation requirements of the current generation result for the color features of different reference images, and maintains the stability of weight evolution through a continuous smoothing mechanism between time steps. As a result, the generation process can selectively absorb the dominant color information, local accent color information, and hierarchical color matching relationships from different reference images, reducing the risk of color contamination, turbidity, or local color conflicts when directly fusing multiple reference images with large color differences, thereby obtaining a comprehensive color expression effect with practical design significance.

[0016] This invention maps the current generated state and each reference image to a uniform color space, and performs color statistical analysis, compensation requirement assessment and distribution alignment constraints in a metric domain that is more in line with the laws of visual perception. This reduces the color shift, artifacts and unnatural colors caused by the coupling of brightness and chromaticity in the traditional RGB color space, and improves the perceptual rationality and visual naturalness of the comprehensive color transfer results.

[0017] The multi-scale slice Wasserstein distance loss in this invention is used to apply coarse-to-fine layered constraints to the comprehensive color constraint target. The coarse scale stabilizes the overall comprehensive color direction and dominant color distribution, the intermediate scale coordinates color transitions between local regions, and the fine scale maintains the stability of edge regions, local accent colors, and detail color textures. Thus, the comprehensive color target formed by multi-reference image fusion can be constrained and expressed simultaneously at different levels, resulting in a synergistic improvement in the overall comprehensive color unity, local hierarchical coordination, and detail color continuity of the generated result.

[0018] This invention applies a binary mask of the target region simultaneously to the color distribution alignment process and the latent space-guided update process, thereby concentrating color migration constraints on the target clothing area and suppressing unexpected color disturbances to the background area. This improves the accuracy of local editing boundaries, the controllability of regional effects, and the spatial consistency of the generated results.

[0019] This invention directly incorporates the comprehensive color constraint target, formed by dynamic fusion of multiple reference images, into the inverse denoising sampling process of the diffusion model. It continuously iterates and updates the color adjustment tensor using color distribution alignment loss, and then gates the updated results onto the potential generation trajectory, ensuring that comprehensive color control is continuously guided throughout the entire denoising path. This not only improves the coordination between comprehensive color representation and structure preservation but also enhances the controllability of the generation process, the stability of trajectory convergence, and the consistent achievement of the comprehensive color target across time steps.

[0020] Through the synergy of the above-mentioned technologies, the present invention enables the generated result to maintain the structural outline, texture details and spatial layout stability of the content image while achieving the fusion expression of color dominant information, local accent color information and hierarchical color matching relationship in multiple reference images, thereby obtaining a controllable image generation effect that combines structural fidelity, comprehensive color stability and design expressiveness. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the dynamic adaptive fusion process of multiple reference images in this invention.

[0022] Figure 2 This is a schematic diagram of the multi-scale slice Wasserstein distance loss calculation process in this invention.

[0023] Figure 3 This is a schematic diagram of the potential spatial gradient-guided update process in this invention.

[0024] Figure 4 This is a diagram showing the generated effect of the method of the present invention. Detailed Implementation

[0025] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0026] Example 1 This embodiment describes in detail the multi-reference image dynamic adaptive fusion process in this invention, such as... Figure 1 As shown.

[0027] S1. Multimodal input data acquisition and preprocessing The system receives the content image X c Reference image set And a binary mask image M. Wherein, the content image X... c Used to provide the structural layout, textural basis, and local fold patterns of the target garment; reference image set Used to provide reference color information to be transferred; binary mask image M is used to indicate the target clothing area for color transfer.

[0028] In some implementations, the system can uniformly adjust the content image and reference image set to a resolution that meets the input requirements of the diffusion model, such as 512×512, 768×768, 1024×1024, or other suitable sizes, and perform normalization processing. For binary masked images, the system can simultaneously generate an image spatial mask and a latent spatial mask, wherein the image spatial mask is used for target region color pixel extraction, and the latent spatial mask is used for latent space update gating.

[0029] In some implementations, the latent space mask can be obtained by mapping a binary mask image according to a downsampling ratio corresponding to the latent feature resolution of the diffusion model. The downsampling ratio can be 4, 8, 16, or other ratios consistent with the latent feature size.

[0030] In some implementations, to reduce the local discontinuities caused by hard mask boundaries, the system can perform Gaussian smoothing, mean smoothing, edge feathering, or morphological boundary processing on the boundary regions of the binary mask image. The smoothing radius can be 1 to 10 pixels, preferably 2 to 5 pixels.

[0031] In this embodiment, the system will display the content image X. c and reference image set The input image is uniformly scaled to a size of 1024×1024 and normalized. For a binary masked image M, the system synchronously generates an image space mask Mi. img and latent space mask M lat Among them, the image space mask M img Used for extracting color pixels from a target region, latent space mask M lat Used for latent space update gating. The latent space mask M lat The binary mask image M is obtained by downsampling it according to the latent feature resolution of the diffusion model. In this embodiment, the downsampling ratio of the latent space of the diffusion model relative to the input image is 8. Therefore, when the input image resolution is 1024×1024, the latent space mask M... lat The resolution is 128×128.

[0032] To reduce the local discontinuities caused by hard mask boundaries, the system performs Gaussian smoothing on the boundary regions of the binary mask image M, with a smoothing radius set to 3 pixels.

[0033] S2. Diffusion Model Initialization and Structural Condition Injection The system loads a pre-trained latent diffusion model as the backbone network for image generation and a pre-trained structure control network as the structure conditional branch.

[0034] In this embodiment, the system will display the content image X.c The corresponding edge map is used as input to the structural control network to extract structural features F that can characterize the shape of the garment, the boundaries of the fabric pieces, the direction of the folds, and the geometric details. s During the reverse denoising sampling process, the structural feature F s A denoising network is continuously injected into the diffusion model to constrain the generated result to maintain consistency with the content image X during color transfer. c Consistent clothing silhouettes and spatial structural details.

[0035] In the reverse denoising sampling process of the diffusion model, for the t-th time step, the system is based on the current latent variable Z. t Generate the corresponding denoised image estimate Then the denoised image is estimated. and various reference images Converting from RGB color space to CIELAB color space yields the following results: and .

[0036] In this embodiment, the system utilizes image spatial mask M img right and Pixel sampling is performed on the target clothing area to obtain the color set of the target area estimated in the current denoised image. and the target region color set corresponding to each reference image. By using the above method, the color distribution comparison between the generated result and the reference image is no longer performed directly in the RGB color space, but in the CIELAB color space, which conforms to the characteristics of human eye perception, thereby improving the stability of color statistics and distribution alignment.

[0037] S3. Dynamic adaptive weight fusion of multiple reference images Obtain the color set of the target area and Then, the system calculates the color statistical differences between the current denoised image estimate and each reference image to obtain the color compensation requirement score s for each reference image. t,i In some implementations, the color compensation requirement fraction s t,i The color compensation requirement score s can be calculated based on differences in color mean, variance, covariance, quantile statistics, color histogram, projected distribution, or any combination thereof within the target area. In this embodiment, the color compensation requirement score s is... t,i The color compensation requirement score (s) is calculated based on a combination of color mean, variance, quantile statistics, and color histogram statistics within the target area. t,iThe larger the value, the more significant the deviation of the current denoised image estimate from the color features corresponding to the i-th reference image, or the more significant the inadequacy of the representation, and therefore the higher the demand for color compensation of the reference image.

[0038] In some implementations, the system may use the Softmax function, an exponential normalization function with temperature scaling, a linear normalization function, or other monotonic mapping methods to generate the initial normalized weights for each reference image at the current time step; wherein, the temperature coefficient... The value can be 0.05 to 1.0, preferably 0.1 to 0.5. In this embodiment, the system uses a temperature coefficient... The Softmax function generates the initial normalized weights for each reference image at the current time step. :

[0039] In this embodiment, the temperature coefficient Take 0.3. s t,i The color compensation score for the i-th reference image at the current time step; the temperature coefficient is used to control the concentration of the control weight distribution. The smaller the value, the more the weight is concentrated on higher scores; The larger the value, the more evenly weighted the values. The denominator is the sum of the exponential terms of the color compensation requirement scores of multiple reference images. In this embodiment, there are 3 reference images.

[0040] In some implementations, to avoid drastic fluctuations in weights between adjacent time steps, the system can use an exponential moving average, sliding average, or momentum update mechanism to smoothly update the initial normalized weights; wherein, the smoothing coefficient α can be 0.5 to 0.95, preferably 0.7 to 0.9. In this implementation, the system uses an exponential moving average mechanism to update the initial normalized weights. Perform a smooth update to obtain the final fusion weights. :

[0041] in, This represents the final fusion weight from the previous time step. In this implementation, the smoothing coefficient α is set to 0.8; when t=1, let... =1 / 3.

[0042] Subsequently, the system based on the final fusion weight The color distributions of each reference image within the target region are weighted and fused to construct the fused reference color distribution D corresponding to the current time step. mix,t The number of reference images can be at least two, preferably two to five. The fusion reference color distribution D mix,tIt is used to characterize the comprehensive color constraint target of each reference image on the target clothing region at the current time step, and to guide the current generated result to gradually converge towards the comprehensive color target.

[0043] In this embodiment, the number of reference images can be set to two or more depending on the specific application scenario; in embodiment 4, the number of reference images is two.

[0044] Example 2 This embodiment details a mask-guided multi-scale slice Wasserstein distance loss construction method, such as... Figure 2 As shown.

[0045] In order to simultaneously constrain the overall color tone and local fine-grained color matching of the target clothing area, the system constructs the target area representation at multiple scales by estimating and fusing the reference color distribution of the current denoised image.

[0046] In some implementations, the multi-scale representation can be constructed by pyramid downsampling, region pooling, local block statistics, or a combination thereof, and the number of scales can be 2 to 6, preferably 3 to 4.

[0047] At each scale, the system can sample multiple unit direction vectors and project the target region representation corresponding to the current denoised image estimation and the target region representation corresponding to the fused reference color distribution onto each direction vector, respectively, to obtain the corresponding one-dimensional projection distribution. The number of direction vectors can be 8 to 128, preferably 16 to 64.

[0048] For each projection direction, the system calculates the one-dimensional Wasserstein distance between the one-dimensional projection distribution corresponding to the current denoised image estimate and the one-dimensional projection distribution corresponding to the fused reference color distribution; then, it sums, averages, or weights the one-dimensional Wasserstein distances for all projection directions and all scales to obtain the color distribution alignment loss for the current time step.

[0049] In this embodiment, the system performs pyramid downsampling on the target area to obtain the target area color representation at three scales, and randomly samples 32 unit direction vectors at each scale.

[0050] At each scale, the system randomly samples 32 unit direction vectors. The target region representations estimated from the current denoised image and the target region representations corresponding to the fused reference color distribution are projected onto the respective directional vectors. The corresponding one-dimensional projection distribution is obtained from the above.

[0051] For each projection direction The system calculates the one-dimensional Wasserstein distance between the one-dimensional projection distribution corresponding to the current denoised image estimate and the one-dimensional projection distribution corresponding to the fused reference color distribution; then, it performs weighted aggregation of the one-dimensional Wasserstein distances at all projection directions and all scales to obtain the color distribution alignment loss L at the current time step. mgm-sw :

[0052] Where, λ s Let represent the aggregation weight at the s-th scale. In this embodiment, the aggregation weights for the three scales are λ1=0.5, λ2=0.3, and λ3=0.2, respectively, and satisfy λ 1+ λ 2+ λ3=1. Represents the one-dimensional Wasserstein distance. This represents the color representation of the target region at the s-th scale in the current denoised image. This represents the target region's color representation at the s-th scale, indicating the fusion reference color distribution. Represents the direction vector The projection operation. Due to the color distribution alignment loss L... mgm-sw It is constructed solely based on the target clothing area corresponding to the mask, thus effectively focusing on the color statistical alignment of the target clothing area and reducing interference with the background area.

[0053] Gradient-guided update based on latent spatial perturbation In obtaining the color distribution alignment loss L mgm-sw Then, the system at the current time step is the latent variable z. t Introduce a color adjustment tensor Δz with the same size as its dimensions. t and the color adjustment tensor Δz t This is an optimizable variable at the current time step. In other words, in this implementation, the optimization target is the color adjustment tensor Δz. t Instead of directly dealing with the current latent variable z t It itself is being updated.

[0054] In some implementations, the color adjustment tensor can be initialized independently at each time step, or inherited based on the update result of the previous time step. The system participates in the calculation of the color distribution alignment loss at the current time step with the perturbed latent representation, and calculates the gradient of the color adjustment tensor based on the color distribution alignment loss.

[0055] In some implementations, the system can update the color adjustment tensor according to gradient descent, gradient update of momentum, or adaptive step size update; wherein the update step size can be 0.001 to 0.1, preferably 0.005 to 0.05.

[0056] Example 3 This embodiment describes in detail a method for guiding color adjustment tensor updates in the latent space, specifically as follows: Figure 3 As shown, the system uses a latent spatial mask to perform a gating operation on the updated color adjustment tensor to retain only the valid update amount within the target clothing area. The gating operation can be a binary hard gating or a continuous soft gating.

[0057] Finally, the system superimposes the gated color adjustment tensor onto the current latent variable according to the color guidance intensity coefficient to obtain the guidance latent variable at the current time step; wherein, the color guidance intensity coefficient can be 0.1 to 2.0, preferably 0.3 to 1.0.

[0058] In this embodiment, the system updates the color adjustment tensor according to the gradient descent method, with an update step size of 0.01 and a color guiding intensity coefficient of 0.6.

[0059] The system uses the perturbed latent representation z t +Δz t The color distribution alignment loss is calculated for the current time step, and based on the color distribution alignment loss L... mgm-sw Color adjustment tensor Δz t Calculating the gradient, we get:

[0060] Subsequently, the system updates the color adjustment tensor Δz according to the gradient descent method. t This is based on the gradient descent optimization algorithm, whose core idea is to update variables along the negative direction of the loss function's gradient to reduce the loss function's value. (Gradient) Since the direction of the fastest increase in loss is in the direction we want to minimize, we should move in the opposite direction:

[0061] In this embodiment, the update step size is updated. Take 0.01.

[0062] Then, the system utilizes the latent space mask M lat Updated color adjustment tensor Perform a gating operation to retain only the effective update amount within the target garment area, and obtain the gating color adjustment tensor. :

[0063] in, This indicates element-wise multiplication.

[0064] Finally, the system adjusts the color tensor after gating. According to color guidance intensity coefficient Superimposed on the current latent variable z t To obtain the guiding latent variables at the current time step :

[0065] During this process, Δz t , , These represent the original, updated, and gated color adjustment tensors, respectively. In this embodiment, the color guiding intensity coefficient... We set it to 0.6. Using the above method, the color distribution alignment loss L... mgm-sw It can directly affect the potential generation trajectory of the diffusion model within the target clothing area, thereby making the generation result gradually approach the fusion reference color distribution at each time step.

[0066] The system repeats the generation steps throughout the entire reverse denoising sampling period of the diffusion model, so that the generated image gradually maintains the clothing structure corresponding to the content image under the constraints of the structure control network, and gradually converges to the fusion reference color distribution under the guidance of mask and multi-scale distribution alignment.

[0067] In some implementations, the number of reverse denoising sampling steps can be set to 20 to 50 steps, preferably 20 to 30 steps; the sampling termination condition can be reaching a preset time step, or the color distribution alignment loss satisfying a preset convergence condition.

[0068] In this embodiment, the number of reverse denoising sampling steps is set to 30. After the sampling reaches the termination time step, the system inputs the final latent variable z0 into the diffusion model decoder to obtain the output image X. out The output image combines the garment structural details provided by the content image with the comprehensive color information provided by multiple reference images, and can be applied to scenarios such as garment design image generation, product visual redrawing, and fashion creative solution generation.

[0069] Example 4 To clearly illustrate the effectiveness of the multi-reference image dynamic adaptive fusion mechanism in this invention and its resulting technical effects, a specific embodiment is given below, with the specific effects as follows: Figure 4 As shown.

[0070] Input data preparation In this embodiment, the content image is a front view of a double-breasted suit jacket, and its structural outline, lapel shape, placket position, and double-breasted layout are shown in the attached figure. Figure 4 As shown in (a) above. The reference image set includes two reference images, the first of which is shown in the attached image. Figure 4 As shown in (b), this is a stylistic image with a soft, abstract topographical layering effect. Its overall color scheme is dominated by mint green, light coral pink, off-white, and warm gray, with soft, luminous gradient areas and circular highlight elements, presenting a light, soft, and airy overall color characteristic. The second reference image is attached. Figure 4 As shown in (c), this is an abstract image with a sunset motif. Its overall color scheme is dominated by highly saturated orange-red, blue-gray, light gray, and sandy brown. The orange-red circular area forms the main visual center, while the blue-gray wavy areas create strong layering and overall color contrast. A binary mask image is used to precisely cover the outer main body area of ​​the content image, thus limiting the scope of color migration and distribution alignment.

[0071] In this embodiment, to quantitatively characterize the comprehensive color features of the reference images, the pixels of the two reference images within the target region are first statistically analyzed in the CIELAB color space, and the centers and area proportions of their respective dominant color clusters are extracted using a K=4 clustering method. For the attached... Figure 4 In (b), the four main color cluster centers within the target area are: C1-1=(74.8,-16.2,8.5), accounting for 36%; C1-2 = (81.3, 14.7, 8.9), accounting for 24%; C1-3 = (88.5, 1.8, 9.7), accounting for 22%; C1-4 = (67.2, 3.6, 6.1), accounting for 18%.

[0072] As can be seen, the overall color characteristics of Figure 4(b) are mainly manifested as a soft transition relationship between mint green, light pink and off-white under high brightness and low to medium saturation.

[0073] For Figure 4(c), the centers of the four main color clusters within the target area are as follows: C2-1 = (62.1, 38.4, 42.8), accounting for 34%; C2-2=(48.7,-2.9,-16.5), accounting for 28%; C2-3 = (79.6, 1.1, 4.3), accounting for 20%; C2-4=(57.4,10.2,22.4), accounting for 18%.

[0074] As can be seen, the overall color characteristics of (c) in Figure 4 are mainly characterized by the coexistence of bright orange-red main tone and low blue-gray contrasting color, supplemented by light gray and sand brown transitional layers.

[0075] Dynamic fusion process at key time steps In the 15th time step of the diffusion model reverse denoising, the system first obtains the denoised image estimate for the current time step. According to statistics, the average CIELAB color value estimated in the denoised image within the target coat area is (70.6, 4.8, 12.1). This result indicates that the target area has formed a preliminary garment outline and stripe layout, and the overall brightness structure is basically established. However, the overall color still leans towards a low-saturation light warm gray, and has not yet fully expressed the mint green to light pink soft gradient features in Figure 4(b), nor has it fully expressed the orange-red highlight to blue-gray contrast layer features in Figure 4(c).

[0076] In this embodiment, the color compensation requirement score is calculated by weighting the differences in the mean color of the target area, the differences in color dispersion, and the differences in projection distribution, and can be expressed as:

[0077] The above formula is one implementation method for obtaining the weighted color compensation requirement score in Example 1, used to quantitatively assess color differences. The weighting coefficients can be adjusted according to specific application requirements; this embodiment is only for illustrative purposes. In this embodiment, the weights corresponding to the weighted differences in the target area's mean color value, color dispersion, and projection distribution are 0.4, 0.3, and 0.3, respectively. Δμ i This represents the difference in the mean color value of the current denoised image estimate relative to the i-th reference image within the target region. Δσ i This indicates the degree of color dispersion. Δp i This represents the one-dimensional distribution difference term along multiple random projection directions.

[0078] For Figure 4(b), the current denoised image estimation is still insufficient in representing the soft, comprehensive color bands such as mint green, light pink, and off-white. The corresponding three difference values ​​are as follows:

[0079] Therefore, the corresponding color compensation requirement score is calculated as follows:

[0080] For Figure 4(c), the current denoised image estimation also shows significant deficiencies in the representation of the orange-red highlight tone and blue-gray contrast levels, with the corresponding three difference values ​​as follows:

[0081] Therefore, the corresponding color compensation requirement score is calculated as follows:

[0082] Based on this, the system uses a Softmax function with a temperature coefficient τ=0.3 to normalize the color compensation requirement scores, obtaining the initial weights of the two reference images at the current time step as follows:

[0083]

[0084] Furthermore, combining the fusion weights [0.40, 0.60] from the previous time step and the smoothing coefficient α = 0.8, the fusion weights for the current time step are updated using an exponential moving average mechanism to obtain the final fusion weights:

[0085]

[0086] Based on the aforementioned final fusion weights, the system performs a weighted fusion of the comprehensive color distributions in Figure 4(b) and Figure 4(c) within the target area, thereby constructing the fusion reference color distribution corresponding to the current time step. The calculated average CIELAB color value of the fusion reference color distribution within the target area is: (67.95, 7.95, 12.21).

[0087] Furthermore, the main composite color clusters after fusion and their proportions are approximately as follows: orange-red highlight 20.1%, blue-gray contrast 16.6%, mint green 14.7%, light coral pink 9.8%, off-white highlight 20.8%, sand brown transition color 10.7%, and warm gray balancing color 7.3%.

[0088] As can be seen, the fusion reference color distribution retains the orange-red highlight main tone and blue-gray contrast layer in Figure 4(c), while superimposing the soft transition characteristics of mint green, light pink and beige in Figure 4(b), thus forming a comprehensive color target that combines the unity of the overall color tone, local layering changes and a soft and dreamy feel.

[0089] Final output results and technical effect analysis After a complete 30-step reverse denoising sampling process, the final output image is shown in the attached image. Figure 4As shown in (d) of Figure 4, the generated double-breasted coat maintains the original lapel structure, placket direction, double-breasted layout, and overall pattern stability while successfully integrating complementary color features from the two reference images. Specifically, the orange-red main tone in Figure 4(c) is prominently reflected in the collar and upper striped area, giving the output a clear visual center and strong overall color tension; the mint green, light pink, off-white, and soft glowing gradient features in Figure 4(b) are transferred to the striped areas of the body and sleeves, resulting in a lighter, softer, and more layered overall color transition effect. At the same time, the stripe boundaries, coat outline, and button positions remain stable, with no obvious color overflow, local breaks, or structural distortion.

[0090] in conclusion As demonstrated in this embodiment, the dynamic adaptive fusion mechanism of this invention does not simply perform an arithmetic average of the colors of multiple reference images. Instead, it dynamically and selectively extracts complementary color elements from multiple reference images with significant differences based on the current generation state, and constructs a fusion reference color distribution with practical constraints. This approach not only effectively fuses the dominant color information and local accent color information, but also generates output results with overall color unity, local design appeal, and high visual naturalness while maintaining the structural stability of the target object and the undisturbed background area. This effectively overcomes the problems of overall color distortion and style confusion that are easily caused by traditional fixed-weight fusion methods under multi-reference image conditions.

Claims

1. An image generation method based on dynamic adaptive fusion of multiple reference images, characterized in that, Includes the following steps: S1. Obtain the content image, reference image, and mask image; S2. Extract the structural features of the content image and inject them into the initialized diffusion model; S3. Based on the content image and the mask image, generate the target image through the reverse denoising process using the diffusion model; An adaptive fusion operation is performed at any time step of the reverse denoising process: S31. At the current time step, based on the target area indicated by the mask image, calculate the color feature difference between the current generation state and each reference image, thereby obtaining the color compensation requirement intensity corresponding to each reference image. S32. Generate corresponding dynamic weights for each reference image based on the intensity of the color compensation requirement; S33. The color distribution of each reference image is weighted and fused according to the dynamic weight to construct a hybrid reference color distribution; S34. Calculate the distribution difference loss between the current generated state and the mixed reference color distribution, update the latent variables of the inverse denoising process, and obtain the generated state of the next time step.

2. The method according to claim 1, characterized in that, The calculation process of color feature differences described in step S31 is carried out in the CIELAB perceptual uniform color space.

3. The method according to claim 1, characterized in that, The greater the difference in color features between the reference image and the current generated state, the higher the intensity of its corresponding color compensation requirement, and the higher the dynamic weight obtained by the reference image.

4. The method according to claim 1, characterized in that, The dynamic weights are calculated by using a normalized exponential function to assess the intensity of the color compensation requirement. The normalized exponential function includes a temperature coefficient that controls the smoothness of the weight distribution.

5. The method according to claim 1, characterized in that, Step S32 further includes smoothing the dynamic weights using an exponential moving average algorithm, and weighting and fusing the dynamic weights of the current time step with the dynamic weights of the previous time step to obtain the final dynamic weights of the current time step.

6. The method according to claim 1, characterized in that, The calculation process of the distribution difference loss in step S34 includes: The set of color pixels within the target area is extracted using the mask image; The target region is represented in layers using different spatial scales; At each spatial scale, the hierarchical representation is projected onto multiple random sampling direction vectors to obtain a one-dimensional projection distribution; Calculate the one-dimensional Wasserstein distance between the current generated state and the mixed reference color distribution; The distribution difference loss is obtained by weighted aggregation of the Wasserstein distances at all scales and in all directions.

7. The method according to claim 1, characterized in that, Step S34 also sets a color adjustment tensor, which has the same size as the current latent variable; The gradient of the color adjustment tensor is calculated based on the distribution difference loss, and the color adjustment tensor is updated using the gradient. The updated color adjustment tensor is superimposed on the current latent variable to obtain the latent variable for the next time step.

8. The method according to claim 7, characterized in that, Also includes: The updated color adjustment tensor is gated using a mask corresponding to the latent space, so that it only applies to the target region.

9. The method according to claim 1, characterized in that, The structural features described in step S2 are extracted through a pre-trained structural control network.

10. The method according to claim 1, characterized in that, The number of reference images is at least two; the total number of time steps in the reverse denoising process is 20 to 50.