Diffusion converter method for realizing material migration only by using image condition

By using the diffusion transformer method to uniformly process material, lighting, and depth image labels in a shared latent space, this approach addresses the lack of diversity and text dependency issues in existing material transfer methods. It achieves high-quality material transfer that is efficient and training-free, making it suitable for digital content creation and industrial design.

CN120876336APending Publication Date: 2025-10-31TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510981120.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing material transfer methods rely on parametric modeling and complex architectures, resulting in insufficient diversity of generated results, failing to meet the customized needs of digital art creation, and suffering from text dependence, high computational costs, and feature misalignment.

Method used

The diffusion transformer method is adopted to uniformly process material, lighting and depth image labels by sharing a latent spatial encoder. It utilizes multimodal attention mechanism and cross-bias modulation to achieve cross-modal interaction, combines low-rank adaptive module to enhance depth image control, performs background-preserving blending operation, and generates high-quality material transfer results.

Benefits of technology

It achieves high-quality material transfer without text prompts and model fine-tuning, reduces architectural complexity and inference time, improves the generalization ability and efficiency of the model, and generates materials with good consistency with the structure, making it suitable for digital content creation and industrial design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876336A_ABST
    Figure CN120876336A_ABST
Patent Text Reader

Abstract

The invention provides a diffusion converter method for realizing material migration only by using an image condition, which comprises the following steps of: acquiring a target image and a material image, generating an illumination image and a depth image, and converting the illumination image and the depth image into a unified image mark through a shared potential space encoder; and performing unified processing through a multi-modal attention mechanism of a diffusion converter, realizing cross-modal interaction in combination with cross deviation modulation, enhancing depth control by using a low-rank adaptive module, performing background maintenance, mixing and fusing a foreground and a background, and finally generating an output image inheriting a target structure and a material texture. According to the MaTe framework of the method, input images are integrated at the token level, potential space is shared, multi-modal attention unified processing is carried out, dependence on an adapter, ControlNet and the like is eliminated, text prompt and model fine adjustment are not needed, zero-shot high-quality material generation is achieved, framework complexity is reduced, efficiency and quality are improved, and the method is suitable for the fields of digital creation, industrial design and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to computer vision technology, and in particular to a diffusion transformer method for achieving material transfer using only image conditions. Background Technology

[0002] Material transfer is a technique that precisely maps the properties of a specific material sample onto the surface of a target object. Due to its broad application prospects in digital content creation and industrial design, this technology has attracted considerable attention in recent years. Traditional material transfer frameworks primarily rely on parametric modeling paradigms, such as optical reflection models based on bidirectional scattering distribution functions or procedural texture generation algorithms. However, due to the limited size of the underlying material library, these frameworks restrict the diversity of generated results and cannot meet the customized needs of non-uniform composite materials in digital art creation.

[0003] Recent diffusion-based material transport methods rely on image fine-tuning or complex architectures with auxiliary networks, but face challenges including text dependency, additional computational cost, and feature misalignment.

[0004] In recent years, inspired by the breakthroughs achieved by diffusion models in conditional generation tasks, diffusion-based material transfer methods have made significant improvements in high-quality material transfer results. Current state-of-the-art methods typically employ fine-tuning of the diffusion model and binding the sample set to textual identifiers (such as...). Figure 2 (a) implicitly encodes material features through conceptual semantic binding. While these methods address some drawbacks of traditional pipelines, their heavy reliance on textual cues limits fine-grained control over material properties. Furthermore, the full-parameter fine-tuning paradigm significantly increases training costs and the risk of overfitting. Recent research has also introduced pre-trained general image encoder IP-Adapters to extract material features, and ControlNet to inject depth information, such as... Figure 2 As shown in (b). However, these methods often result in hierarchical decoupling between material and structural information during the generation process, rather than seamless integration, and are also affected by increased inference time.

[0005] It should be noted that the information disclosed in the background section above is only for understanding the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The main objective of this invention is to overcome the deficiencies in the aforementioned background technology and provide a diffusion transformer method for achieving material transfer using only image conditions.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] A diffusion transform method for achieving material transfer using only image conditions includes the following steps:

[0009] S1. Obtain the target image and material image, and generate the corresponding lighting image and depth image;

[0010] S2. The material image, lighting image, and depth image are uniformly converted into image tags using a shared latent spatial encoder;

[0011] S3. A multimodal attention mechanism using a diffusion converter model is employed to perform unified sequence processing on the image tags, and cross-modal interaction is achieved through cross-bias modulation.

[0012] S4. Enhance the control of depth image labeling over the generation process by using a low-rank adaptive module;

[0013] S5. Perform background-preserving blending operation to blend the foreground generated during diffusion with the background of the target image in the latent space;

[0014] S6. Generate an output image that inherits the target image structure and material texture based on the processed markers.

[0015] Further, in step S1, generating the illumination image includes:

[0016] Extract the foreground mask from the target image;

[0017] Convert the foreground area to grayscale to eliminate interference from the underlying color while preserving geometric shadows;

[0018] By blending the grayscale foreground with the background area of ​​the target image, a lighting image decoupled from material interference is generated.

[0019] Further, in step S3, the unified sequence processing includes:

[0020] The material image markers, lighting image markers, and depth image markers are directly concatenated into a joint sequence;

[0021] The correlation between any markers in the joint sequence is calculated using the multimodal attention block of the diffusion transformer;

[0022] The cross-bias modulation controls the cross-modal attention weights through an adjustable logarithmic bias matrix, which satisfies:

[0023] Main diagonal blocks maintain in-modal attention invariance;

[0024] Off-diagonal blocks adjust the interaction intensity between material markers and lighting / depth markers using a logarithmic intensity factor.

[0025] Furthermore, in step S3, the multimodal attention mechanism specifically includes:

[0026] In the diffusion converter block, spatial information is preserved through rotational position encoding;

[0027] Project the concatenated joint sequence into a query, key, and value vector;

[0028] Material, lighting, and depth features are fused based on cross-modal attention weights.

[0029] Furthermore, in step S3, the cross-bias modulation dynamically adjusts the condition strength during the testing phase:

[0030] When the logarithmic intensity factor approaches negative infinity, cross-modal interaction between material markers and lighting / depth markers is suppressed;

[0031] When the logarithmic intensity factor is greater than zero, the effect of the material marker on the lighting / depth marker is enhanced.

[0032] Furthermore, the training optimization of the diffusion converter model adopts a conditional flow matching objective function, and the corrected flow velocity is estimated through a noise prediction network.

[0033] Further, in step S4, the operation of the low-rank adaptive module includes:

[0034] Freeze the weight matrix of the pre-trained diffusion transformer;

[0035] Introducing trainable low-rank matrices to enhance the features of deep image labels;

[0036] The influence of depth information on the generation process is adjusted by using hyperparameter weights.

[0037] Further, in step S5, the background-preserving blending operation includes:

[0038] At each step of the diffusion process, the conditionally generated latent noise representation is fused with the corresponding noise version of the target image using a foreground mask weighted fusion.

[0039] In the final step, the foreground region of the generated image is directly stitched together with the original background region of the target image.

[0040] Furthermore, the method achieves material transfer without model fine-tuning or text prompts, including:

[0041] An image encoder that reuses a pre-trained diffusion transformer processes all input images;

[0042] Replace additional adapters or control networks with unified sequence processing and multimodal attention mechanisms.

[0043] A computer program product includes a computer program that, when executed by a processor, implements the diffusion transformer method for material migration using only image conditions.

[0044] The present invention has the following beneficial effects:

[0045] This invention proposes a diffusion transformer method for material transfer using only image conditions. A concise and efficient architecture (called MaTe) is designed, integrating the input image at the token level and achieving unified processing using multimodal attention in a shared latent space. This eliminates the need for other adapters, ControlNet, inverse sampling, or model fine-tuning, significantly reducing architectural complexity and inference time. This method requires no task-specific model training or fine-tuning, nor does it require manually setting text information. It infers all necessary information from the input image using an existing pre-trained diffusion model, achieving high-quality material generation in zero-shot, training-free paradigms. This enhances the model's generalization ability and adapts to various materials and scenes. Simultaneously, the natural fusion of material, geometry, and lighting information is achieved through multimodal attention mechanisms and cross-bias modulation, avoiding problems such as material geometry misalignment and inconsistent lighting reflections, resulting in high-quality material transfer results with consistent detail. Furthermore, LoRA enhances the integration of depth information in the diffusion model, flexibly controlling the influence of depth on the generation process. Combined with unified sequence processing and cross-bias modulation mechanisms, the denoising network jointly processes all image signals, further improving the model's efficiency and performance, providing broader application prospects for digital content creation and industrial design. Extensive experiments have shown that this invention achieves high-quality material generation in a zero-shot, training-free paradigm, outperforming existing methods in both visual quality and efficiency, while maintaining precise detail alignment, thus greatly simplifying inference prerequisites.

[0046] Compared with the prior art, the significant advantages of the present invention are reflected in the following aspects:

[0047] Significantly reduced architectural complexity and inference time: This invention proposes a concise material transport architecture that integrates input images at the token level and achieves unified processing through multimodal attention in a shared latent space, eliminating the need for other adapters, ControlNet, inverse sampling, or model fine-tuning, thereby significantly reducing architectural complexity and inference time.

[0048] No fine-tuning or manual text setting required: This invention simplifies the inference process, eliminating the need for task-specific model training or fine-tuning, and also eliminating the need for manual text setting. It utilizes existing pre-trained diffusion models to infer all necessary information from input images, achieving high-quality material generation in a zero-shot, training-free paradigm.

[0049] Producing high-quality material transfer results: This invention constructs a diverse and extensive real-world material transfer evaluation dataset. Its method produces high-quality material transfer results with consistent detail, outperforming state-of-the-art baselines in both qualitative and quantitative analysis. Through a multimodal attention mechanism and cross-bias modulation, it achieves a natural fusion of material, geometry, and lighting information, avoiding problems such as material geometry misalignment and inconsistent lighting reflections.

[0050] Improving Model Efficiency and Performance: By enhancing the integration of depth information in the diffusion model through LoRA, this invention allows for flexible control over the impact of depth on the generation process, thereby improving model efficiency and performance. Simultaneously, unified sequence processing and cross-bias modulation mechanisms enable the denoising network to jointly process all image signals, further enhancing model performance.

[0051] Enhanced model generalization ability: Since this invention achieves high-quality material generation without a training paradigm, it has stronger generalization ability and can adapt to various materials and scenarios, providing a wider range of application prospects for digital content creation and industrial design.

[0052] Other beneficial effects of the embodiments of the present invention will be further described below. Attached Figure Description

[0053] Figure 1 This is the overall flowchart of the diffusion transformer method for material transfer using only image conditions according to the present invention.

[0054] Figure 2 A structural comparison diagram of different material migration methods.

[0055] Figure 3 This is a schematic diagram of the architecture of the MaTe method according to an embodiment of the present invention.

[0056] Figure 4 This is an ablation experiment result diagram of an embodiment of the present invention when the light-related input information is replaced with different images.

[0057] Figure 5 A comparison chart showing the effects of different methods on complex materials.

[0058] Figure 6 These are illustrations showing the effects of material transfer using the method of the present invention in various processing scenarios. Detailed Implementation

[0059] The embodiments of the present invention will be described in detail below. It should be emphasized that the following description is merely exemplary and not intended to limit the scope and application of the present invention.

[0060] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of the present invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0061] See Figure 1 This invention provides a diffusion transformer method for achieving material transfer using only image conditions, comprising the following steps:

[0062] Step S1: Obtain the target image and material image, and generate the corresponding lighting image and depth image.

[0063] In some embodiments, step S1, generating the illumination image includes:

[0064] Extract the foreground mask from the target image;

[0065] Convert the foreground area to grayscale to eliminate interference from the underlying color while preserving geometric shadows;

[0066] By blending the grayscale foreground with the background area of ​​the target image, a lighting image decoupled from material interference is generated.

[0067] Step S2: Convert the material image, lighting image, and depth image into image tags using a shared latent spatial encoder.

[0068] Step S3: Using the multimodal attention mechanism of the diffusion transformer model, the image tags are processed in a unified sequence, and cross-modal interaction is achieved through cross-bias modulation.

[0069] In some embodiments, step S3, the unified sequence processing includes: directly concatenating material image tags, lighting image tags, and depth image tags into a joint sequence; calculating the correlation between any tags in the joint sequence using a multimodal attention block of a diffusion transformer; and controlling the cross-modal attention weights through an adjustable logarithmic bias matrix, which satisfies the following: the main diagonal block maintains the intramodal attention unchanged; and the off-diagonal block adjusts the interaction intensity between the material tag and the lighting / depth tag through a logarithmic intensity factor.

[0070] In some embodiments, in step S3, the multimodal attention mechanism specifically involves: preserving spatial information through rotational position encoding in the diffusion transform block; projecting the concatenated joint sequence into a query, key, and value vector; and fusing material, lighting, and depth features based on cross-modal attention weights.

[0071] In some embodiments, in step S3, the cross-bias modulation dynamically adjusts the conditional intensity during the testing phase: when the logarithmic intensity factor approaches negative infinity, it suppresses cross-modal interaction between the material marker and the illumination / depth marker; when the logarithmic intensity factor is greater than zero, it enhances the influence of the material marker on the illumination / depth marker.

[0072] In some embodiments, the training optimization of the diffusion converter model employs a conditional flow matching objective function, and the corrected flow velocity is estimated through a noise prediction network.

[0073] Step S4: Enhance the control of depth image labeling over the generation process using a low-rank adaptive module.

[0074] In some embodiments, in step S4, the operation of the low-rank adaptive module includes: freezing the weight matrix of the pre-trained diffusion transformer; introducing a trainable low-rank matrix to enhance the features of the depth image label; and adjusting the influence of depth information on the generation process through hyperparameter weights.

[0075] Step S5: Perform a background-preserving blending operation to blend the foreground generated during the diffusion process with the background of the target image in the latent space.

[0076] In some embodiments, step S5, the background-preserving blending operation includes: in each step of the diffusion process, fusing the conditionally generated latent noise representation with the corresponding noise version of the target image by foreground mask weighting; and in the final step, directly stitching the foreground region of the generated image with the original background region of the target image.

[0077] Step S6: Generate an output image that inherits the target image structure and material image texture based on the processed markers.

[0078] In some embodiments, the diffusion transformer method for material transfer achieves material transfer without model fine-tuning or text prompts, including: reusing the image encoder of the pre-trained diffusion transformer to process all input images; and replacing additional adapters or control networks with unified sequence processing and multimodal attention mechanisms.

[0079] The following further describes specific embodiments of the present invention, algorithm examples, and experimental verification.

[0080] A diffusion transformer method for material transfer using only image conditions is proposed. A training-free, text-hint-free zero-shot image-to-image framework, MaTe, is designed to enable structure and material control. The MaTe architecture handles image conditions uniformly through multimodal attention, eliminating reliance on complex control modules. The unified sequence processing and cross-bias modulation mechanism in MaTe enable the denoising network to jointly process all image signals, requiring only a lightweight low-rank adaptive (LORA) layer to enhance depth information.

[0081] Given target image I input and material image I M The output image I generated by MaTe o Inherited from I input Structure and I M The material.

[0082] DiT (Diffusion Transformer) is a diffusion model denoising network based on the Transformer architecture. It processes image labels through a self-attention mechanism to achieve iterative refinement of noisy images. As the basic model of the MaTe framework, DiT handles multimodal conditions (such as image, lighting, and depth) to achieve joint control of materials and structures.

[0083] like Figure 3 As shown, the method of this embodiment of the invention uses the multimodal attention mechanism of DIT to simply combine three types of image tags (material image tags C) M Depth image labeling C D and lighting image marker C I This allows for interactive transfer of materials, ensuring high-quality material transfer and maintaining them within the same feature space throughout the diffusion process. It eliminates the need for unnecessary image fine-tuning and textual guidance, thus simplifying the reasoning process.

[0084] In providing target image I input and material image I M Subsequently, the present invention obtains illumination image I and depth image I. D I M , I and I D The unified sequence processing and cross-bias modulation in the MaTe architecture designed in this invention enable control over the structure and materials. Subsequently, LoRA is used to enhance depth effects, and background-preserving blending is employed to strengthen the consistency between the foreground and background.

[0085] Model Basics

[0086] Rectification model

[0087] The generative model aims to define a mapping from a sample x1 of a noise distribution p1 to a sample x0 of a data distribution p0, where p0 represents the real image in the image generation task. The correction flow defines a forward process that constructs a path between distributions p0 and p1, like a straight line trajectory, as shown in Equation 1, where... Here, the forward process is time-dependent due to the time step t.

[0088] x t =(1-t)x0+t∈,∈~N(0,1) (1)

[0089] To learn this mapping, a network is trained with parameters θ to estimate the velocity v of the corrected flow, denoted as v0. θ Then, using a reparameterization method, this velocity prediction network can be used as a noise prediction network. θ The objective function is optimized using conditional flow matching, as shown in Equation 2.

[0090]

[0091] Where, λ' t w represents the signal-to-noise ratio in reparameterization. t It is a time-dependent weight function.

[0092] Among them, the conditional flow matching objective function trains the noise prediction network in the MaTe framework by matching the flow fields of the conditional distribution and the target distribution, ensuring that the model responds accurately to the image conditions.

[0093] Multimodal diffusion converter

[0094] DiT models are applied in architectures such as FLUX.1, Stable Diffusion 3, and PixArt, using a transformer as a denoising network to iteratively refine noisy image labels.

[0095] like Figure 2 As shown in (a), the DiT model handles two types of labels: noisy image labels. and text conditional tags Where d is the embedding dimension, and N and M are the number of image and text tags, respectively. These tags maintain a consistent shape throughout the network as they pass through multiple transformer blocks.

[0096] As the implementation basis for the MaTe framework reference, it uses the DiT-based generative model FLUX.1 and employs multimodal attention to process image and text conditions to achieve high-quality image generation.

[0097] In FLUX.1, each DiT block consists of a layer normalized followed by a multimodal attention MMA, where Rotational Position Embedding (RoPE) is used to encode spatial information. MMA is an attention mechanism that allows interaction between different modal features (such as images, text, and depth), achieving information fusion by calculating the correlation between cross-modal labels. The MaTe framework uses MMA as the core mechanism for uniformly processing material, lighting, and depth image labels, eliminating dependence on additional adapters. RoPE is a technique for adding positional information to sequential data, encoding the vector representation of labels through a rotation matrix, enabling the model to perceive the spatial order of the input. The DiT model uses RoPE to encode the spatial location of image labels, improving the model's understanding of image structure.

[0098] The multimodal attention mechanism then projects the position-encoded tags into the query Q, key K, and value V representations. This allows for the computation of attention among all tags:

[0099]

[0100] Where [X; C] T The symbol [] represents a connection between an image and a text tag. This form enables bidirectional attention. Based on the DiT architecture and implemented using FLUX.1, this invention aims to develop MaTe, a framework that balances superior performance and minimalist design, guiding material transfer solely through visual conditions.

[0101] Depth and lighting guidance

[0102] Illumination consistency

[0103] Diffusion models typically start sampling from random noise, which can lead to the loss of illumination information from the optical priors in the initial image. To address this issue, this invention employs a foreground illumination image I that preserves illumination information without the underlying color priors. The denoising process is initialized using the following formula:

[0104] I=F⊙I gray +(1-F)⊙I input (4)

[0105] Here, F represents the target mask, and I... gray This represents a grayscale image. The formula decouples illumination preservation from material interference: (1-F)⊙I input Some ambient lighting characteristics (direction, intensity, and hue) were preserved, while F⊙I gray It eliminates the inherent color pollution of the original target while preserving geometric shadow information.

[0106] Figure 4 Showing C IAblation experiment results when set to different images. For example... Figure 4 The ablation study in the study showed that, compared with random noise, the illumination image I can better preserve illumination information.

[0107] To address the inherent challenges of decoupling geometric and material properties in images and the need for additional training data, this invention proposes an enhancement solution that integrates the depth image D as a geometric prior into the diffusion model, thereby strengthening the structural representation of the target object in a lightweight, standard manner. Experimental results confirm that this design achieves a 1.8x inference speedup in material transfer tasks while maintaining geometric fidelity.

[0108] MaTe

[0109] Architecture Design

[0110] To achieve versatile architectural changes, MaTe first reuses the variational autoencoder (VAE) from the base DiT framework, treating the illumination image I as a noisy image and projecting the material image into the same latent space labeled as the noisy image. This approach contrasts sharply with previous methods that relied on separate feature extractors, significantly reducing architectural complexity. The VAE, as a generative model, learns a distributed representation of the data by mapping and reconstructing images into the latent space. The MaTe framework reuses its VAE encoder portion, projecting the illumination and material images into a shared latent space, further reducing architectural complexity.

[0111] This invention achieves material transfer by utilizing the spatial correspondence in multimodal attention blocks, without requiring a specific theme.

[0112] Encoded depth condition marker C D With illumination image label C I Having the same dimensions and latent space allows them to be processed directly by the transformer block. Since the conditions and image labels reside in the same latent space, MaTe utilizes existing DiT blocks to process them jointly.

[0113] It is worth noting that this invention innovatively proposes cross-bias modulation, a novel mechanism that achieves controllable cross-modal interaction through structured logarithmic bias, thereby enabling free control between conditions.

[0114] A diffusion model component combining UNets and DiT can be used to achieve image denoising and feature extraction.

[0115] LoRA

[0116] The method enhances parameter efficiency by freezing the pre-trained weight matrix and introducing an additional trainable low-rank matrix into the neural network. This approach is based on the observation that the pre-trained model has low "intrinsic dimensionality." Specifically, for the diffusion model ∈ θ A weight matrix Introducing the LoRA module involves updating W to W', defined as W' = W + BA. Here, and It is a matrix with low-rank factor r, satisfying r << min(n,m).

[0117] In this invention, a LoRA module is used to control depth information. By assigning specific weights w to the LoRA module, this invention can adjust its influence on the generation process:

[0118] W' = W + w × BA. (5)

[0119] The weight w is typically an empirically tuned hyperparameter used to balance the relationship between depth information and other generated features. This allows LoRA to effectively integrate depth information into the diffusion model, generating images that conform to depth structure, unlike previous material transfer methods using ControlNet.

[0120] Unified Sequence Processing

[0121] Previous methods, such as ControlNet and T2I-Adapter, introduced conditional images into the model by directly adding features:

[0122] X←X+C D (6)

[0123] Among them, conditional feature C D Align the image with the noisy image tag X spatially and add them together. While this approach is effective for spatial alignment tasks, it faces two limitations: (1) a lack of flexibility in unaligned scenarios where no spatial correspondence exists, and (2) a rigid addition operation that limits the potential interaction between the condition and the image tag.

[0124] In contrast, MaTe directly concatenates conditional tags with image tags [C] M C I C D This new operation enables flexible tag interaction through DiT's multimodal attention mechanism, allowing direct relationships to be established between any pair of tags without imposing strict spatial constraints.

[0125] Cross-bias modulation

[0126] Although MaTe's unified sequence processing and multimodal attention can achieve effective labeled interactions during training, practical applications often require adjustable conditional strength at test time.

[0127] This invention achieves this by introducing a bias term into the multimodal attention calculation. Specifically, for a given intensity factor γ, this invention modifies the attention operation in Equation 3 as follows:

[0128]

[0129] Where B(γ) is an adjustment connection marker [C M C I C D The attentional bias matrix between [ ]. Given and The structure of the deviation matrix is ​​as follows:

[0130]

[0131] The structure of the bias matrix ensures that the attention patterns within each modality (main diagonal block) remain unchanged, while the interaction strength (C) between modalities remains constant. M and C I / C D This is adjusted by introducing a log(γ) bias term. During the testing phase, when γ→0... + When β is minimized, the influence of conditions can be eliminated, while when β > 1, conditional guidance is enhanced by increasing cross-modal attention weights. This approach allows for flexible control over the integration strength between image conditions and generated results without retraining the model.

[0132] Background remains blended

[0133] A simple way to preserve the background is to replace the generated background with the original background from the input image: x'⊙m+x⊙(1-m). Clearly, combining two images in this way does not produce a coherent, seamless result.

[0134] The key assumption of this invention is that, at each step of the diffusion process, the latent representation of noise is projected onto a natural image manifold, which is noise-enhanced to a certain level. Although mixing two noisy images (from the same level) may produce a result that could be outside the manifold, the next diffusion step projects the result onto the next level manifold, thereby improving inconsistency.

[0135] Therefore, at each stage, from the potential representation x t Initially, a diffusion step is performed based on the conditional prompts, resulting in a latent representation, denoted as x. genThis is different from the noisy version x of the input image obtained from the input image. in Combine them. Now use a mask m to blend these two potential representations:

[0136] x t-1 =x t-1,gen ⊙m+x t-1,in ⊙(1-m), (9)

[0137] This process is repeated. In the final step, the entire area outside the mask is replaced with the corresponding region from the input image, thus strictly preserving the background. This simple yet effective method is used to achieve foreground-background fusion without requiring training on a diffusion architecture based on inscription.

[0138] In summary, to address the limitations of previous work, this invention re-examines the necessity of using additional image encoders in material transfer tasks, aligning with cutting-edge trends in image generation research. A core objective of material transfer is to ensure that the texture properties on the surface of the target object are consistent with the visual details of the material instance. When multiple independent conditional control signals are injected in parallel without cross-modal interaction mechanisms, they often induce intermodal decoupling. This manifests as the isolated encoding of material features, geometry, and lighting information in the latent space, where each modality independently influences the generation process rather than achieving feature fusion through collaborative optimization. Such limitations often lead to imperfections such as material geometry misalignment (material texture does not conform to surface curvature) and inconsistent lighting reflections (specular highlights deviate from the light source direction). These insights, coupled with the emergence of the Diffusion Transformer (DiT) paradigm, prompted this invention to re-examine the fundamental approach to material transfer. While IP-Adapter and ControlNet can effectively represent material and structural information separately, limitations such as inconsistent feature spaces, lack of interaction modules, and information competition during generation hinder the natural fusion of material, depth, and input image in diffusion models (e.g., Figure 5 (as shown in (d)-(g)). This invention proposes prioritizing semantic alignment across multimodal features and their interaction potential within a shared latent space, thereby projecting all image conditions into a unified latent representation. This invention designs a novel MaTe architecture that processes image conditions uniformly through multimodal attention, eliminating reliance on complex control modules. The unified sequence processing and cross-bias modulation mechanism in MaTe enable the denoising network to jointly process all image signals, requiring only a lightweight low-rank adaptive LORA to enhance depth information. Extensive experiments demonstrate that MaTe achieves high-quality material generation in a zero-shot, training-free paradigm. It outperforms existing state-of-the-art methods in both visual quality and efficiency, while maintaining accurate detail alignment, thus significantly simplifying inference prerequisites.

[0139] like Figure 5 As shown, the present invention demonstrates significant advantages in processing complex materials. Figure 5 (d)-(g) are based on results generated using an IP adapter and ControlINet, while (h) and (i) are based on results generated by fine-tuning the input image and text.

[0140] Figure 6 The images demonstrate the effectiveness of the method of this invention in material transfer across various processing scenarios. It can transform textures from a single real-world image without any prior knowledge. This method can successfully extract texture information not only from antiques thousands of years old, but also handle popular computer graphics, jewelry, and fur materials, providing strong support for design work.

[0141] Further reference Figure 2 This invention simplifies the structural comparison of different material transfer methods. It does not rely on fine-tuning of image sets / single images, nor does it require additional image encoding via IP adapters. It only requires basic image information as input, rather than complex textual guidance, to obtain high-quality material transfer results.

[0142] In summary, this invention proposes a diffusion transformer method for achieving material transfer using only image conditions. The innovative contributions and key features of this invention mainly include:

[0143] 1. Simple Material Transfer Architecture: This architecture eliminates the need for additional adapters, ControlNet, inverse sampling, or model fine-tuning by integrating input images at the token level and achieving unified processing through multimodal attention in a shared latent space.

[0144] 2. Inference process without fine-tuning or manual setting of text information: This invention utilizes existing pre-trained diffusion models to infer all necessary information from input images, achieving high-quality material generation under zero-shot and no training paradigm.

[0145] 3. Multimodal attention mechanism: Information fusion is achieved by calculating the correlation between cross-modal tags, allowing interaction between different modal features (such as images, text, and depth), thereby achieving a natural fusion of material, geometric structure, and lighting information.

[0146] 4. Cross-bias modulation mechanism: By introducing a bias term into the multimodal attention calculation, controllable cross-modal interaction is achieved, thereby enabling free control between conditions. This allows for flexible control of the integration strength between image conditions and generated results without retraining the model.

[0147] Compared with the prior art, the significant advantages of the present invention are reflected in the following aspects:

[0148] Significantly reduced architectural complexity and inference time: This invention proposes a concise material transport architecture that integrates input images at the token level and achieves unified processing through multimodal attention in a shared latent space, eliminating the need for other adapters, ControlNet, inverse sampling, or model fine-tuning, thereby significantly reducing architectural complexity and inference time.

[0149] No fine-tuning or manual text setting required: This invention simplifies the inference process, eliminating the need for task-specific model training or fine-tuning, and also eliminating the need for manual text setting. It utilizes existing pre-trained diffusion models to infer all necessary information from input images, achieving high-quality material generation in a zero-shot, training-free paradigm.

[0150] Producing high-quality material transfer results: This invention constructs a diverse and extensive real-world material transfer evaluation dataset. Its method produces high-quality material transfer results with consistent detail, outperforming state-of-the-art baselines in both qualitative and quantitative analysis. Through a multimodal attention mechanism and cross-bias modulation, it achieves a natural fusion of material, geometry, and lighting information, avoiding problems such as material geometry misalignment and inconsistent lighting reflections.

[0151] Improving Model Efficiency and Performance: By enhancing the integration of depth information in the diffusion model through LoRA, this invention allows for flexible control over the impact of depth on the generation process, thereby improving model efficiency and performance. Simultaneously, unified sequence processing and cross-bias modulation mechanisms enable the denoising network to jointly process all image signals, further enhancing model performance.

[0152] Enhanced model generalization ability: Since this invention achieves high-quality material generation without a training paradigm, it has stronger generalization ability and can adapt to various materials and scenarios, providing a wider range of application prospects for digital content creation and industrial design.

[0153] This invention also provides a storage medium for storing a computer program, which, when executed, performs at least the methods described above.

[0154] This invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein the processor executes the computer program by performing at least the method described above.

[0155] This invention also provides a processor that executes a computer program, at least performing the methods described above.

[0156] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc or CD-ROM; magnetic surface memory can be disk storage or magnetic tape storage. The storage media described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable types of memory.

[0157] In the several embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0158] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0159] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0160] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0161] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0162] The methods disclosed in the several method embodiments provided by this invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0163] The features disclosed in the several product embodiments provided by this invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0164] The features disclosed in the several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0165] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various equivalent substitutions or obvious modifications can be made without departing from the concept of the present invention, and all such modifications, achieving the same performance or application, should be considered within the scope of protection of the present invention.

Claims

1. A diffusion transformer method for achieving material transfer using only image conditions, characterized in that, Includes the following steps: S1. Obtain the target image and material image, and generate the corresponding lighting image and depth image; S2. The material image, lighting image, and depth image are uniformly converted into image tags using a shared latent spatial encoder; S3. A multimodal attention mechanism using a diffusion converter model is employed to perform unified sequence processing on the image tags, and cross-modal interaction is achieved through cross-bias modulation. S4. Enhance the control of depth image labeling over the generation process using a low-rank adaptive module; S5. Perform background-preserving blending operation to blend the foreground generated during diffusion with the background of the target image in the latent space; S6. Generate an output image that inherits the target image structure and material texture based on the processed markers.

2. The method as described in claim 1, characterized in that, In step S1, generating the illumination image includes: Extract the foreground mask from the target image; Convert the foreground area to grayscale to eliminate interference from the underlying color while preserving geometric shadows; By blending the grayscale foreground with the background area of ​​the target image, a lighting image decoupled from material interference is generated.

3. The method as described in claim 1 or 2, characterized in that, In step S3, the unified sequence processing includes: The material image markers, lighting image markers, and depth image markers are directly concatenated into a joint sequence; The correlation between any markers in the joint sequence is calculated using the multimodal attention block of the diffusion transformer; The cross-bias modulation controls the cross-modal attention weights through an adjustable logarithmic bias matrix, which satisfies: Main diagonal blocks maintain in-modal attention invariance; Off-diagonal blocks adjust the interaction intensity between material markers and lighting / depth markers using a logarithmic intensity factor.

4. The method according to any one of claims 1 to 3, characterized in that, In step S3, the multimodal attention mechanism specifically includes: In the diffusion converter block, spatial information is preserved through rotational position encoding; Project the concatenated joint sequence into a query, key, and value vector; Material, lighting, and depth features are fused based on cross-modal attention weights.

5. The method according to any one of claims 1 to 4, characterized in that, In step S3, the cross-bias modulation dynamically adjusts the condition strength during the testing phase: When the logarithmic intensity factor approaches negative infinity, cross-modal interaction between material markers and lighting / depth markers is suppressed; When the logarithmic intensity factor is greater than zero, the effect of the material marker on the lighting / depth marker is enhanced.

6. The method according to any one of claims 1 to 5, characterized in that, The training and optimization of the diffusion converter model adopts a conditional flow matching objective function, and the corrected flow velocity is estimated through a noise prediction network.

7. The method according to any one of claims 1-6, characterized in that, In step S4, the operation of the low-rank adaptive module includes: Freeze the weight matrix of the pre-trained diffusion transformer; Introducing trainable low-rank matrices to enhance the features of deep image labels; The influence of depth information on the generation process is adjusted by using hyperparameter weights.

8. The method according to any one of claims 1-7, characterized in that, In step S5, the background-preserving blending operation includes: At each step of the diffusion process, the conditionally generated latent noise representation is fused with the corresponding noise version of the target image using a foreground mask weighted fusion. In the final step, the foreground region of the generated image is directly stitched together with the original background region of the target image.

9. The method according to any one of claims 1-8, characterized in that, The method achieves material transfer without model fine-tuning or text prompts, including: An image encoder that reuses a pre-trained diffusion transformer processes all input images; Replace additional adapters or control networks with unified sequence processing and multimodal attention mechanisms.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the diffusion transformer method for material transfer using only image conditions as described in any one of claims 1-9.

Citation Information

Cited By

  • Model training method, effect picture generation method, device and equipment

    CN122223199A