An adaptive weight migration method and device for diffusion model channel expansion

CN122737533APending Publication Date: 2026-09-11EVERYTHING MIRROR (BEIJING) COMPUTER SYST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610919807.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0005]本公开实施例的目的是提供一种扩散模型通道扩展的自适应权重迁移方法及装置,以解决预训练知识崩塌、参数冗余和训练成本高的技术问题

Benefits of technology

[0029]In this way, by precisely inheriting the parameters of the original convolutional kernel to the front channels of the new convolutional kernel, the expanded diffusion model can retain the diffusion denoising network's ability to process the input of the front channels, avoiding knowledge collapse caused by random parameter initialization and achieving zero-sample distribution preservation. By independently cloning the parameters of the original convolutional kernel to the back channels of the new convolutional kernel, the initial response of the new channels can be kept symmetrical with that of the original channels, thus ensuring that the diffusion denoising network can produce stable and predictable outputs when facing new input channels. At the same time, only the channel dimensions of the input convolutional layer are modified, without adding additional encoder branches or introducing a large number of additional parameters, avoiding parameter redundancy. Furthermore, the entire parameter processing process only involves parameter inheritance and cloning operations, without any additional training data, gradient calculations, or optimization iterations, achieving zero-training-cost model channel expansion and avoiding the high computational cost of training from scratch. Meanwhile, due to the precise inheritance and symmetrical initialization of parameters, the expanded diffusion model can effectively reduce the caching of invalid intermediate features during inference, reduce memory usage and data transmission volume, improve hardware processing speed, and meet the requirements of rapid deployment, low power consumption, and real-time response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122737533A_ABST
    Figure CN122737533A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and device for adaptive weight migration of diffusion model channel expansion. In the method, first, the original convolution kernel parameters of the input convolution layer in the diffusion denoising network of the pre-trained diffusion model are obtained, the original convolution kernel parameters being original input weights and having a first input channel number; then, in response to the latent space channel number of the pre-trained diffusion model being expanded from the first channel number to a second channel number, new convolution kernel parameters are created, the new convolution kernel parameters having a second input channel number which is an integer multiple of the first input channel number; then, the original convolution kernel parameters are inherited to the front channels of the new convolution kernel parameters, and the original convolution kernel parameters are cloned to the rear channels of the new convolution kernel parameters; finally, the new convolution kernel parameters are loaded to the input convolution layer to replace the original convolution kernel parameters, obtaining an expanded diffusion model. In this way, through the mode of accurate inheritance of the front channels and symmetric cloning of the rear channels, zero-shot distribution preservation can be achieved without additional training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, specifically to an adaptive weight transfer method and apparatus for diffusion model channel expansion. Background Technology

[0002] With the widespread application of diffusion models in tasks such as image processing and natural language understanding, the scale and complexity of these models are constantly increasing. In practical applications, it is often necessary to adjust the structure or extend the capabilities of existing models according to new task requirements.

[0003] However, when it is necessary to expand the latent space channels of a diffusion model from the standard dimension to a higher dimension to support multimodal input, if a random initialization strategy is used to assign weights to the new channels, the output distribution of the expanded model will be severely shifted during the zero-shot inference stage, leading to the collapse of pre-trained knowledge. Furthermore, methods such as full-parameter fine-tuning, low-rank adaptation (LoRA), and adapters introduce a large number of additional parameters, resulting in parameter redundancy and making them unsuitable for scenarios where the channel dimension changes. If the pre-trained model is abandoned and trained from scratch, the hardware must repeatedly cache invalid intermediate features during inference, increasing memory usage and data transfer volume, resulting in high training costs.

[0004] Therefore, there is an urgent need for an adaptive parameter processing scheme for diffusion model channel expansion to effectively solve the technical problems of pre-trained knowledge collapse, parameter redundancy, and high training costs. Summary of the Invention

[0005] The purpose of this disclosure is to provide an adaptive weight transfer method and apparatus for diffusion model channel expansion, so as to solve the technical problems of pre-training knowledge collapse, parameter redundancy and high training cost.

[0006] In a first aspect, embodiments of this disclosure provide an adaptive weight transfer method for channel expansion of a diffusion model. The method includes: obtaining the original convolutional kernel parameters of the input convolutional layer in a diffusion denoising network of a pre-trained diffusion model, wherein the original convolutional kernel parameters have a first number of input channels and are the original input weights; in response to the expansion of the latent space channel number of the pre-trained diffusion model from the first number of channels to a second number of channels, creating new convolutional kernel parameters, wherein the new convolutional kernel parameters have a second number of input channels, the second number of input channels being an integer multiple of the first number of input channels; inheriting the original convolutional kernel parameters to the front channels of the new convolutional kernel parameters, the number of channels in the front channels being the first number of input channels; cloning the original convolutional kernel parameters to the rear channels of the new convolutional kernel parameters, the number of channels in the rear channels being the difference between the second number of input channels and the first number of input channels; and loading the new convolutional kernel parameters into the input convolutional layer, replacing the original convolutional kernel parameters, to obtain the expanded diffusion model.

[0007] In an optional implementation, the method further includes: obtaining bias parameters corresponding to the original convolution kernel parameters, the bias parameters being used to adjust the bias of each output channel of the input convolutional layer; and using the bias parameters as bias parameters for the new convolution kernel parameters.

[0008] In one optional implementation, the diffusion denoising network further includes an output convolutional layer and a backbone network. Compared to the pre-trained diffusion model, the number of input channels in the input convolutional layer of the expanded diffusion model's diffusion denoising network changes, or the number of input channels in the input convolutional layer and the number of output channels in the output convolutional layer changes, while the structure and parameters of the backbone network in the diffusion denoising network remain unchanged. The backbone network includes a downsampling module, an intermediate module, an upsampling module, and an attention module.

[0009] In an optional implementation, when the number of input channels of the input convolutional layer and the number of output channels of the output convolutional layer change, the method further includes: obtaining the original output weights of the output convolutional layer in the diffusion denoising network, the original output weights having a first number of output channels; in response to the expansion of the number of latent space channels from the first number of channels to a second number of channels, creating new output weights, the new output weights having a second number of output channels, the second number of output channels being an integer multiple of the first number of output channels; inheriting the original output weights to the front channels of the new output weights, the number of channels in the front channels being the first number of output channels; cloning the original output weights to the rear channels of the new output weights, the number of channels in the rear channels being the difference between the second number of output channels and the first number of output channels; and loading the new output weights into the output convolutional layer, replacing the original output weights, to obtain the expanded diffusion model.

[0010] In one alternative implementation, after loading new output weights into the output convolutional layer to replace the original output weights, the method further includes: jointly fine-tuning at least two of the input convolutional layer, the output convolutional layer, and the backbone network.

[0011] In one optional implementation, the integer multiple is N, where N is an integer greater than 1; cloning the original convolution kernel parameters to the rear channels of the new convolution kernel parameters includes: when N equals 2, cloning the original convolution kernel parameters to the rear channels, where the number of channels in the rear channels is the number of the first input channels; when N is greater than 2, repeatedly cloning the original convolution kernel parameters to all rear channels; wherein the number of channels filled in each cloning is the number of the first input channels.

[0012] In one alternative implementation, the first number of channels is 4 and the second number of channels is 8.

[0013] In one optional implementation, inheriting the original convolution kernel parameters to the front channel of the new convolution kernel parameters includes: inheriting the original convolution kernel parameters to the front channel of the new convolution kernel parameters through a deep copy operation; and / or, cloning the original convolution kernel parameters to the rear channel of the new convolution kernel parameters includes: cloning the original convolution kernel parameters to the rear channel of the new convolution kernel parameters through a deep copy operation, wherein the deep copy operation is an operation of creating a copy that is completely identical to the original convolution kernel parameters and has independent storage space.

[0014] In an optional implementation, the method further includes: performing zero-shot distribution consistency verification on the expanded diffusion model, wherein the zero-shot distribution consistency verification includes at least one of the following: the mean square error (MSE) of the output feature maps of the expanded diffusion model and the pre-trained diffusion model under the same input conditions is less than 10. -6 ; and / or, the Fraser initial distance (FID) increment of the extended diffusion model is less than 1.0; and / or, the learned perceptual image patch similarity (LPIPS) of the extended diffusion model is less than 0.05.

[0015] In one alternative implementation, the original convolution kernel parameters are used to process a single-modal image, and the new convolution kernel parameters are used to process a multimodal image, which includes depth images and color images.

[0016] Secondly, embodiments of this disclosure provide an adaptive weight transfer apparatus for diffusion model channel expansion, the apparatus comprising: The acquisition module is used to acquire the original convolution kernel parameters of the input convolutional layer in the diffusion denoising network of the pre-trained diffusion model. The original convolution kernel parameters have the first number of input channels and are the original input weights. Create a module to create new convolutional kernel parameters in response to the expansion of the number of latent space channels of the pre-trained diffusion model from the first number of channels to the second number of channels. The new convolutional kernel parameters have a second number of input channels, which is an integer multiple of the first number of input channels. The inheritance module is used to inherit the original convolution kernel parameters to the front channels of the new convolution kernel parameters. The number of channels in the front channels is the number of the first input channels. The cloning module is used to clone the original convolution kernel parameters to the back channel of the new convolution kernel parameters. The number of channels in the back channel is the difference between the number of the second input channels and the number of the first input channels. The replacement module is used to load new convolutional kernel parameters into the input convolutional layer, replacing the original convolutional kernel parameters, and obtaining the expanded diffusion model.

[0017] In one optional implementation, the device further includes a bias module for obtaining bias parameters corresponding to the original convolution kernel parameters, the bias parameters being used to adjust the bias of each output channel of the input convolutional layer; and using the bias parameters as bias parameters for the new convolution kernel parameters.

[0018] In one optional implementation, the diffusion denoising network further includes an output convolutional layer and a backbone network. Compared to the pre-trained diffusion model, the number of input channels in the input convolutional layer of the expanded diffusion model's diffusion denoising network changes, or the number of input channels in the input convolutional layer and the number of output channels in the output convolutional layer changes, while the structure and parameters of the backbone network in the diffusion denoising network remain unchanged. The backbone network includes a downsampling module, an intermediate module, an upsampling module, and an attention module.

[0019] In an optional implementation, the apparatus further includes an extension module for obtaining the original output weights of the output convolutional layer in the diffusion denoising network, the original output weights having a first number of output channels; in response to the expansion of the number of latent space channels from the first number of channels to a second number of channels, creating new output weights, the new output weights having a second number of output channels, the second number of output channels being an integer multiple of the first number of output channels; inheriting the original output weights to the front channels of the new output weights, the number of channels in the front channels being the first number of output channels; cloning the original output weights to the rear channels of the new output weights, the number of channels in the rear channels being the difference between the second number of output channels and the first number of output channels; and loading the new output weights into the output convolutional layer, replacing the original output weights, to obtain the extended diffusion model.

[0020] In one alternative implementation, the device further includes a fine-tuning module for jointly fine-tuning at least two of the input convolutional layer, the output convolutional layer, and the backbone network.

[0021] In one optional implementation, the integer multiple is N, where N is an integer greater than 1; the cloning module is used to: clone the original convolution kernel parameters to the rear channels when N equals 2, the number of channels in the rear channels being the number of the first input channels; and to repeatedly clone the original convolution kernel parameters to all rear channels when N is greater than 2; wherein the number of channels filled in each clone is the number of the first input channels.

[0022] In one alternative implementation, the first number of channels is 4 and the second number of channels is 8.

[0023] In one optional implementation, the inheritance module is used to: inherit the original convolution kernel parameters to the front channel of the new convolution kernel parameters through a deep copy operation; and / or, the cloning module is used to: clone the original convolution kernel parameters to the rear channel of the new convolution kernel parameters through a deep copy operation; wherein, the deep copy operation is an operation to create a copy that is exactly the same as the original convolution kernel parameters and has independent storage space.

[0024] In an optional implementation, the apparatus further includes a verification module for performing zero-shot distribution consistency verification on the expanded diffusion model. Zero-shot distribution consistency verification includes at least one of the following: the mean square error (MSE) of the output feature maps of the expanded diffusion model and the pre-trained diffusion model under the same input conditions is less than 10. -6 ; and / or, the Fraser initial distance (FID) increment of the extended diffusion model is less than 1.0; and / or, the learned perceptual image patch similarity (LPIPS) of the extended diffusion model is less than 0.05.

[0025] In one alternative implementation, the original convolution kernel parameters are used to process a single-modal image, and the new convolution kernel parameters are used to process a multimodal image, which includes depth images and color images.

[0026] Thirdly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, characterized in that, when executed by a processor, the program implements the steps of the method described in the first aspect above.

[0027] Fourthly, embodiments of this disclosure provide a computing device, including: Memory, used to store computer program products; A processor for executing a computer program product stored in memory, wherein, when the computer program product is executed, it implements the method described in the first aspect above.

[0028] In this embodiment, the original convolutional kernel parameters of the input convolutional layer in the diffusion denoising network of the pre-trained diffusion model are first obtained. These original convolutional kernel parameters have a first number of input channels and are the original input weights. Then, in response to the expansion of the number of latent space channels of the pre-trained diffusion model from the first number of channels to the second number of channels, a new convolutional kernel parameter is created. This new convolutional kernel parameter has a second number of input channels, which is an integer multiple of the first number of input channels. Then, the original convolutional kernel parameter is inherited to the front channels of the new convolutional kernel parameter, where the number of channels in the front channels is the first number of input channels. The original convolutional kernel parameter is also cloned to the rear channels of the new convolutional kernel parameter, where the number of channels in the rear channels is the difference between the second number of input channels and the first number of input channels. Finally, the new convolutional kernel parameter is loaded into the input convolutional layer, replacing the original convolutional kernel parameter, to obtain the expanded diffusion model.

[0029] In this way, by precisely inheriting the parameters of the original convolutional kernel to the front channels of the new convolutional kernel, the expanded diffusion model can retain the diffusion denoising network's ability to process the input of the front channels, avoiding knowledge collapse caused by random parameter initialization and achieving zero-sample distribution preservation. By independently cloning the parameters of the original convolutional kernel to the back channels of the new convolutional kernel, the initial response of the new channels can be kept symmetrical with that of the original channels, thus ensuring that the diffusion denoising network can produce stable and predictable outputs when facing new input channels. At the same time, only the channel dimensions of the input convolutional layer are modified, without adding additional encoder branches or introducing a large number of additional parameters, avoiding parameter redundancy. Furthermore, the entire parameter processing process only involves parameter inheritance and cloning operations, without any additional training data, gradient calculations, or optimization iterations, achieving zero-training-cost model channel expansion and avoiding the high computational cost of training from scratch. Meanwhile, due to the precise inheritance and symmetrical initialization of parameters, the expanded diffusion model can effectively reduce the caching of invalid intermediate features during inference, reduce memory usage and data transmission volume, improve hardware processing speed, and meet the requirements of rapid deployment, low power consumption, and real-time response. Attached Figure Description

[0030] Figure 1 This is a schematic diagram of the architecture of a multimodal conditional image generation system provided in an embodiment of the present disclosure; Figure 2 A flowchart illustrating an adaptive weight transfer method for diffusion model channel expansion provided in this embodiment of the present disclosure; Figure 3 A schematic diagram of an adaptive parameter migration method provided in an embodiment of this disclosure; Figure 4 A flowchart illustrating another adaptive weight transfer method for diffusion model channel expansion provided in this embodiment of the present disclosure; Figure 5 A schematic diagram comparing the U-Net network architecture before and after channel expansion is provided in an embodiment of this disclosure; Figure 6 A schematic diagram of a zero-sample distribution preservation verification process provided for embodiments of this disclosure; Figure 7 A structural block diagram of an adaptive weight transfer device for diffusion model channel expansion provided in this embodiment of the present disclosure; Figure 8 This is a structural block diagram of a computing device provided in an embodiment of the present disclosure. Detailed Implementation

[0031] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. Through these descriptions, the features and advantages of the present application will become clearer and more apparent.

[0032] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments. Although various aspects of embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless specifically indicated otherwise.

[0033] Furthermore, the technical features involved in the different embodiments of this application described below can be combined with each other as long as they do not conflict with each other.

[0034] To facilitate understanding of the technical solutions of this disclosure, the application scenarios of the technical solutions provided in the embodiments of this disclosure are illustrated below.

[0035] With the development of deep learning models, diffusion models, as one of the core architectures in the current field of generative artificial intelligence, have been widely used in tasks such as image processing and natural language understanding. Latent Diffusion Models (LDMs), represented by StableDiffusion, significantly reduce computational complexity by performing denoising iterations in a low-dimensional latent space, making real-time generation of high-resolution images possible.

[0036] In practical applications, it is often necessary to adjust the structure or expand the capabilities of existing pre-trained models according to new task requirements. Therefore, the transfer and adaptation of parameters (such as original input weights and original output weights) of pre-trained diffusion models has become a key technical aspect.

[0037] Currently, mainstream parameter transfer techniques, such as full parameter fine-tuning, LoRA, adapter, and ControlNet, still have many shortcomings when transferring and adapting parameters to pre-trained diffusion models.

[0038] Full parameter fine-tuning adapts to new tasks by retraining all parameters end-to-end, but it has extremely high training costs, large storage overhead, and is prone to catastrophic forgetting. For large-scale diffusion models, it often requires hundreds of GB of GPU memory and several days of training, making it less feasible in practice. LoRA achieves efficient parameter fine-tuning by introducing a low-rank factorization matrix, but it assumes that the weight update has a low-rank structure. When the number of input or output channels changes, the dimension of the original weight matrix changes, and the low-rank factorization cannot be directly applied. Adapter inserts a lightweight bottleneck layer into the network layer for feature adaptation, but it is only suitable for cases where the network topology remains unchanged and cannot handle changes in the dimension of the first convolutional kernel. Although the above methods such as LoRA and Adapter have certain parameter efficiency advantages, they cannot be directly applied in scenarios with expanded channel dimensions. If forced to modify them, additional adaptation structures need to be introduced, which leads to parameter redundancy. Cue word learning does not modify the model parameters at all, and its ability is limited by the inherent representation space of the pre-trained model, making it unable to achieve physical adaptation to new modal inputs. ControlNet achieves conditional control by cloning parts of U-Net and adding parallel control paths, introducing hundreds of millions of trainable parameters for each control mode, resulting in severe parameter redundancy. Its design goal is not channel expansion, but rather to increase parallel control flow.

[0039] Regarding parameter transfer and adaptation for pre-trained diffusion models, such as when expanding the model's channels to accommodate more information, if a conventional random initialization strategy is used to assign weights to the new channels, the output distribution of the expanded model will be severely shifted during the zero-shot inference phase. The original pre-trained knowledge cannot be preserved, and meaningless noise output may even occur, leading to a "pre-trained knowledge collapse." Meanwhile, the aforementioned parameter transfer methods such as full parameter fine-tuning, LoRA, and Adapter, because they assume the network topology remains unchanged, cannot be directly applied to changes in the number of input channels; and ControlNet's goal is to increase parallel control paths rather than expanding the channels themselves. Furthermore, abandoning the pre-trained model and training from scratch requires massive amounts of labeled data and expensive computational resources, resulting in long training cycles, high costs, and the risk of training non-convergence or substandard generation quality.

[0040] To address the technical challenges of pre-training knowledge collapse, parameter redundancy, and high training costs in parameter transfer scenarios, this disclosure provides an adaptive weight transfer method and apparatus for channel expansion in a diffusion model. The method first obtains the original convolutional kernel parameters of the input convolutional layer in the diffusion denoising network of the pre-trained diffusion model. These original convolutional kernel parameters have a first number of input channels and serve as the original input weights. Then, in response to the expansion of the latent space channel number of the pre-trained diffusion model from the first number of channels to a second number of channels, a new convolutional kernel parameter is created. This new convolutional kernel parameter has a second number of input channels, which is an integer multiple of the first number of input channels. Next, the original convolutional kernel parameters are inherited to the front channels of the new convolutional kernel parameter, where the number of channels in the front channels is equal to the first number of input channels. The original convolutional kernel parameters are then cloned to the rear channels of the new convolutional kernel parameter, where the number of channels in the rear channels is the difference between the second and first number of input channels. Finally, the new convolutional kernel parameters are loaded into the input convolutional layer, replacing the original convolutional kernel parameters, to obtain the expanded diffusion model.

[0041] In this way, by precisely inheriting the parameters of the original convolutional kernel to the front channels of the new convolutional kernel, the expanded diffusion model can retain the diffusion denoising network's ability to process the input of the front channels, avoiding knowledge collapse caused by random parameter initialization and achieving zero-sample distribution preservation. By independently cloning the parameters of the original convolutional kernel to the back channels of the new convolutional kernel, the initial response of the new channels can be kept symmetrical with that of the original channels, thus ensuring that the diffusion denoising network can produce stable and predictable outputs when facing new input channels. At the same time, only the channel dimensions of the input convolutional layer are modified, without adding additional encoder branches or introducing a large number of additional parameters, avoiding parameter redundancy. Furthermore, the entire parameter processing process only involves parameter inheritance and cloning operations, without any additional training data, gradient calculations, or optimization iterations, achieving zero-training-cost model channel expansion and avoiding the high computational cost of training from scratch. Meanwhile, due to the precise inheritance and symmetrical initialization of parameters, the expanded diffusion model can effectively reduce the caching of invalid intermediate features during inference, reduce memory usage and data transmission volume, improve hardware processing speed, and meet the requirements of rapid deployment, low power consumption, and real-time response.

[0042] The adaptive weight transfer method for diffusion model channel expansion provided in this disclosure will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0043] The adaptive weight transfer method for diffusion model channel expansion provided in this disclosure can be applied to a multimodal conditional image generation system. This system is based on LDM and can generate target images from multimodal images. The method provided in this disclosure is mainly used to expand the channels and process the parameters of the input convolutional layer of the diffusion denoising network in the system to adapt to multimodal input. This system can be deployed in autonomous driving platforms, intelligent robot environmental perception systems, virtual reality content generation platforms, or intelligent transportation scene data synthesis systems.

[0044] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of a multimodal conditional image generation system provided in an embodiment of this disclosure. Figure 1 As shown, the multimodal conditional image generation system 10 may include an encoder 11, a latent space 12, a diffusion denoising network 13, and a decoder 14.

[0045] For example, the system 10 can receive multimodal images to be processed, such as conditional modal images (e.g., LiDAR depth rendering maps) used to provide scene geometric constraints, reference modal images (e.g., real red-green-blue (RGB) texture images) used to provide texture color references, etc., and process the multimodal images through encoder 11, latent space 12, diffusion denoising network 13 and decoder 14, and finally output a target image that fuses the geometric information of the conditional modal image and the texture information of the reference modal image.

[0046] The encoder 11 is used to encode the input multimodal images (such as conditional modal images and reference modal images) respectively, generating latent vectors (such as conditional modal latent vectors and reference modal latent vectors) corresponding to each modal image. In an optional implementation, the encoder 11 may adopt a variational autoencoder (VAE) structure, and its specific architecture can be selected according to actual application requirements. This disclosure does not limit this aspect.

[0047] The latent space 12 can be obtained by expanding the original latent space (usually 4 channels) of a pre-trained diffusion model (such as the Stable Diffusion series of LDMs), with a greater number of channels than the original latent space (e.g., expanding from 4 channels to 8 channels). This latent space 12 can be divided into multiple subspaces along the channel dimension, such as a first subspace corresponding to the conditional modality latent vector and a second subspace corresponding to the reference modality latent vector. The latent space 12 can store the conditional modality latent vector output by encoder 11 in the first subspace, store the reference modality latent vector in the second subspace, and directly concatenate the contents of the first and second subspaces along the channel dimension to form a joint latent vector.

[0048] The diffusion denoising network 13 can be a U-Net network, with the same number of input channels as the latent space 12 (e.g., both are 8 channels). The adaptive weight transfer method for diffusion model channel expansion provided in this embodiment inherits the original convolutional kernel parameters to the front channels of the new convolutional kernel and clones them to the rear channels, thereby achieving lossless expansion of the number of input convolutional layers of U-Net. This allows it to directly receive the joint latent vector output by the latent space 12, and performs multi-step denoising on the joint latent vector through internal multi-layer convolution operations, residual blocks, and self-attention mechanisms, finally outputting the denoised joint latent vector.

[0049] Decoder 14 can be a VAE decoder. Decoder 14 is used to receive the denoised joint latent vector output by the diffusion denoising network 13 and decode it to restore the target image in pixel space, that is, the target color image that fuses the geometric structure of the input conditional modality image and the texture and color of the reference modality image.

[0050] It should be noted that, Figure 1 The system architecture shown is merely an illustrative example and does not constitute a limitation on the system. In practical applications, the modules in the above system can be combined, split, or omitted according to different needs.

[0051] For example, the system may also include a text encoder and a time-step embedding module. The text encoder receives externally input textual prompts (such as scene category, object attributes, style, etc.), encodes them into text feature vectors, and inputs them into the diffusion denoising network. The time-step embedding module receives temporal information representing the current denoising time step (the step number in the denoising iteration process), encodes it into a time-step embedding vector, and inputs it into the diffusion denoising network. This guides the diffusion denoising network to consider auxiliary information such as text semantic constraints and time-step characteristics in each denoising step, generating a target image that better matches the expected content.

[0052] The above is an exemplary description of a multimodal conditional image generation system. The following section, in conjunction with the appendix... Figure 1 The system shown provides a detailed description of the adaptive weight transfer method for diffusion model channel expansion provided in this embodiment.

[0053] See appendix Figure 2 , Figure 2 This is a flowchart illustrating an adaptive weight transfer method for channel expansion in a diffusion model, provided in an embodiment of this disclosure. Figure 2 As shown, the method may include steps S201 to S205.

[0054] Step S201: Obtain the original convolution kernel parameters of the input convolutional layer in the diffusion denoising network of the pre-trained diffusion model. The original convolution kernel parameters have a first number of input channels and are the original input weights.

[0055] In one alternative implementation, before performing channel expansion on the input convolutional layers of a diffusion denoising network (such as U-Net), the original convolutional kernel parameters of the input convolutional layers in the diffusion denoising network of the pre-trained diffusion model are first obtained.

[0056] Among them, the pre-trained diffusion model refers to the diffusion model that has been trained on a large-scale dataset (such as the Stable Diffusion series of LDMs), also known as the original model. The original model usually includes a diffusion denoising network (U-Net) and related latent spatial encoders and decoders.

[0057] The original convolutional kernel parameters refer to the weights in the original convolutional layer (i.e., the original input weights). Specifically, they can be a weight parameter matrix, typically stored and manipulated in computer programs as tensors. A tensor can be understood as a multidimensional array where each value is a learnable parameter. This tensor of the original convolutional kernel parameters has the number of initial input channels (usually 4, corresponding to the number of channels in the original latent space), as well as dimensional information such as the number of output channels and the spatial dimensions of the convolutional kernel (e.g., height and width). For example, the shape of an original convolutional kernel parameter tensor can be (number of output channels, number of input channels, kernel height, kernel width).

[0058] Taking the U-Net input convolutional layer in Stable Diffusion v1.5 as an example, its original convolutional kernel parameters can be represented as a four-dimensional tensor. .in, The number of output channels (e.g., 320); =4, which is the number of the first input channels, and k is the kernel size (e.g., 3).

[0059] The original convolution kernel parameters are used to process single-modal images, that is, to receive only the latent representation of 4 channels, such as the compressed representation of a single RGB image obtained by an encoder.

[0060] It should be noted that the specific value of the first input channel number mentioned above is only for illustrative purposes. In practical applications, the number of the first input channels of the original convolution kernel parameters may vary depending on the pre-trained model, and this disclosure does not limit this.

[0061] Step S201 provides the original weight parameters for subsequent channel expansion operations, providing a data foundation for subsequent inheritance and cloning operations.

[0062] Step S202: In response to the expansion of the number of latent space channels of the pre-trained diffusion model from the first number of channels to the second number of channels, a new convolutional kernel parameter is created, which has a second number of input channels, which is an integer multiple of the first number of input channels.

[0063] In one alternative implementation, when the number of latent space channels in the pre-trained diffusion model is expanded from a first number of channels (e.g., 4) to a second number of channels (e.g., 8), the number of input channels of the input convolutional layer of the diffusion denoising network also needs to be expanded accordingly to receive the joint latent vector output from the latent space. For this purpose, a new convolutional kernel parameter tensor needs to be created to hold the expanded weights.

[0064] For example, the new convolutional kernel parameters have a second number of input channels (e.g., 8) corresponding to the second number of channels in the latent space, and its output channel number and kernel size are consistent with the original convolutional kernel parameters. The second number of input channels (i.e., the expanded number of channels) is usually an integer multiple of the first number of input channels, such as 2 times (4→8), 3 times (4→12), etc., to facilitate subsequent inheritance and cloning operations.

[0065] Taking the expansion from 4 channels to 8 channels as an example, the dimension of the new convolutional kernel parameters can be expressed as follows: .in, , which is the number of the second input channels; , is the expansion factor; The parameters k are the same as those of the original convolution kernel. The parameters of the new convolution kernel can be initialized to a tensor with all zeros during creation. The specific weight values ​​are then filled in through inheritance and cloning operations.

[0066] The new convolution kernel parameters can be used to process multimodal images. The joint latent vector formed by encoding and stitching the multimodal images input to the system will be input into the extended diffusion denoising network, and convolution operations will be performed through the new convolution kernel parameters.

[0067] For example, multimodal images may include depth images (conditional modal images) and color images (reference modal images), etc. Depth images are used to provide three-dimensional geometric information of the scene, and may be LiDAR rasterized depth rendering maps, structured light depth maps, time-of-flight (ToF) camera depth maps, binocular stereo matching depth maps, or simulation-generated depth maps, etc.; color images are used to provide texture and color information, and may be RGB texture images of real scenes, luminance-chrominance (YUV) images, hue-saturation-brightness (HSV) images, etc.

[0068] It is understood that the embodiments of this disclosure do not limit the specific source and color space of the multimodal image, as long as geometric constraints and texture references can be provided. In addition, the specific values ​​of the number of second input channels and the multiplication relationships mentioned above are only examples. In practical applications, other values ​​(such as expanding to 12 channels, 16 channels, etc.) can be selected according to the needs of potential spatial expansion, as long as the integer multiple relationship is satisfied. The embodiments of this disclosure do not limit this.

[0069] Step S202 provides a data container for subsequent weight transfer by creating new convolutional kernel parameters that match the number of channels in the expanded latent space. This gives subsequent inheritance and cloning operations a clear target, laying the foundation for maintaining the zero-sample distribution. Simultaneously, the output channel number and kernel size of the new convolutional kernel parameters are consistent with the original parameters, ensuring network structure compatibility.

[0070] Step S203: Inherit the original convolution kernel parameters to the front channels of the new convolution kernel parameters; wherein, the number of channels in the front channels is the number of the first input channels.

[0071] In one alternative implementation, after creating new convolutional kernel parameters, the original convolutional kernel parameters can be precisely inherited into the first few channels of the new convolutional kernel parameters. The first few channels refer to the first few consecutive channels along the input channel dimension in the tensor of the new convolutional kernel parameters, and their number is equal to the number of the first input channels (i.e., the number of the original input channels, for example, 4).

[0072] Through this inheritance operation, the front channels of the new convolutional kernel parameters obtain the exact same weight values ​​as the original convolutional kernel parameters, thus ensuring that when the extended diffusion denoising network uses only the front channel input (with the rear channels padded with zeros), its output is completely consistent with the original model.

[0073] For example, in a scenario scaling from 4 channels to 8 channels, the original convolution kernel parameters... The shape is New convolution kernel parameters The shape is Through inheritance, the values ​​in the original weight matrix can be copied element-wise to the corresponding positions in the first four input channels (e.g., channel indices 0 to 3) of the new convolutional kernel. .

[0074] In one alternative implementation, the original convolution kernel parameters can be inherited to the front channel of the new convolution kernel parameters through a deep copy operation.

[0075] Deep copy, also known as deep copy, refers to creating a copy in computer memory that is identical in content to the original convolution kernel parameters (i.e., the original tensor) but with independent storage space. Unlike shallow copy, deep copy not only copies the references to the data itself but also recursively copies all sub-objects, ensuring that the new tensor is completely separated from the original tensor in memory. Through deep copy, the new convolution kernel parameters obtain independent memory space from the original convolution kernel parameters, avoiding interference from subsequent gradient updates or modifications.

[0076] For example, in Python's deep learning framework PyTorch, deep copying can be achieved using the `clone()` method, i.e., `W_new[:, 0:4, :, :] = W_orig.clone()`. This method returns a tensor with the same values ​​as `W_orig` but with independent storage.

[0077] It should be noted that the specific implementation of the inheritance operation can be set according to the actual deployment environment, as long as the values ​​of the original convolution kernel parameters can be completely transferred to the front channel of the new convolution kernel. This embodiment does not limit this.

[0078] Step S203 inherits the original convolutional kernel parameters precisely to the front channels of the new convolutional kernel, ensuring that the output of the input convolutional layer in the expanded model is element-wise equal to that of the original model when only the front channel input is used (i.e., the back channel input is zero). This property is called "zero-shot compatibility," which fully preserves the pre-trained model's ability to process the original input, avoiding knowledge collapse caused by random initialization. Simultaneously, deep copying ensures the independence of the old and new parameters, providing a clean and isolated environment for subsequent fine-tuning or gradient updates.

[0079] Step S204: Clone the original convolution kernel parameters to the rear channel of the new convolution kernel parameters; wherein the number of channels in the rear channel is the difference between the number of the second input channels and the number of the first input channels.

[0080] After inheriting the original convolution kernel parameters to the front channels, it is also necessary to initialize the back channels of the new convolution kernel parameters. For example, the original convolution kernel parameters can be cloned independently and filled into the back channel region of the new convolution kernel parameters. The back channels refer to the last several consecutive channels along the input channel dimension in the tensor of the new convolution kernel parameters. Their number is equal to the difference between the number of the second input channels and the number of the first input channels. For example, when expanding from 4 channels to 8 channels, the number of back channels is 8-4=4.

[0081] Through this cloning operation, the new channel obtains the exact same initial weight values ​​as the original channel, thus giving the new channel a response characteristic that is symmetrical to the original channel in the initial state.

[0082] For example, in a scenario expanding from 4 channels to 8 channels, with 4 rear channels, the cloning operation can be represented as follows: The cloning operation allows the values ​​in the original weight matrix to be inherited to the positions corresponding to the last four input channels (e.g., channel indices 4 to 7) of the new convolutional kernel.

[0083] In one alternative implementation, a deep copy operation can be used to clone the original convolutional kernel parameters into the later channels of the new convolutional kernel parameters. This ensures that the later channels of the new convolutional kernel parameters are independent of the original convolutional kernel parameters in memory space, avoiding interference between subsequent gradient updates. For example, in PyTorch, this can be achieved using W_new[:, 4:8, :, :] = W_orig.clone().

[0084] In one optional implementation, the number of second input channels is N times the number of first input channels, where N is an integer greater than 1. Further, when N equals 2, the original convolutional kernel parameters can be directly cloned to the subsequent channels, with the number of channels in the subsequent channels being the number of first input channels; when N is greater than 2, the original convolutional kernel parameters need to be repeatedly cloned to all subsequent channels; wherein the number of channels filled each time is the number of first input channels, until the subsequent channels are filled.

[0085] For example, when N is greater than 2, the rear channels are divided into (N-1) expansion blocks along the channel dimension, and the number of channels in each expansion block is equal to the number of the first input channels. In this case, the original convolution kernel parameters need to be cyclically cloned and filled into each expansion block.

[0086] Specifically, for the i-th expansion block (i=1,2,…,N-1), the following cloning operation can be performed: .in, For the new convolution kernel parameters, These are the original convolution kernel parameters. This is the number of the first input channels. For extended block indexes (values ​​range from 1 to N-1), : represents the entire range on the corresponding dimension.

[0087] The cloning operation described above means that the parameters of the original convolutional kernel are precisely inherited to the channel positions corresponding to the i-th expansion block of the new convolutional kernel, where the number of channels in each expansion block is... Through the aforementioned cyclic cloning operation, the original weights can be gradually filled into all the back channels.

[0088] For example, when expanding from 4 channels to 12 channels (N=3), the number of rear channels is 8. The original 4-channel weights need to be cloned twice and filled into channels 4 to 7 (channel indices 3 to 6) and channels 8 to 12 (channel indices 7 to 11), respectively.

[0089] It should be noted that the specific implementation method of repeated cloning (such as loop cloning, using repeat, etc.) can be selected according to actual needs, as long as all channels in the later channels obtain the values ​​from the original convolution kernel parameters. This disclosure does not limit this.

[0090] Step S204 clones the original convolutional kernel parameters independently into the later channels of the new convolutional kernel, ensuring that the new channels have the exact same response characteristics as the original channels in the initial state. This guarantees that the expanded model can produce stable and predictable outputs when faced with the new input channels. Furthermore, since the initial response of the new channels is symmetrical to that of the original channels, it provides a good initialization prior for subsequent optional fine-tuning, helping to accelerate convergence. In addition, the repeated cloning mechanism allows this method to flexibly adapt to channel expansion requirements of any multiple.

[0091] For example, the inheritance and cloning operations shown in steps S203 to S204 above can be referred to the appendix. Figure 3 , Figure 3 This is a schematic diagram of an adaptive parameter migration method provided in an embodiment of the present disclosure.

[0092] like Figure 3 As shown, the original convolution kernel parameters are first obtained. , shape is Then, the original weights are inherited to the first 4 channels of the new convolutional kernel through inheritance, and the original weights are cloned to the last 4 channels of the new convolutional kernel through symmetric cloning (the inheritance and cloning operations can use depth copying), finally obtaining the parameters of the expanded new convolutional kernel. , shape is The first four channels are 100% compatible with the original model, while the last four channels have symmetrical initial responses.

[0093] After completing the construction of the new convolution kernel parameters, step S205 can be executed.

[0094] Step S205: Load the new convolution kernel parameters into the input convolutional layer, replacing the original convolution kernel parameters, to obtain the expanded diffusion model.

[0095] In one alternative implementation, after completing the inheritance and cloning operations of the front and back channels of the new convolutional kernel parameters, the new convolutional kernel parameters can be loaded into the input convolutional layer of the diffusion denoising network, replacing the original convolutional kernel parameters.

[0096] This replacement operation can be understood as: changing the weight parameters of the input convolutional layer from the original... Updated to After the replacement, the number of input channels in the input convolutional layer becomes the second number of channels (e.g., 8), while the number of output channels and the kernel size remain unchanged, thus obtaining the expanded diffusion model. In this way, the diffusion denoising network of the expanded diffusion model can directly receive the joint latent vector with the second number of channels (e.g., 8 channels) as input.

[0097] Step S205 loads the new convolutional kernel parameters and replaces the original weight parameters to obtain the expanded diffusion model. The diffusion denoising network of this model is perfectly matched to the expanded latent space in terms of the number of input channels, enabling it to directly receive the joint latent vector. Furthermore, since the original weights are retained in the front channels and the rear channels are reasonably initialized through symmetric cloning, the expanded model maintains a consistent output distribution with the original model during zero-shot inference (with zeros padded in the rear channels), thus achieving complete transfer of pre-trained knowledge. Simultaneously, since the entire parameter processing requires no training data or gradient calculations, the expanded model can be immediately used for inference, meeting the need for rapid deployment.

[0098] In the adaptive weight transfer method for diffusion model channel expansion provided in this embodiment, in addition to transferring the convolution kernel weights, the bias parameters of the convolution layer can also be processed accordingly to maintain the integrity of the model output distribution.

[0099] In an optional implementation, the bias parameters corresponding to the original convolution kernel parameters can also be obtained and used as the bias parameters for the new convolution kernel parameters. These bias parameters are used to adjust the bias of each output channel of the input convolutional layer.

[0100] For example, when obtaining the original convolutional kernel parameters, their corresponding bias parameters (if they exist) can be obtained simultaneously. Taking the U-Net input convolutional layer of Stable Diffusion v1.5 as an example, this layer contains bias parameters, the shape of which is... (Number of output channels, such as 320).

[0101] Furthermore, after creating the new convolutional kernel parameters, the original bias parameters can be used as the bias parameters for the new convolutional kernel parameters. The inheritance of bias parameters can be represented as: .

[0102] Since the bias parameter is independent of the number of input channels, no expansion or transformation is needed; it can be directly inherited. If the original convolutional layer does not contain a bias parameter (e.g., the bias is set to False), this operation is unnecessary.

[0103] By inheriting the bias parameters of the original convolutional layers, the bias distribution of the extended model on the output channels is made completely consistent with that of the original model, further ensuring the integrity of the output distribution during zero-shot inference. Simultaneously, in conjunction with the precise inheritance of the weight parameters, bias inheritance ensures that the extended model's output is element-wise equal to that of the original model when only the front-channel input is used.

[0104] To gain a deeper understanding of the theoretical basis of the above parameter processing methods, the principle of the front and rear channel weight allocation strategy will be explained below.

[0105] Two-dimensional convolution operations are linearly decomposable along the input channel dimension. For the input feature map... and convolution kernel The convolution output can be expressed as follows (1): (1) in, Input the number of channels. Number of output channels The spatial dimensions of the feature map. The kernel size is [size]. This indicates that the kernel corresponding to the first convolution kernel... The weight submatrix of each input channel, Represents the first... One channel, This represents a two-dimensional convolution operation. is the bias term. Equation (1) shows that the convolution operation can be decomposed into the sum of independent convolutions of each input channel, that is, it has linear decomposability.

[0106] When the input channels are expanded to 8, and the input feature map is composed of two 4-channel feature maps concatenated together ( ,in When ), the convolution output can be decomposed into the following equation (2): (2) in, The parameters for the expanded new convolutional kernel are 8 input channels; For the corresponding bias term. Equation (2) decomposes the convolution output into the sum of the first 4 channels and the last 4 channels.

[0107] When the last 4 channels are zero-filled, that is Then, the above equation (2) can be simplified to the following equation (3): (3) If the weights of the first 4 channels satisfy and ,in, These are the original 4-channel convolution kernel parameters. Given the original bias parameters, equation (3) above is further transformed into equation (4): (4) That is, the output of the extended diffusion model (hereinafter referred to as the extended model) is completely consistent with the original model. This mathematical equivalence is the theoretical basis for achieving "zero-sample compatibility" in the embodiments of this disclosure.

[0108] Furthermore, when the inputs to the two consecutive 4-channel subspaces are exactly the same, that is... Substituting into equation (2) above and using the exact inheritance condition, we can obtain the following equation (5): (5) This symmetric response ensures that when the input data to the two subspaces are identically distributed, the model's response is a linear scaling of the original response (considering the bias shift), rather than an unpredictable random output. Symmetric cloning provides the new channel with an initial prior that is "consistent with the original channel," significantly accelerating convergence for subsequent optional fine-tuning training.

[0109] For a more rigorous description, we can define the input feature map. The standardized form is ,in .

[0110] make This represents the convolution operation of the original model. This represents the convolution operation of the extended model. The symmetric cloning strategy of this embodiment satisfies the following two key properties: Property 1 (Compatibility): When hour, That is, the output remains unchanged when zeros are padded in the rear channel.

[0111] Property 2 (Symmetry): When hour, That is, a linear transformation in which the output is the original output when the inputs of the two subspaces are the same.

[0112] These two properties together form the theoretical foundation of the weight transfer strategy in the embodiments of this disclosure.

[0113] Furthermore, the technical goal of weight transfer is to achieve "zero-sample distribution preservation," that is, after weight transfer is completed and before any additional training is performed, the generated output distribution of the extended model remains approximately consistent with that of the original model.

[0114] In one possible implementation, zero-shot distribution consistency verification can be performed on the extended diffusion model. This zero-shot distribution consistency verification may include: the Fréchet Inception Distance (FID) increment of the extended diffusion model being less than 1.0; and / or, the Learned Perceptual Image Patch Similarity (LPIPS) of the extended diffusion model being less than 0.05.

[0115] For example, let the original model parameters be... The extended model parameters are For data distribution Potential representation of samples The conditional generation distribution of the original model is defined as follows: The conditional generation distribution of the extended model is .in To be The representation is embedded in an 8-channel space. At this point, the zero-sample distribution preservation requirement is given by equation (6): (6) In the U-Net architecture of the diffusion model, the first convolutional layer typically undergoes group normalization (GroupNorm), activation functions, attention mechanisms, and multiple downsampling / upsampling operations. The weight transfer strategy provided in this disclosure ensures that the output feature map of the first convolutional layer is completely consistent with the original model. Since the parameters of subsequent network layers do not change (the number of channels remains unchanged), the input distribution of these layers is completely consistent with the input distribution of the original model, thus preserving the output distribution of the entire network.

[0116] More strictly, make This represents the original U-Net network. This represents the expanded U-Net network. For the first convolutional layer... and subsequent networks We have the following equations (7) and (8): (7) (8) in, and These represent the original and expanded input convolutional layers, respectively.

[0117] because The parameters and structure remain unchanged, and it is guaranteed that... Therefore, we can obtain the following equation (9): (9) This equation (9) holds strictly in the sense that U-Net is a deterministic function, providing a strict mathematical guarantee for "zero sample distribution preservation".

[0118] Furthermore, in practical evaluation, metrics such as FID and LPIPS can be used as quantitative measures of distribution preservation. FID measures the distance between the generated image distribution and the real image distribution; a smaller FID value indicates a closer similarity. LPIPS measures the perceptual similarity between generated image pairs, used to evaluate the consistency between the outputs of the original model and the extended model under the same input conditions.

[0119] In one alternative implementation, as previously described, the FID increment can be required during the zero-sample inference phase after weight transfer. LPIPS distance .

[0120] For example, the above weight transfer and bias inheritance operations can be implemented using the following code (using PyTorch as an example): def adaptive_weight_migration(W_orig, b_orig=None): # Parameter description: W_orig is the original convolution kernel parameter, with a shape of (C_out, 4, k, k) #b_orig is the original bias parameter, with shape (C_out,), and is optional. # Return value: W_new is the new convolution kernel parameter, with shape (C_out, 8, k, k) #b_new is the new bias parameter, with shape (C_out,). If b_orig is not None, C_out, C_in_orig, k_h, k_w = W_orig.shape # Get the dimensions of the original convolutional kernel (number of output channels, number of input channels, kernel height, kernel width) assert C_in_orig == 4 # Ensures the original input channel count is 4 (default condition) C_in_new = 8# Sets the number of expanded input channels to 8 # Step 1: Initialize the new convolutional kernel tensor W_new = torch.zeros(C_out, C_in_new, k_h, k_w, dtype=W_orig.dtype, device=W_orig.device)# All elements are initialized to 0, and the shape is (C_out, 8, k_h, k_w) # Step 2: Precise inheritance of the first 4 channels W_new[:, 0:4, :, :] = W_orig.clone() # Copies the parameters of the original convolution kernel to the first 4 input channels of the new convolution kernel. # Step 3: Symmetric cloning of the last 4 channels W_new[:, 4:8, :, :] = W_orig.clone() # Copies the parameters of the original convolution kernel to the last 4 input channels of the new convolution kernel. # Step 4: Biased Inheritance b_new = b_orig.clone() if b_orig is not None else None # If the original bias exists, copy it to the new bias; otherwise, the new bias is None return W_new, b_new # Returns the new convolution kernel parameters and the new bias parameters The code above fully demonstrates the creation of a new convolutional kernel, the precise inheritance of the front channels, the symmetric cloning of the back channels, and the inheritance of bias parameters.

[0121] The above mainly explains the channel expansion of the input convolutional layer. In practical applications, diffusion denoising networks also include output convolutional layers. When the number of latent spatial channels is expanded, the number of output channels of the output convolutional layer may also need to be expanded accordingly to match the expanded number of latent spatial channels.

[0122] In other words, compared to the pre-trained diffusion model, the number of input channels in the input convolutional layer of the expanded diffusion model changes, or the number of input channels in the input convolutional layer and the number of output channels in the output convolutional layer changes, while the structure and parameters of the backbone network in the diffusion denoising network remain unchanged; the backbone network may include downsampling modules, intermediate modules, upsampling modules, and attention modules, etc.

[0123] See appendix Figure 4 , Figure 4 This is a flowchart illustrating another adaptive weight transfer method for channel expansion of a diffusion model provided in this embodiment of the disclosure. Figure 4 As shown, when the number of input channels of the input convolutional layer and the number of output channels of the output convolutional layer change, the method may further include steps S401 to S405.

[0124] Step S401: Obtain the original output weights of the output convolutional layer in the diffusion denoising network. The original output weights have a first number of output channels.

[0125] For example, when it is necessary to expand the output convolutional layer, the original output weights (i.e., convolutional kernel parameters) of the output convolutional layer can be obtained first. The original number of output channels is usually consistent with the original number of latent spatial channels (e.g., 4).

[0126] Step S402: In response to the expansion of the number of potential space channels from the first number of channels to the second number of channels, create a new output weight. The new output weight has a second number of output channels, which is an integer multiple of the first number of output channels.

[0127] For example, when the number of latent spatial channels is expanded from a first number of channels (e.g., 4) to a second number of channels (e.g., 8), the number of output channels of the output convolutional layer also needs to be expanded accordingly to output prediction noise that matches the expanded number of latent spatial channels.

[0128] To this end, new output weights can be created with the same number of output channels as the second output weights and the same number of input channels as the original output weights. The number of second output channels is usually an integer multiple of the number of first output channels (such as 2 times, 3 times, etc.).

[0129] Step S403: Inherit the original output weights to the front channel of the new output weights. The number of channels in the front channel is the same as the number of the first output channels.

[0130] For example, the original output weights can be precisely inherited to the first few channels of the new output weights. The first few channels refer to the first few consecutive channels along the output channel dimension in the new output weight tensor, and their number is equal to the number of the first output channels (e.g., 4).

[0131] Through this inheritance operation, the front channel of the new output weights obtains the exact same value as the original output weights.

[0132] Step S404: Clone the original output weights to the rear channel of the new output weights. The number of channels in the rear channel is the difference between the number of the second output channels and the number of the first output channels.

[0133] For example, the original output weights can be cloned and populated into the back channels of the new output weights. The number of back channels is equal to the difference between the number of second output channels and the number of first output channels.

[0134] Through symmetric cloning, the newly added output channel has the same response characteristics as the original output channel in the initial state.

[0135] Step S405: Load the new output weights into the output convolutional layer to replace the original output weights, and obtain the expanded diffusion model.

[0136] For example, after the inheritance and cloning operations are completed, the new output weights can be loaded into the output convolutional layer, replacing the original output weights.

[0137] After the replacement, the number of output channels of the output convolutional layer becomes the number of the second output channels, while the number of input channels remains unchanged, thus obtaining the complete extended diffusion model.

[0138] The above embodiment expands the output convolutional layer symmetrically with the input convolutional layer, enabling the expanded diffusion model to output predictive noise matching the number of channels in the latent space. Specifically, the first channels retain the original output weights, ensuring model compatibility when the original dimensions are required for output; the last channels achieve reasonable initialization through symmetrical cloning, providing a symmetrical initial response basis for multimodal output. Furthermore, the entire expansion process involves only parameter inheritance and cloning operations, requiring no training data or gradient calculations, resulting in a minimal number of new parameters, and maintaining the same zero-sample distribution preservation characteristic as the expansion of the input convolutional layer.

[0139] The adaptive weight transfer method for diffusion model channel expansion provided in this disclosure follows the "minimum intrusion principle," meaning that modifications to the original model are limited to the minimum necessary scope. Specifically, only the input convolutional layers in the diffusion denoising network that are directly related to the latent representation of the input are modified, and optionally the output convolutional layers are modified.

[0140] See appendix Figure 5 , Figure 5 This is a schematic diagram comparing the U-Net network architecture before and after channel expansion, provided as an embodiment of this disclosure. Figure 5 As shown, the left side shows the original 4-channel input architecture, and the right side shows the expanded 8-channel input architecture.

[0141] The input latent representation of the original 4-channel input architecture on the left. After input convolutional layer Then, the signal passes sequentially through a downsampling path (containing 4 downsampling blocks, each of which is a residual block + attention (ResBlock + Attn), with the number of channels changing from 320→640→1280→1280), an intermediate block (MidBlock ResBlock + Attn), an upsampling path (containing 4 upsampling blocks, each of which is a residual block + attention + skip connection (ResBlock + Attn + Skip)), and finally through the output convolutional layer. Output The spatial resolution of the feature map changes from 64×64 to 32×32 to 16×16 to 8×8 (downsampling), and then is recovered in reverse (upsampling).

[0142] Input latent representation of the extended 8-channel architecture on the right After input convolutional layer Afterwards, the structure of the downsampling path, intermediate block, and upsampling path is completely consistent with the left side, and finally passes through the output convolutional layer. Output .

[0143] It is understood that the ResBlock, Attn, Skip, MidBlock, and channel numbers such as 320ch / 640ch / 1280ch in the figure are all standard components and parameters of the U-Net network. The embodiments disclosed in this disclosure only modify the channel dimensions of the input and output convolutional layers, and the structure and parameters of these internal modules remain unchanged.

[0144] Furthermore, from Figure 5 As shown in the diagram, the input convolutional layer The weight shape from Adjusted to .in, This refers to the number of channels in the intermediate layers of U-Net, typically 320, 640, etc., depending on the model version; This refers to the kernel size, typically 3. If it's necessary to expand the output convolutional layer... (If an 8-channel latent representation is required, then its weight shape is changed from...) Adjusted to .

[0145] In addition to the two convolutional layers mentioned above, U All downsampling modules, intermediate blocks, upsampling modules, attention modules, and the structure and parameters of the time-step embedded projection layers in the Net backbone network remain unchanged. This "modify only the beginning and end, keep the middle" design results in extremely low parameter increments.

[0146] In one optional implementation, the detailed parameter configuration of the modified U-Net network is as follows: the number of input channels is expanded from 4 to 8, and the number of output channels is expanded from 4 to 8 (optional); the number of intermediate layer channels ( =320), kernel size (k=3), stride (stride=1), padding (padding=1), bias term (bias=False), number of downsampled blocks (4), number of upsampled blocks (4), number of attention heads (8), attention dimension (40), and time step embedding dimension (1280) remain unchanged.

[0147] Taking Stable Diffusion v1.5 as an example, let the parameters of the input convolutional layer be... The parameters of the output convolutional layer are .in, ; (Standard 3x3 convolution kernel), the calculation yields: the original input convolution parameters are 320×4×3×3=11,520; the new input convolution parameters are 320×8×3×3=23,040; the increase in input layer parameters is 23,040-11,520=11,520; the increase in output layer parameters (if modified) is 23,040-11,520=11,520; the total increase in parameters is 23,040 (approximately 23K).

[0148] The parameter details of the modified layers are as follows: the original weight shape of the input convolutional layer is (320, 4, 3, 3), which is expanded to (320, 8, 3, 3); the original weight shape of the output convolutional layer is (4, 320, 3, 3), which is expanded to (8, 320, 3, 3); the normalization layer keeps the standard GroupNorm unchanged, and the number of parameters changes by 0, for a total increase of 23K parameters.

[0149] For scenarios where the output channels are expanded from 4 to 8, the total number of new parameters is approximately 46K, or about 0.5M parameter increments (considering larger-scale U-Net variants). This increment is only about 0.013% of the hundreds of millions of new parameters added by ControlNet, demonstrating a significant advantage in parameter efficiency.

[0150] To more intuitively illustrate the changes in parameter counts across different model versions, Table 1 below lists the parameter increments for common diffusion models (such as Stable Diffusion v1.4 / v1.5 / v2.1, SDXL Base, and SDXL Refiner).

[0151] Table 1

[0152] As can be seen from Table 1, regardless of the model version, the number of new parameters is far less than 0.01% of the number of parameters in the original model, which shows extremely high parameter efficiency.

[0153] Since only the channel dimensions of the input and output convolutional layers are modified, the number of channels in all feature maps within U-Net remains unchanged. In other words, the number of channels in skip connections within U-Net does not need to be adjusted, the query, key, and value projection dimensions of the attention module remain unchanged, the grouping parameters of GroupNorm remain unchanged, and the projection dimensions of the time-step embeddings also remain unchanged. Therefore, all internal components of the original model can be directly reused without any additional adaptation, fully demonstrating compatibility with pre-trained models.

[0154] Through this compatible design, the parameter processing method provided in this disclosure embodiment can maximize the retention of the pre-trained model's generation capability while expanding the number of model channels, and significantly reduce the demand for hardware resources.

[0155] Furthermore, the entire parameter processing involves only the inheritance and cloning of the original convolutional kernel parameters, requiring no training data, gradient calculations, or optimization iterations. The transfer process is extremely short, with a very small number of new parameters (not exceeding 0.5M), accounting for no more than 0.1% of the total parameters of the original diffusion model. The parameter processing takes less than 10 milliseconds in a CPU environment, truly achieving model expansion with zero additional training cost. In addition, this method is applicable to various pre-trained diffusion models, such as StableDiffusion v1.4, v1.5, v2.1, and SDXL Base, and can automatically adapt the corresponding number of intermediate layer channels by reading the U-Net configuration parameters of each model version. and kernel size It has good model compatibility.

[0156] It is understood that the core objective of this disclosure is to achieve parameter transfer under zero-training conditions. However, in actual deployment, lightweight fine-tuning can be selectively performed to further improve the multimodal fusion effect.

[0157] In one alternative implementation, at least two of the input convolutional layers, output convolutional layers, and backbone of the diffusion denoising network can be jointly fine-tuned.

[0158] For example, after completing the parameter inheritance and cloning operations, a low-rank adaptation (LoRA) module can be inserted into the backbone of the diffusion denoising network. The LoRA module achieves efficient parameter fine-tuning by adding a trainable low-rank matrix next to the attention layer or fully connected layer of the backbone network.

[0159] Furthermore, the input convolutional layer, output convolutional layer, and inserted LoRA module in the diffusion denoising network can be jointly fine-tuned. During fine-tuning, most parameters of the original network can be frozen, and only the parameters of the input convolutional layer, output convolutional layer, and LoRA module can be updated.

[0160] For example, the rank of a LoRA module can be ≤128 (e.g., 64) to balance performance and efficiency; the scaling factor alpha of the LoRA module can be set to 128; fine-tune the learning rate ≤10. -5To prevent overfitting, a lower learning rate is used; the batch size is adjusted based on GPU memory (e.g., 8), and the number of training epochs is typically 5-10 epochs to achieve rapid convergence in short cycles; the optimizer can be AdamW with weight decay of 0.01, and cosine annealing is used for learning rate scheduling with a warm-up ratio of 0.1 (i.e., linear warm-up with 10% of the steps); the gradient clipping threshold is 1.0 to prevent gradient explosion; mixed precision is achieved using fp16 to accelerate training; the fine-tuning dataset size is ≤100,000 sample pairs (e.g., 10,000-50,000 sample pairs), much smaller than training from scratch. It is understood that the above parameters are only examples and can be adjusted according to specific tasks and hardware conditions in practical applications.

[0161] Through the aforementioned joint fine-tuning, the extended model's ability to process multimodal inputs can be further optimized, resulting in more accurate generated images in terms of geometric structure and texture details. Furthermore, since the adaptive weight transfer method for diffusion model channel expansion provided in this embodiment already provides symmetric initialization priors for the new channels, the convergence speed of the fine-tuning process is significantly faster than random initialization, typically requiring only a small number of iterations to achieve satisfactory results.

[0162] The following detailed description of the adaptive weight transfer method for diffusion model channel expansion provided in this embodiment, with reference to specific implementation steps, provides a detailed explanation of the process.

[0163] It is understandable that this parameter processing procedure can be performed without any training data, relying entirely on the inheritance and cloning operations based on the parameters of the pre-trained model itself. Therefore, during the parameter processing phase, there is no need to prepare a dataset or perform data preprocessing. The entire operation only involves reading, copying, and assigning model parameters, without the need for forward or backward propagation.

[0164] If subsequent joint fine-tuning is chosen (e.g., using the LoRA module to further optimize model performance), a small amount of multimodal pairing data will be required. For example, the Microsoft COCO dataset, such as the COCO2014 dataset (approximately 128,000 images, 512×512 resolution, used as a general fine-tuning benchmark), a subset of the Large Scale Artificial Intelligence Open Network - 5 Billion Dataset (LAION-5B) (1 million to 5 million images, 512×512 resolution, used for large-scale fine-tuning), or a custom multimodal dataset (10,000 to 100,000 pairs of depth maps and color images, 512×512 resolution, used for domain adaptation).

[0165] In one optional implementation, the data preprocessing process may include: first, randomly cropping the input image to 512×512 pixels, then randomly horizontally flipping it with a probability of 0.5, and then normalizing the pixel values ​​to the [0,1] interval; the normalized image is encoded using VAE to obtain a 4×64×64 latent space representation, and then this representation is concatenated with the latent representation of another modality to form an 8×64×64 input tensor. It is understood that the aforementioned preprocessing process is merely illustrative and can be adjusted according to requirements in practical applications.

[0166] The following uses the U-Net input convolutional layer in Stable Diffusion v1.5 as an example to illustrate the adaptive weight transfer method for diffusion model channel expansion provided in this embodiment.

[0167] For example, the adaptive weight transfer method for diffusion model channel expansion can include seven steps, each of which can be adaptively adjusted according to the actual model and framework.

[0168] Step 1: Load the pre-trained model.

[0169] In one alternative implementation, a pre-trained diffusion model can be loaded first, and the code can be implemented based on the Python and PyTorch frameworks, specifically using the Hugging Faces diffuses library.

[0170] For example, this can be achieved through the following code: from diffusers import StableDiffusionPipeline # Import diffusion model pipeline import torch # Import PyTorch model_id = "runwayml / stable-diffusion-v1-5" # Specifies the name of the pre-trained model pipe = StableDiffusionPipeline.from_pretrained( # Load the pipeline) model_id, torch_dtype=torch.float16, # Use half-precision floating-point numbers to save video memory variant="fp16", # Specifies the use of the fp16 variant use_safetensors=True # Use the safe tensors format. ) pipe = pipe.to("cuda") # Move the model to the GPU unet = pipe.unet # Get U Net Networks conv_in_orig = unet.conv_in # Get U .NET input convolutional layer object This code loads the Stable Diffusion v1.5 model from Hugging Face, obtains its U-Net network, and extracts the input convolutional layers conv_in_orig. Half-precision (float16) is used here to reduce GPU memory usage.

[0171] Step 2: Extract the original convolution kernel parameters.

[0172] In one alternative implementation, the weight parameters (and bias parameters, if present) can be extracted from the original input convolutional layer, and their key attributes can be recorded.

[0173] For example, this can be achieved through the following code: W_orig = conv_in_orig.weight.data # Original convolutional kernel weight tensor, shape (C_out, 4, k, k) b_orig = conv_in_orig.bias.data if conv_in_orig.bias is not None elseNone # Bias parameter (if it exists) dtype_orig = W_orig.dtype # Records the data type of the original weights, such as torch.float16 device_orig = W_orig.device # Records the device where the original weights reside, for example, cuda:0 C_out, C_in_orig, k_h, k_w = W_orig.shape # Get the number of output channels, the number of input channels, the kernel height, and the kernel width. Where C_out is the number of output channels (usually 320), C_in_orig = 4, k_h = k_w = 3 (kernel size), and W_orig is the original kernel parameter tensor.

[0174] Step 3: Create a new convolutional layer.

[0175] In one alternative implementation, a new convolutional layer can be created with 8 input channels (the second number of channels) and the same number of output channels and kernel size as the original layer.

[0176] For example, this can be achieved through the following code: import torch.nn as nn # Import the neural network module C_in_new = 8# Number of expanded input channels # Create a new convolutional layer with 8 input channels; other parameters are the same as the original layer. conv_in_new = nn.Conv2d( in_channels=C_in_new, # Set the number of input channels to 8 out_channels=C_out, # The number of output channels is the same as the original layer. kernel_size=k_h, # The kernel size is the same as the original layer. stride=conv_in_orig.stride, # The stride is the same as the original layer padding=conv_in_orig.padding, # Padding is the same as the original layer dilation=conv_in_orig.dilation, # The dilation coefficient is the same as the original layer groups=conv_in_orig.groups, # The number of grouped convolutions is the same as the original layer. bias=(b_orig is not None), # The bias is consistent with the original layer (if the original layer has a bias, then enable it). padding_mode=conv_in_orig.padding_mode, # The padding mode is the same as the original layer. dtype=dtype_orig, # The data type is the same as the original layer. device=device_orig # The device is the same as the original layer ) The weights and biases of the new convolutional layer conv_in_new have not yet been assigned valid values ​​(currently they are randomly initialized or default values), and will be overridden in subsequent steps.

[0177] Step 4: Perform adaptive weight transfer.

[0178] In one alternative implementation, the original convolutional kernel weights can be copied to the front channel of the new convolutional kernel (exact inheritance) and cloned to the back channel (symmetric cloning), while inheriting the bias.

[0179] For example, this can be achieved through the following code: def execute_adaptive_migration(conv_orig, C_in_new=8): W_orig = conv_orig.weight.data # Original weights b_orig = conv_orig.bias.data if conv_orig.bias is not None else None # Original bias C_out, C_in_orig, k_h, k_w = W_orig.shape # Get the original convolution kernel dimensions # Step 4.1: Initialize the new convolutional kernel tensor with all zeros. W_new = torch.zeros( C_out, C_in_new, k_h, k_w, dtype=W_orig.dtype, device=W_orig.device ) # Step 4.2: Precise inheritance of the first 4 channels W_new[:, 0:C_in_orig, :, :] = W_orig.clone() # Step 4.3: Rear 4-channel symmetric cloning (supports multiple expansion) num_clone = C_in_new - C_in_orig for i in range(0, num_clone, C_in_orig): end_idx = min(i + C_in_orig, C_in_new) copy_len = end_idx - i W_new[:, i:end_idx, :, :] = W_orig[:, :copy_len, :, :].clone() # Step 4.4: Bias Inheritance b_new = b_orig.clone() if b_orig is not None else None return W_new, b_new W_new, b_new = execute_adaptive_migration(conv_in_orig, C_in_new=8) This fully implements precise inheritance and symmetric cloning, and supports cyclic cloning when the expansion factor is greater than 2. Specifically, the use of `clone()` for deep copying ensures that the new tensor is independent of the original tensor in memory, avoiding interference from subsequent gradient updates.

[0180] Step 5: Replace the original convolutional layer.

[0181] In one alternative implementation, U can be The original input convolutional layer in the .NET is replaced with a newly created convolutional layer.

[0182] For example, this can be achieved through the following code: conv_in_new.weight.data = W_new # Assign the new convolutional kernel weights to the new convolutional layer if b_new is not None: conv_in_new.bias.data = b_new # Assign the new bias to the new convolutional layer (if it exists). # Replace original conv_in with new conv_in in U-Net unet.conv_in = conv_in_new # Replace U with a new convolutional layer Raw input convolutional layer in .NET After replacing U using the steps described above The number of input convolutional layers in Net has been increased from 4 to 8, while the number of other layers remains unchanged.

[0183] Step 6: Verify the correctness of the weight transfer.

[0184] In one alternative implementation, the correctness of the transfer can be verified by comparing the outputs of the original convolutional layer and the new convolutional layer with the same input (with the last 4 channels padded with zeros).

[0185] For example, this can be achieved through the following code: def verify_migration(conv_orig, conv_new, C_in_orig=4, C_in_new=8): # Check if the weights of the front channels (input channels 0 to C_in_orig-1) are equal to the original weights element-wise. match_front = torch.allclose( conv_new.weight[:, :C_in_orig, :, :], # Weights of the front channels of the new convolutional kernel conv_orig.weight, # Original convolutional kernel weights atol=1e-6# Absolute error tolerance 1e-6 ) # Check if the weights of the back channels (input channels C_in_orig to C_in_new-1) are equal to the original weights element-wise. match_rear = torch.allclose( conv_new.weight[:, C_in_orig:C_in_new, :, :], # Weights of the back channels of the new convolutional kernel conv_orig.weight, # Original convolutional kernel weights atol=1e-6 ) # If the original convolutional layer has a bias, check if the bias of the new convolutional layer is consistent with the original bias. if conv_orig.bias is not None: match_bias = torch.allclose(conv_new.bias, conv_orig.bias, atol=1e-6) else: match_bias = True # If there is no bias initially, the bias validation will pass automatically. # Return True only if the weights and biases of the current and previous channels match. return match_front and match_rear and match_bias is_valid = verify_migration(conv_in_orig, conv_in_new) print(f"Weight migration verification: {'PASSED' if is_valid else 'FAILED'}") This verification step ensures that when the input of the last 4 channels is zero, the output of the new convolutional layer is equal to the original convolutional layer element-wise, with the mean squared error (MSE) close to 0 and the maximum absolute error (MAE) extremely low. This proves that the accurate inheritance of the front channels and the symmetrical cloning of the back channels are correct.

[0186] Step 7: Zero-shot inference verification In one alternative implementation, performing zero-sample distribution consistency verification may include the aforementioned FID incremental verification, LPIPS verification, and may also include MSE verification.

[0187] Taking the zero-padding expansion of the original 4-channel input to 8 channels as an example, during MSE validation, the output feature maps of the first convolutional layer of the model before and after expansion can be compared to verify whether the model's processing capability for the original input remains unchanged after weight transfer. If the MSE of the output feature maps of the expanded diffusion model and the pre-trained diffusion model under the same input conditions is less than 10... -6 And the maximum absolute error (MaxAE) is less than 10. -5 This indicates that the front channel is accurately inherited correctly, and the model achieves zero-sample compatibility.

[0188] For example, this can be achieved through the following code: batch_size = 1 height = width = 64 # Generate random 4-channel input to simulate the original latent representation z_orig = torch.randn(batch_size, 4, height, width, dtype=torch.float16, device="cuda") # Construct an 8-channel input, pad the last 4 channels with zeros. z_new = torch.zeros(batch_size, 8, height, width, dtype=torch.float16, device="cuda") # Input the original 4-channel input into the first 4 channels of the new tensor, leaving the last 4 channels as 0. z_new[:, 0:4, :, :] = z_orig # Compare with the output of the first convolutional layer, and disable gradient calculation. with torch.no_grad(): h_orig = conv_in_orig(z_orig) # Original convolutional layer output h_new = conv_in_new(z_new) # Output of the new convolutional layer # Calculate the mean square error of the output feature map mse = torch.mean((h_orig - h_new) 2).item() # Calculate the maximum absolute error of the output feature map max_ae = torch.max(torch.abs(h_orig - h_new)).item() print(f"First conv MSE: {mse:.8f} (target<1e-6)") print(f"First conv MaxAE: {max_ae:.8f} (target<1e-5)") print(f"Migration complete. Model ready for inference.") It should be noted that the above verification only checks the output feature map of the first convolutional layer. More comprehensive zero-shot verification (including the consistency of end-to-end generated images, distribution-level alignment, etc.) will be explained in detail below, and will not be elaborated here.

[0189] By following the seven steps above, the channels of the input convolutional layer of the diffusion denoising network can be expanded, and the expanded model can maintain the generation capability of the pre-trained model without any training data.

[0190] Furthermore, in order to fully verify whether the extended model (the extended diffusion model) retains the generative ability of the original model (the pre-trained diffusion model) without any training, multi-level zero-shot validation can be performed.

[0191] See appendix Figure 6 , Figure 6 This is a schematic diagram of a zero-sample distribution preservation verification process provided for an embodiment of this disclosure. Figure 6As shown, this verification process can be performed from four dimensions: feature-level verification, image-level verification, distribution-level verification, and numerical stability.

[0192] Verification Process 1 (Sub-node A): Feature-level Verification This verification is used to confirm the feature. Figure 1 Consistency refers to whether the first-layer convolutional output of the extended model is element-wise equal to that of the original model under the same input conditions.

[0193] For example, a random 4-channel latent representation can be generated first. Then construct an 8-channel input The last four channels are all zero, and then they are respectively... and Input the first convolution of the original model and the extended model to obtain the output feature map.

[0194] Furthermore, the MSE and MaxAE of the two can be compared, requiring MSE < 1e-6 and MaxAE < 1e-5, to verify the consistency of the feature maps. If the verification passes, child node A is determined to be "elementally equivalent".

[0195] Verification Process Two (Child Node B): Image-Level Verification This verification is used to confirm end-to-end generation consistency, that is, whether the extended model produces perceptually consistent images with the original model during the complete denoising generation process.

[0196] For example, one can first select a number (e.g., 100) standard text prompts covering diverse topics, and then use the same random seed s and text prompt p to generate 100 images using the original model, and then generate another 100 images using the extended model. When expanding the model input, the 4-channel latent representation needs to be zero-padded to expand it to 8 channels.

[0197] Furthermore, the Learned Perceptual Image Patch Similarity (LPIPS) and Structural Similarity Index Measure (SSIM) can be calculated between the two batches of images, requiring LPIPS < 0.05 and SSIM > 0.95, to verify the perceptual consistency of the generated images. If they pass, child node B is determined to be "perceptually consistent".

[0198] Verification Process 3 (Child Node C): Distributed-Level Verification This verification is used to confirm distribution-level consistency, that is, whether the distribution of images generated by the extended model is consistent with that of the original model.

[0199] For example, an extended model can be used to generate 5000 images (such as text prompts based on the COCO2014 validation set). The FID of the generated images is then calculated, using the COCO2014 dataset as a reference, requiring an FID < 16.0. Simultaneously, the FID of the images generated by the original model is compared, requiring an FID increment < 1.0. Furthermore, the initial score (InceptionScore, IS) retention rate can be verified to be greater than 99%, thus verifying the consistency of the distribution of the images generated by the extended model with the original model through the aforementioned metrics. If successful, child node C is determined to be "distribution aligned".

[0200] Verification Process 4 (Child Node D): Numerical Stability Verification This verification is used to confirm the numerical stability and memory safety of the extended model under various input conditions.

[0201] For example, forward propagation can be performed on inputs of different scales (such as latent representations of different resolutions) to ensure that the output does not contain non-numerical (NaN) or infinity (Inf); inputs of different data types (such as fp16 and fp32) can be tested to ensure consistent results; extreme value inputs (such as all zeros or all maxima) can be tested to ensure no abnormal output, and continuous inference can be performed more than 1000 times to check for memory leaks. If all tests are passed, the child node D is determined to be "numerically stable".

[0202] Furthermore, if all four sub-nodes pass the verification, it can be determined that "zero-sample distribution preservation verification has passed." If any verification fails, the diagnostic feedback process can be triggered (e.g., Figure 6 (As indicated by the dashed arrow in the middle), this indicates whether the weight transfer step has been executed correctly.

[0203] Through the above verification process, it can be confirmed that the expanded diffusion model retains the generative capability of the original model even without training, achieving zero-sample distribution preservation. Simultaneously, the verification system can promptly identify issues in weight transfer, ensuring model reliability. It is understood that the above verification process is only an illustrative example. In practical applications, some or all of the verifications can be selected according to specific needs, and this disclosure does not limit this.

[0204] This disclosure achieves significant improvements in pre-training knowledge preservation, parameter efficiency, inference speed, and optional fine-tuning convergence through a dual-path weight transfer strategy of precise inheritance and symmetric cloning, a zero-training-cost parameter processing method, and a minimally invasive convolutional layer adaptation design. The following, combined with specific experimental data and theoretical analysis, illustrates the beneficial effects of the adaptive weight transfer method for diffusion model channel expansion provided in this disclosure.

[0205] Regarding the preservation of pre-trained knowledge, by precisely inheriting the parameters of the original convolutional kernel to the front channels of the new convolutional kernel, when the input to the rear channels is zero, the output of the expanded model is completely element-wise equal to that of the original model, thus fully preserving the model's ability to process the original input. Experiments show that, compared to the random initialization strategy, the zero-sample FID increment of this embodiment is reduced from 273.4 to 0.1, LPIPS is reduced from 0.847 to 0.003, IS retention rate is 99.2%, CLIP score retention rate is 99.7%, and pre-trained knowledge retention rate exceeds 99%.

[0206] Regarding zero additional training cost, the entire parameter processing involves only parameter inheritance and cloning operations, requiring no training data, gradient calculation, or optimization iteration. The transfer process takes no more than 10 milliseconds in a CPU environment and no more than 1 millisecond in a GPU environment, achieving true zero-cost model expansion. In contrast, training from scratch requires high costs and several weeks, while the training cost of this disclosed embodiment is zero.

[0207] In terms of parameter efficiency, only the channel dimensions of the input and output convolutional layers are modified, resulting in a very small number of new parameters. For example, when expanding from 4 channels to 8 channels, only about 0.5M parameters are added, which accounts for only about 0.06% of the original model parameters. The storage increase is only about 2MB, which has an order-of-magnitude advantage in parameter efficiency. It is especially suitable for autonomous driving applications deployed on edge devices (such as in-vehicle computing units), where the parameter increase is negligible and will not increase model loading time or inference memory usage.

[0208] Regarding inference speed, since only the first and last convolutional layers are modified without introducing additional computational branches, the inference speed of the expanded model is almost identical to that of the original model. Real-world testing shows that the inference speed of this embodiment is only about 2% lower than the original model, while the ControlNet solution is 41% slower. This embodiment demonstrates significant advantages in scenarios with high real-time requirements.

[0209] In the optional fine-tuning scenario, the symmetric cloning initialization strategy provides the new channels with the same initial prior as the original channels, making the subsequent fine-tuning convergence speed 5-10 times faster than random initialization, resulting in better FID metrics and a more stable training process.

[0210] In summary, the adaptive weight transfer method for diffusion model channel expansion provided in this disclosure can fully retain the generation capability of the pre-trained model with zero training cost, add very few parameters, achieve almost lossless inference speed, and provide a good initialization basis for optional fine-tuning, fully meeting the engineering requirements of rapid deployment, low power consumption, and real-time response.

[0211] The above is a description of the adaptive weight transfer method for diffusion model channel expansion provided in the embodiments of this disclosure.

[0212] It is understood that all the above-mentioned optional technical solutions can be combined arbitrarily to form the optional embodiments of this disclosure, and will not be described in detail here.

[0213] Based on the same concept, this invention also provides an adaptive weight transfer device for diffusion model channel expansion. Since the principle by which this device solves the problem is similar to the aforementioned adaptive weight transfer method for diffusion model channel expansion, the implementation of this device can refer to the implementation of the aforementioned adaptive weight transfer method for diffusion model channel expansion; repeated details will not be elaborated further.

[0214] See appendix Figure 7 , Figure 7 This is a structural block diagram of an adaptive weight transfer device for diffusion model channel expansion provided in an embodiment of this disclosure. Figure 7 As shown, the adaptive weight transfer device 700 for diffusion model channel expansion may include: an acquisition module 701, a creation module 702, an inheritance module 703, a cloning module 704, and a replacement module 705. Among them, The acquisition module 701 is used to acquire the original convolution kernel parameters of the input convolutional layer in the diffusion denoising network of the pre-trained diffusion model. The original convolution kernel parameters have a first number of input channels and are the original input weights. Create module 702 to create new convolutional kernel parameters in response to the expansion of the number of latent space channels of the pre-trained diffusion model from the first number of channels to the second number of channels. The new convolutional kernel parameters have a second number of input channels, which is an integer multiple of the first number of input channels. Inheritance module 703 is used to inherit the original convolution kernel parameters to the front channels of the new convolution kernel parameters, and the number of channels in the front channels is the number of the first input channels; The cloning module 704 is used to clone the original convolution kernel parameters to the rear channel of the new convolution kernel parameters, wherein the number of channels in the rear channel is the difference between the number of the second input channels and the number of the first input channels; Replacement module 705 is used to load new convolutional kernel parameters into the input convolutional layer, replacing the original convolutional kernel parameters, to obtain the expanded diffusion model.

[0215] In one alternative implementation, such as Figure 7 As shown, the device also includes a bias module 706, which is used to obtain the bias parameters corresponding to the original convolution kernel parameters. The bias parameters are used to adjust the bias of each output channel of the input convolutional layer; and the bias parameters are used as the bias parameters of the new convolution kernel parameters.

[0216] In one optional implementation, the diffusion denoising network further includes an output convolutional layer and a backbone network. Compared to the pre-trained diffusion model, the number of input channels in the input convolutional layer of the expanded diffusion model's diffusion denoising network changes, or the number of input channels in the input convolutional layer and the number of output channels in the output convolutional layer changes, while the structure and parameters of the backbone network in the diffusion denoising network remain unchanged. The backbone network includes a downsampling module, an intermediate module, an upsampling module, and an attention module.

[0217] In one alternative implementation, such as Figure 7 As shown, the device also includes an extension module 707, used to obtain the original output weights of the output convolutional layer in the diffusion denoising network, the original output weights having a first number of output channels; in response to the expansion of the number of latent space channels from the first number of channels to a second number of channels, to create new output weights, the new output weights having a second number of output channels, the second number of output channels being an integer multiple of the first number of output channels; to inherit the original output weights to the front channels of the new output weights, the number of channels in the front channels being the first number of output channels; to clone the original output weights to the rear channels of the new output weights, the number of channels in the rear channels being the difference between the second number of output channels and the first number of output channels; and to load the new output weights into the output convolutional layer, replacing the original output weights, to obtain the expanded diffusion model.

[0218] In an optional implementation, the device further includes a fine-tuning module 708 for jointly fine-tuning at least two of the input convolutional layer, the output convolutional layer, and the backbone network.

[0219] In one optional implementation, the integer multiple is N, where N is an integer greater than 1; the cloning module 704 is used to: clone the original convolution kernel parameters to the rear channels when N equals 2, the number of channels in the rear channels being the number of the first input channels; and to repeatedly clone the original convolution kernel parameters to all rear channels when N is greater than 2; wherein the number of channels filled in each clone is the number of the first input channels.

[0220] In one alternative implementation, the first number of channels is 4 and the second number of channels is 8.

[0221] In one optional implementation, the inheritance module 703 is used to: inherit the original convolution kernel parameters to the front channel of the new convolution kernel parameters through a deep copy operation; and / or, the cloning module 704 is used to: clone the original convolution kernel parameters to the rear channel of the new convolution kernel parameters through a deep copy operation; wherein, the deep copy operation is an operation of creating a copy that is exactly the same as the original convolution kernel parameters and has independent storage space.

[0222] In an optional implementation, the apparatus further includes a verification module 709 for performing zero-shot distribution consistency verification on the expanded diffusion model. Zero-shot distribution consistency verification includes at least one of the following: the mean square error (MSE) of the output feature maps of the expanded diffusion model and the pre-trained diffusion model under the same input conditions is less than 10. -6 ; and / or, the Fraser initial distance (FID) increment of the extended diffusion model is less than 1.0; and / or, the learned perceptual image patch similarity (LPIPS) of the extended diffusion model is less than 0.05.

[0223] In one alternative implementation, the original convolution kernel parameters are used to process a single-modal image, and the new convolution kernel parameters are used to process a multimodal image, which includes depth images and color images.

[0224] It should be noted that the specific data processing methods of each unit in the above-mentioned adaptive weight transfer device 700 for diffusion model channel expansion can be referred to the method embodiment above, and will not be repeated here.

[0225] Based on the same concept, embodiments of this disclosure also provide a computing device. Figure 8 This is a structural block diagram of a computing device provided in an embodiment of this disclosure. Figure 8 As shown, the computing device 800 may include a processor 801 and a memory 802; the memory 802 may be coupled to the processor 801. It is worth noting that... Figure 8 This is an example; other types of structures can also be used to supplement or replace this structure to achieve telecommunications functions or other functions. Optionally, the computing device 800 can be a server or a local computing device, thereby enabling operation in the cloud or locally.

[0226] In one possible implementation, the functionality of the adaptive weight transfer device 700 for diffusion model channel expansion can be integrated into the processor 801. The processor 801 can be configured to perform the following operations: Obtain the original convolution kernel parameters of the input convolutional layer in the diffusion denoising network of the pre-trained diffusion model. The original convolution kernel parameters have the first number of input channels and are the original input weights. In response to the expansion of the number of latent space channels of the pre-trained diffusion model from the first number of channels to the second number of channels, a new convolutional kernel parameter is created, which has a second number of input channels, which is an integer multiple of the first number of input channels; The original convolution kernel parameters are inherited to the front channels of the new convolution kernel parameters, and the number of channels in the front channels is the same as the number of the first input channels. The original convolution kernel parameters are cloned into the rear channel of the new convolution kernel parameters. The number of channels in the rear channel is the difference between the number of the second input channels and the number of the first input channels. The new convolutional kernel parameters are loaded into the input convolutional layer, replacing the original convolutional kernel parameters, to obtain the expanded diffusion model.

[0227] In another possible implementation, the adaptive weight transfer device 700 for diffusion model channel extension can be configured separately from the processor 801. For example, the adaptive weight transfer device 700 for diffusion model channel extension can be configured as a chip connected to the processor 801, and the adaptive weight transfer method for diffusion model channel extension in the previous embodiment can be implemented through the control of the processor 801.

[0228] Furthermore, in some alternative implementations, the computing device 800 may also include: a communication module, an input unit, an audio processor, a display, a power supply, etc. It is worth noting that the computing device 800 is not necessarily required to include these components. Figure 8 All components shown; in addition, the computing device 800 may also include Figure 8 For components not shown, please refer to existing technologies.

[0229] In some alternative implementations, the processor 801, sometimes also referred to as a controller or operation control, may include a microprocessor or other processor device and / or logic device, which receives input and controls the operation of various components of the computing device 800.

[0230] The memory 802 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store information related to the adaptive weight transfer device 700 for diffusion model channel expansion, and may also store a program for executing that information. The processor 801 may execute the program stored in the memory 802 to perform information storage or processing, etc.

[0231] An input unit can provide input to the processor 801. This input unit may be, for example, a button or touch input device. A power supply can be used to provide power to the computing device 800. A display can be used to display images and text, etc. This display may be, for example, an LCD display, but is not limited to this.

[0232] Memory 802 can be a solid-state memory, such as read-only memory (ROM), random access memory (RAM), SIM card, etc. It can also be a memory that retains information even when power is off, can be selectively erased, and contains more data; examples of this type of memory are sometimes referred to as EPROM, etc. Memory 802 can also be some other type of device. Memory 802 includes buffer memory (sometimes referred to as a buffer). Memory 802 may include an application / function storage section for storing application programs and function programs or processes for executing operations of computing device 800 via processor 801.

[0233] The memory 802 may also include a data storage section for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the computing device. The driver storage section of the memory 802 may include various drivers for the computer device for communication functions and / or for performing other functions of the computer device (such as messaging applications, address book applications, etc.).

[0234] The communication module is a transmitter / receiver that sends and receives signals via an antenna. The communication module (transmitter / receiver) is coupled to the processor 801 to provide input signals and receive output signals, which is the same as in a conventional mobile communication terminal.

[0235] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the various processes of the above-described adaptive weight migration method embodiment for diffusion model channel expansion, and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0236] The aforementioned readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0237] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and system embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0238] While one or more embodiments of this specification provide method operation steps as shown in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only execution order. In actual device or client product execution, the method can be executed in the order shown in the embodiments or drawings or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0239] In the description of this application, it should be noted that the terms "upper", "lower", "inner", "outer", "front", "rear", "left", "right", etc., indicate the orientation or positional relationship based on the orientation or positional relationship in the working state of this application. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0240] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. Those skilled in the art can understand the specific meaning of these terms in this application based on the specific circumstances.

[0241] The present application has been described above with reference to preferred embodiments; however, these embodiments are merely exemplary and illustrative. Various substitutions and modifications can be made to the present application based on these embodiments, all of which fall within the protection scope of the present application.

Claims

1. An adaptive weight migration method for diffusion model channel expansion, characterized in that, The method includes: Obtain the original convolution kernel parameters of the input convolutional layer in the diffusion denoising network of the pre-trained diffusion model. The original convolution kernel parameters have a first number of input channels and are the original input weights. In response to the expansion of the number of latent spatial channels of the pre-trained diffusion model from the first number of channels to the second number of channels, a new convolutional kernel parameter is created, the new convolutional kernel parameter having a second number of input channels, the second number of input channels being an integer multiple of the first number of input channels; The original convolution kernel parameters are inherited to the front channel of the new convolution kernel parameters, and the number of channels in the front channel is the number of the first input channels; The original convolution kernel parameters are cloned into the rear channel of the new convolution kernel parameters, and the number of channels in the rear channel is the difference between the number of the second input channels and the number of the first input channels; The new convolutional kernel parameters are loaded into the input convolutional layer, replacing the original convolutional kernel parameters, to obtain the expanded diffusion model.

2. The method according to claim 1, characterized in that, The method further includes: Obtain the bias parameters corresponding to the original convolution kernel parameters. The bias parameters are used to adjust the bias of each output channel of the input convolutional layer. The bias parameter is used as the bias parameter for the new convolution kernel parameter.

3. The method according to claim 1, characterized in that, The diffusion denoising network also includes an output convolutional layer and a backbone network; Compared to the pre-trained diffusion model, the number of input channels of the input convolutional layer in the diffusion denoising network of the extended diffusion model changes, or the number of input channels of the input convolutional layer and the number of output channels of the output convolutional layer changes, while the structure and parameters of the backbone network in the diffusion denoising network remain unchanged; wherein, the backbone network includes a downsampling module, an intermediate module, an upsampling module, and an attention module.

4. The method according to claim 3, characterized in that, When the number of input channels of the input convolutional layer and the number of output channels of the output convolutional layer change, the method further includes: Obtain the original output weights of the output convolutional layer in the diffusion denoising network, wherein the original output weights have a first number of output channels; In response to the expansion of the number of potential spatial channels from a first number of channels to a second number of channels, a new output weight is created, the new output weight having a second number of output channels, the second number of output channels being an integer multiple of the first number of output channels; The original output weights are inherited to the front channel of the new output weights, and the number of channels in the front channel is the same as the number of the first output channels. The original output weight is cloned to the rear channel of the new output weight, and the number of channels in the rear channel is the difference between the number of the second output channels and the number of the first output channels; The new output weights are loaded into the output convolutional layer, replacing the original output weights, to obtain the expanded diffusion model.

5. The method according to claim 4, characterized in that, After loading the new output weights into the output convolutional layer to replace the original output weights, the method further includes: At least two of the input convolutional layer, the output convolutional layer, and the backbone network are jointly fine-tuned.

6. The method according to claim 1, characterized in that, The integer multiple is N, where N is an integer greater than 1; cloning the original convolution kernel parameters to the rear channel of the new convolution kernel parameters includes: When N equals 2, the original convolution kernel parameters are cloned into the rear channel, and the number of channels in the rear channel is the same as the number of the first input channels; When N is greater than 2, the original convolution kernel parameters are repeatedly cloned to all the rear channels; wherein the number of channels filled each time is the number of the first input channels.

7. The method according to claim 6, characterized in that, The first channel has 4 channels, and the second channel has 8 channels.

8. The method according to claim 1, characterized in that, The step of inheriting the original convolution kernel parameters to the front channel of the new convolution kernel parameters includes: Through a deep copy operation, the original convolution kernel parameters are inherited to the front channel of the new convolution kernel parameters; and / or, The step of cloning the original convolution kernel parameters to the rear channel of the new convolution kernel parameters includes: The original convolution kernel parameters are cloned to the back channel of the new convolution kernel parameters through the deep copy operation. The deep copy operation is an operation that creates a copy of the original convolution kernel with identical parameters but independent storage space.

9. The method according to claim 1, characterized in that, The method further includes: Perform zero-sample distribution consistency verification on the extended diffusion model, wherein the zero-sample distribution consistency verification includes at least one of the following: The extended diffusion model has an output feature map mean square error (MSE) with the pre-trained diffusion model under the same input condition less than 10 -6 ; and / or, The expanded diffusion model has an initial Fraser distance (FID) increment of less than 1.0; and / or, The learned perceptual image patch similarity (LPIPS) of the extended diffusion model is less than 0.

05.

10. The method according to claim 1, characterized in that, The original convolution kernel parameters are used to process single-modal images, and the new convolution kernel parameters are used to process multimodal images, including depth images and color images.

11. An adaptive weight transfer device for diffusion model channel expansion, characterized in that, The device includes: The acquisition module is used to acquire the original convolution kernel parameters of the input convolutional layer in the diffusion denoising network of the pre-trained diffusion model. The original convolution kernel parameters have a first number of input channels and are the original input weights. A module is created to create new convolutional kernel parameters in response to the expansion of the number of latent spatial channels of a pre-trained diffusion model from a first number of channels to a second number of channels. The new convolutional kernel parameters have a second number of input channels, which is an integer multiple of the first number of input channels. An inheritance module is used to inherit the original convolution kernel parameters to the front channel of the new convolution kernel parameters, wherein the number of channels in the front channel is the number of the first input channels; The cloning module is used to clone the original convolution kernel parameters to the rear channel of the new convolution kernel parameters, wherein the number of channels in the rear channel is the difference between the number of the second input channels and the number of the first input channels; The replacement module is used to load the new convolutional kernel parameters into the input convolutional layer, replacing the original convolutional kernel parameters, to obtain the expanded diffusion model.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method of any one of claims 1 to 10.

13. A computing device, characterized in that, include: Memory, used to store computer program products; A processor is configured to execute a computer program product stored in the memory, wherein, when the computer program product is executed, it implements the method described in any one of claims 1 to 10.