Multi-operator isomorphic diffusion model-based acceleration system and method

Through the multi-operator isomorphic diffusion model acceleration system, the problem of resource limitation of diffusion model at edge devices is solved, efficient image generation and resource utilization are achieved, and computing efficiency and image quality are improved.

CN120338005APending Publication Date: 2025-07-18TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510321811.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The application of the diffusion model on edge devices with limited resources is mainly due to the large number of parameters and high computational complexity, which leads to insufficient hardware resources and difficult to meet computing needs.

Method used

A multi-operator isomorphic diffusion model acceleration system is adopted, including memory buffer module, pixel component array control module, accumulation module and quantization module. By caching model parameters and noise image features, batch operations and quantization processing are performed to reduce complex hardware control logic and improve computing resource utilization.

Benefits of technology

It improves computing efficiency, reduces resource waste, improves image generation quality and stability, improves the utilization rate of computing resources, and reduces the demand for hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338005A_ABST
    Figure CN120338005A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides an acceleration system and method based on a multi-operator isomorphic diffusion model, and the system comprises a memory buffer module which is used for caching the model parameters of the trained diffusion model and the inputted noise image features; the pixel element array control module is used for generating input data according to the model parameters and the noise image features, performing batch operation in combination with an algorithm of a diffusion model to obtain an intermediate result of a corresponding operation batch, and inputting the intermediate result of the corresponding operation batch into the accumulation module; the accumulation module is used for obtaining an accumulation result according to all the input intermediate results and inputting the accumulation result into the quantization module; and the quantization module is used for carrying out quantization processing on the accumulation result to obtain an image generation result after quantization processing. Various operators of the diffusion model can be supported, complex control logic of hardware is reduced, and the utilization rate of computing resources is improved while the image generation quality is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to an acceleration system and method for a multi-operator isomorphic diffusion model. Background Art

[0002] In today's digital age, the amount of data has grown explosively, and artificial intelligence technology has also developed rapidly. Among them, the technology of Artificial Intelligence Generated Content (AIGC) has attracted wide attention from all walks of life and has become an important force driving innovation and development. In the field of AIGC technology, the Diffusion Model has become the most advanced model at present due to its excellent data generation and analysis capabilities, and occupies a mainstream position in various generation tasks.

[0003] Currently, the diffusion model has the characteristics of a large number of parameters and high computational complexity. In a single inference process, its internal basic operators need to perform hundreds of operations, and the data scale involved in the calculation is as high as hundreds of millions of parameters, which requires the model to be supported by a large-capacity storage and high-performance computing resources. However, the traditional general-purpose processors pursue the design goals of high energy efficiency and low area, which conflicts with the high demand for computing resources of the diffusion model, thus restricting the application of the diffusion model on resource-constrained edge devices. For example, in some edge devices such as Internet of Things devices and mobile terminals, due to limited hardware resources, it is difficult to meet the operating requirements of the diffusion model, resulting in its inability to fully play its role. Summary of the Invention

[0004] The present invention provides an acceleration system and method for a multi-operator isomorphic diffusion model, which are used to solve the defect that the design goal of the existing general-purpose processor restricts the application of the diffusion model on resource-constrained edge devices, can support various operators of the diffusion model, reduce the complex control logic of the hardware, and improve the utilization rate of computing resources while ensuring the quality of image generation.

[0005] The present invention provides an acceleration system for a multi-operator isomorphic diffusion model, including: a memory buffer module for caching the model parameters of the diffusion model after training and the input noise image features, where the noise image features are extracted based on a previously obtained noise image; a pixel element array control module for generating input data according to the model parameters and the noise image features, and performing batch operations in combination with the algorithm of the diffusion model to obtain intermediate results corresponding to the operation batches, and inputting the intermediate results corresponding to the operation batches into an accumulation module; an accumulation module for obtaining an accumulation result according to all the input intermediate results and inputting the accumulation result into a quantization module; and a quantization module for performing quantization processing on the accumulation result to obtain a quantized image generation result.

[0006] According to an acceleration system based on a multi-operator isomorphic diffusion model provided by the present invention, the model parameters include weight parameters and index parameters, and the index parameters are used to represent the image index positions corresponding to the weights; the pixel element array control module includes: an array controller, which determines the operation batches and the model parameters and noise image features corresponding to each operation batch participating in the operation according to the model parameters and the noise image features, and inputs the model parameters and noise image features corresponding to each operation batch participating in the operation to the corresponding pixel element array; each processing element block in the pixel element array looks up the corresponding index parameter according to the weight parameter in the input model parameters to determine the image index position, obtains the image reading mode, combines the noise image features to obtain the input data, and according to the input data and the input weight parameters, uses the diffusion model algorithm to obtain an intermediate result, and inputs the intermediate result into the accumulation module.

[0007] According to an acceleration system based on a multi-operator isomorphic diffusion model provided by the present invention, the quantization module includes: a quantization unit, which obtains a multiplication result according to the accumulation result and a preset scale factor, and performs a preset shift operation on the multiplication result to obtain a quantization processing result; a comparison unit, which compares each value in the quantization processing result, if there is a value greater than a preset maximum value, then uses the preset maximum value as the image generation result after quantization and outputs it, if there is a value less than a preset minimum value, then uses the preset minimum value as the image generation result after quantization and outputs it.

[0008] According to an acceleration system based on a multi-operator isomorphic diffusion model provided by the present invention, the system further includes a detection module, wherein: the detection module obtains the attention matrix and the value vector of the attention mechanism of the attention layer in the model parameters, groups the attention matrix based on the rows of the attention matrix, and for each group, obtains the corresponding position integration element according to the elements in the group and their corresponding positions, and combines a preset selection strategy to obtain the position integration elements selected by each group and inputs them into the pixel element array control module, and controls to input the value vector into the pixel element array control module; wherein, the attention matrix is generated according to a preset algorithm based on the query vector and the key vector of the attention mechanism of the attention layer; the pixel element array control module performs batch operations according to the input data, the input position integration element, and the value vector of the attention mechanism, and combines the algorithm of the diffusion model.

[0009] A multi-operator isomorphic diffusion model acceleration system provided according to the present invention, the memory buffer module includes an input memory and a vector memory, the input memory is used to store the value vectors of the attention mechanism of the attention layer of the diffusion model, and the vector memory is used to store the attention matrix; the detection module includes: a control unit, which controls the input of the value vectors into the pixel element array control module, and groups the attention matrix based on the rows of the attention matrix, and for each group, integrates each element in the group with the corresponding position of the element respectively to obtain the integrated elements at the corresponding positions, and controls the input of the integrated elements at the corresponding positions into the element selection unit; the element selection unit, for each group, sequentially selects each position integrated element in the group, compares the selected position integrated element with the remaining position integrated elements one by one to obtain the comparison results corresponding to the position integrated elements, and according to the comparison results of each position integrated element, selects a preset number of position integrated elements from the corresponding group, and inputs the selected position integrated elements into the pixel element array control module.

[0010] A multi-operator isomorphic diffusion model acceleration system provided according to the present invention, the system further includes a pattern pruning training module, wherein: the pattern pruning training module performs pattern pruning on the basic operators in the pre-acquired diffusion model parameters to be trained according to a preset sparsity, and uses the pre-acquired training image features to train the diffusion model to be trained after pattern pruning until a preset end condition is reached, and ends the training to obtain the trained diffusion model; wherein, the basic operators include at least one of a one-dimensional convolutional layer, a two-dimensional convolutional layer, and a fully connected layer.

[0011] A multi-operator isomorphic diffusion model acceleration system provided according to the present invention, the pattern pruning training module includes: a pattern pruning unit, which performs pattern pruning on the basic operators in the pre-obtained diffusion model to be trained according to a preset sparsity, and obtains the parameters of the diffusion model to be trained after pattern pruning; a sparse training unit, which controls the input of the pre-obtained training image features into the diffusion model to be trained after pattern pruning, so as to train the diffusion model to be trained after pattern pruning until a preset end condition is reached; or, a pattern pruning unit, which determines a sparsity target according to a preset sparsity, and performs pattern pruning on the basic operators of the pre-obtained diffusion model to be trained according to the sparsity target, and obtains the diffusion model to be trained after pattern pruning; a sparse training unit, which controls the input of the pre-obtained training image features into the diffusion model to be trained after pattern pruning, so as to train the diffusion model to be trained after pattern pruning until a preset end condition is reached; a model evaluation unit, which performs performance evaluation on the diffusion model after preliminary training, obtains a performance evaluation result, and inputs the performance evaluation result into the pattern pruning unit; a pattern pruning unit, which determines whether the model performance meets a preset fluctuation range and determines whether the model sparsity meets a preset sparsity according to the performance evaluation result. If the model performance meets the preset fluctuation range and the model sparsity meets the preset sparsity, the corresponding trained diffusion model is used as the trained diffusion model; otherwise, according to the performance evaluation result, combined with a preset pruning adjustment strategy, the sparsity target is adjusted, and the model is re-pruned according to the adjusted sparsity target; wherein, the preset pruning adjustment strategy is used to limit the reduced or increased sparsity of the corresponding model according to the degree of improvement and decline of the model performance; a sparse training unit, which controls the input of the training image features into the re-pruned diffusion model to be trained, so as to train the re-pruned diffusion model to be trained until a preset end condition is reached; a model evaluation unit, which performs performance evaluation on the re-trained diffusion model, re-obtains a performance evaluation result, and inputs the re-obtained performance evaluation result into the pattern pruning unit.

[0012] A multi-operator isomorphic diffusion model acceleration system provided according to the present invention, the system further includes a quantization-aware training module, wherein: the quantization-aware training module performs quantization-aware training on the diffusion model to be trained according to the pre-obtained training image features until a preset end condition is reached, and ends the training to obtain the trained diffusion model.

[0013] A multi-operator isomorphic diffusion model acceleration system provided according to the present invention, the quantization-aware training module includes: a quantization processing unit, which inserts quantization operators at preset positions of the diffusion model to be trained; a perception training unit, which inputs the pre-obtained training image features into the diffusion model to be trained to train the diffusion model to be trained.

[0014] The present invention also provides a method for accelerating a multi-operator isomorphic diffusion model, which is applied to any one of the above-mentioned multi-operator isomorphic diffusion model acceleration systems. The method includes: obtaining the noise image features and the model parameters of the trained diffusion model, where the noise image features are extracted based on the previously obtained noise image; generating input data according to the model parameters and the noise image features, and performing batch operations in combination with the algorithm of the diffusion model to obtain intermediate results corresponding to the operation batches; obtaining an accumulation result according to all the input intermediate results; performing quantization processing on the accumulation result to obtain a quantized image generation result.

[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements any one of the above-mentioned methods for accelerating a multi-operator isomorphic diffusion model.

[0016] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above-mentioned methods for accelerating a multi-operator isomorphic diffusion model.

[0017] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements any one of the above-mentioned methods for accelerating a multi-operator isomorphic diffusion model.

[0018] The multi-operator isomorphic diffusion model acceleration system and method provided by the present invention use a memory buffer module to cache model parameters and noise image features to avoid reloading these data every time an image is generated, thereby improving computational efficiency. And the pixel element array control module processes data through batch operations to better utilize computing resources and improve the speed of data processing. Further, an accumulation module synthesizes multiple intermediate results to facilitate improving the quality and stability of image generation. Then, a quantization module performs quantization processing on the accumulation result to reduce the error of image generation, improve the clarity and visual effect of the image, thereby effectively utilizing hardware resources, improving the utilization rate of computing resources, and reducing resource waste. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 is one of the structural schematic diagrams of the multi-operator isomorphic diffusion model acceleration system provided by the present invention; Figure 2 It is a schematic structural diagram of the pixel element array control module provided by the present invention; Figure 3 It is the second schematic structural diagram of the multi-operator isomorphic diffusion model acceleration system provided by the present invention; Figure 4 It is a schematic structural diagram of the quantization module provided by the present invention; Figure 5 It is a schematic structural diagram of the detection module provided by the present invention; Figure 6 It is a schematic diagram of pattern pruning provided by the present invention; Figure 7 It is a schematic flow diagram of the multi-operator isomorphic diffusion model acceleration method provided by the present invention; Figure 8 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0021] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0022] Figure 1 It is a schematic structural diagram of the multi-operator isomorphic diffusion model acceleration system provided by the present invention. As Figure 1 shown, the system includes: A memory buffer module, configured to cache the model parameters of the diffusion model after training and the input noise image features, and the noise image features are extracted based on the previously obtained noise images; A pixel element array control module, which generates input data according to the model parameters and the noise image features, performs batch operations in combination with the diffusion model algorithm, obtains the intermediate results of the corresponding operation batch, and inputs the intermediate results of the corresponding operation batch into the accumulation module; An accumulation module, which obtains an accumulation result according to all the input intermediate results and inputs the accumulation result into the quantization module; A quantization module, which performs quantization processing on the accumulation result to obtain a quantized image generation result.

[0023] In this embodiment, the model parameters include weight parameters and index parameters. The index parameters are used to represent the image index positions corresponding to the weights, and are used to indicate the valid and invalid parts of the image. The index parameters can be determined according to the sparse pattern data during the pattern sparse training of the model. For the specific pattern sparse training, refer to the following text and no further elaboration will be made here.

[0024] Correspondingly, referring to Figure 2 , the pixel element array control module includes: an array controller that determines the operation batches and the model parameters and noise image features corresponding to each operation batch participating in the operation according to the model parameters and the noise image features, and inputs the model parameters and noise image features corresponding to each operation batch participating in the operation to the corresponding pixel element array; each processing element block in the pixel element array looks up the corresponding index parameter according to the weight parameter in the input model parameters, determines the image index position, obtains the image reading mode, combines the noise image features to obtain the input data, and according to the input data and the input weight parameter, uses the diffusion model algorithm to obtain an intermediate result, and inputs the intermediate result into the accumulation module.

[0025] It should be added that the diffusion model algorithm can be determined according to the actual algorithm of the diffusion model involved, such as performing multiplication-accumulation operations, etc., and no further limitation will be made here. In addition, the pixel element array control module can be selected according to the actual design requirements, such as selecting a PE Array, etc., and no further limitation will be made here.

[0026] In addition, continuing to refer to Figure 3 , the memory buffer module includes a weight memory, an index memory, and an input memory. The weight memory is used to cache the weight parameters of the diffusion model after training is completed, the index memory is used to cache the index parameters of the diffusion model after training is completed, and the input memory is used to cache the input noise image features. By caching the weight parameters and index parameters into the corresponding memories, various operators of the diffusion model (Diffusion Model) are comprehensively supported, the complex control logic of the hardware is reduced, the operation efficiency is improved, the storage overhead is reduced, the calculation of different operators on the same hardware architecture is realized, and these data are transferred to the pixel element array control module for multiplication-accumulation calculation during each operation.

[0027] Furthermore, continuing to refer to Figure 3 , the system further includes an off-chip memory and a control module. The off-chip memory is used to store the model parameters, and the control module is used to load the model parameters to be calculated into the corresponding memory buffer module, and input the model parameters and the noise image features into the pixel element array control module.

[0028] Furthermore, the system further includes an output memory, and the output memory is used to store the image generation result output by the quantization module.

[0029] Further, the control module controls to store the image generation result in the output memory into the off-chip memory.

[0030] In an alternative embodiment, referring to Figure 4 , the quantization module includes: a quantization unit that obtains a multiplication result according to the accumulation result and a preset scale factor, and performs a preset shift operation on the multiplication result to obtain a quantization processing result; a comparison unit that compares each value in the quantization processing result, and if there is a value greater than a preset maximum value, uses the preset maximum value as the image generation result after quantization and outputs it, and if there is a value less than a preset minimum value, uses the preset minimum value as the image generation result after quantization and outputs it. It should be noted that the preset scale factor can be determined in advance based on the activation value or weight of each layer of the diffusion model, or can be set based on actual experimental results or prior experience. The preset shift operation can be selected according to actual design requirements, such as selecting a shift right operation. The comparison unit can be selected according to actual design requirements, such as selecting a multiplexer, and no further limitation is made here.

[0031] In an alternative embodiment, continuing to refer to Figure 2 , the system further includes a detection module, where: the detection module obtains the attention matrix and the value vector of the attention mechanism of the attention layer in the model parameters, groups the attention matrix based on the rows of the attention matrix, and for each group, obtains the corresponding position integration element according to the elements in the group and their corresponding positions, and combines a preset selection strategy to obtain the position integration elements selected by each group and inputs them into the pixel element array control module, and controls to input the value vector of the attention mechanism into the pixel element array control module; wherein, the attention matrix is generated according to a preset algorithm based on the query vector and the key vector of the attention mechanism of the attention layer; the pixel element array control module performs batch operations according to the input data, the input position integration elements, and the value vector of the attention mechanism, and in combination with the diffusion model algorithm.

[0032] It should be added that through the above setting of the detection module, the utilization rate of computing resources is improved, thereby enhancing the overall performance. In addition, the diffusion model includes a feature extraction layer, an attention layer, a feature fusion layer, and an image generation layer, where: the feature extraction layer is used to extract the features of the input noise image features to obtain image features; the attention layer is used to identify the image features and generate an attention map; the feature fusion layer is used to fuse the image features and the attention map to obtain fusion features; the image generation layer is used to perform backpropagation according to the fusion features to obtain a denoised image.

[0033] Further, the attention matrix is obtained by multiplying the query vector Q and the transposed matrix of the key vector K and passing through the softmax function.

[0034] Furthermore, the feature extraction layer can adopt, such as, a one-dimensional convolutional layer (Conv1d), a two-dimensional convolutional layer (Conv2d), etc. The output layer in the image generation layer can adopt a fully connected layer, which is used to predict the noise residual and the denoising parameter. Specifically, it can be selected according to actual design requirements and will not be further limited here. It is worth noting that through the above detection module, the computational and storage overhead of the data volume can be greatly reduced.

[0035] In addition, referring to Figure 5 , the memory buffer module includes an input memory and a vector memory. The input memory is used to store the value vector of the attention mechanism of the diffusion model attention layer, and the vector memory is used to store the attention matrix; the detection module includes: a control unit, which controls to input the value vector into the pixel element array control module, and groups the attention matrix based on the rows of the attention matrix, and for each group, integrates each element in the group with the corresponding position of the element respectively to obtain the position integration element of the corresponding group, and controls to input the position integration elements corresponding to each group into the element selection unit; the element selection unit, for each group, selects each position integration element in the group one by one, compares the selected position integration element with the remaining position integration elements one by one to obtain the comparison result corresponding to the position integration element, and according to the comparison results of each position integration element, selects a preset number of position integration elements from the corresponding group, and inputs the selected position integration elements into the pixel element array control module.

[0036] It should be added that the preset number can be set according to the actual number of columns of the involved attention matrix and the involved requirements or prior experience. For example, if the attention matrix has 16 rows, the preset number can be 4, and no further electrical details are provided here.

[0037] In addition, the control unit includes: an encoding subunit, which encodes the corresponding element position for each element in the group to obtain the position encoding of the corresponding element; an integration subunit, which splices the element with the position encoding of the corresponding element to obtain the position integration element of the corresponding element.

[0038] For example, assume that the elements in the attention matrix are 10-bit scores, and each group contains 16 elements, and the preset number is 4. Then for each element, the corresponding position information is encoded into 4-bit position information through position encoding, and the 10-bit score and the 4-bit position information are combined into a new 16-bit value to obtain the position integration element, denoted as Comb, and each Comb is numbered to obtain S0 - S15. The position integration element contains both the value of the element itself and the position information of the element.

[0039] Next, the Comb value will be compared one by one with the other 15 Comb values in the same group; if the current Comb value is greater than the other Comb values, the corresponding Sort flag is incremented by 1; if the current Comb value is less than or equal to the other Comb values, the corresponding Sort flag is incremented by 0. The 16 comparison results obtained through the comparison, namely the Sort values, are used to accurately determine the sorting order of the 16 input values in terms of size. Based on this sorting order, the top 4 Comb values are selected in descending order. Each Comb value not only contains the numerical information of the corresponding element but also carries the position information of the element.

[0040] Finally, the 4 Combs selected from each group and the value vector (Value) are input into the pixel element array control module together, so that the pixel element array control module can perform the multiply-accumulate operation of the matrix based on this information, and finally complete the calculation of the attention layer. In this way, the most relevant values can be effectively screened out and participate in the subsequent model calculation, thereby improving the calculation efficiency and reducing the processing of irrelevant data, and optimizing the calculation performance of the entire neural network.

[0041] In the actual experimental process, for the attention mechanism layer, sparse characteristics are found in the attention matrix (Attention score). Each row of data is divided into a group of 16, and the largest 4 data are taken from each group, and the remaining data become 0. As a result, 75% of the computational amount can be reduced in the attention mechanism layer, and the FID value increases by 0.16.

[0042] In an alternative embodiment, before performing batch operations using the pixel element array control module, the diffusion model needs to be pre-trained. Accordingly, the system further includes a pattern pruning training module, where: the pattern pruning training module performs pattern pruning on the basic operators in the pre-obtained diffusion model to be trained according to a preset sparsity, and uses the pre-obtained training image features to train the diffusion model to be trained after pattern pruning until a preset end condition is reached, and the training is ended to obtain the trained diffusion model; where the basic operators include at least one of a one-dimensional convolutional layer, a two-dimensional convolutional layer, and a fully connected layer.

[0043] It should be noted that by introducing the pattern sparse training method, the sparse characteristics of the data in the diffusion model (DiffusionModel) algorithm are effectively utilized. Without significant difference in the final image generation quality, the neural network of the diffusion model is greatly compressed to significantly reduce the number of model parameters while maintaining good generation effects of the model. In addition, the pattern pruning for the one-dimensional convolutional layer, two-dimensional convolutional layer, and fully connected layer can refer to Figure 6As shown, it calculates the importance of the corresponding pattern for the corresponding layer, sorts according to the importance, and performs structured pruning according to the preset sparsity to remove the least important patterns and fine-tune the model. Among them, in a one-dimensional convolutional layer, a pattern refers to an element of the convolutional kernel; in a two-dimensional convolutional layer, a pattern refers to a row or a column in the convolutional kernel; in a fully connected layer, a pattern refers to the weights of an input neuron connecting to all output neurons or the weights received by an output neuron from all input neurons.

[0044] In addition, the preset sparsity can be set according to prior experience or actual experimental results. It can be known from experiments that the sparsities of different layers in the diffusion model have different effects on the metric Frechet Inception Distance (FID for short) for evaluating the quality of the images generated by the model. For example, in key layers for parameter transmission, such as the skip_connection layer, the impact is greater. Therefore, the sparsity of each layer can be configured according to FID. During the actual experimental process, the overall sparsity of the entire model is as high as 89% (reducing the number of parameters by 89%), and the value of FID increases by 0.26 (the smaller the FID value, the higher the quality of the generated images).

[0045] In a possible implementation, for scenarios where the network structure does not need to be frequently adjusted, one-time pruning can be selected. It performs one-time pruning on the parameters of the pre-specified layers in the model before the start of training, and then uses the training data to train the model until the preset sparsity is reached.

[0046] Specifically, the pattern pruning training module includes: a pattern pruning unit that performs pattern pruning on the basic operators in the pre-obtained diffusion model to be trained according to the preset sparsity to obtain the parameters of the diffusion model to be trained after pattern pruning; a sparse training unit that controls the input of the pre-obtained training image features into the diffusion model to be trained after pattern pruning to train the diffusion model to be trained after pattern pruning until the preset end condition is reached.

[0047] In another possible implementation, for scenarios where the model needs to be finely adjusted to obtain the best performance, iterative pruning can be selected. It re-adjusts the pruning strategy according to the weight distribution and performance of the current network model before each iterative training to gradually explore and optimize the sparsity of the model, helps the model gradually adapt to the sparse structure, finds the optimal sparse structure, while maintaining or improving the performance of the model and improving the pruning effect.

[0048] Specifically, the pattern pruning training module includes: a pattern pruning unit that determines a sparsity target according to a preset sparsity, and performs pattern pruning on the basic operators of the pre-obtained diffusion model to be trained according to the sparsity target, so as to obtain a diffusion model to be trained after pattern pruning; a sparse training unit that controls inputting the pre-obtained training image features into the diffusion model to be trained after pattern pruning to train the diffusion model to be trained after pattern pruning until a preset end condition is reached; a model evaluation unit that performs performance evaluation on the diffusion model after preliminary training to obtain a performance evaluation result, and inputs the performance evaluation result into the pattern pruning unit; the pattern pruning unit determines whether the model performance meets a preset fluctuation range and determines whether the model sparsity meets a preset sparsity according to the performance evaluation result. If the model performance meets the preset fluctuation range and the model sparsity meets the preset sparsity, the corresponding trained diffusion model is used as the diffusion model after training is completed; otherwise, according to the performance evaluation result, combined with a preset pruning adjustment strategy, the sparsity target is adjusted, and the model is re-pruned according to the adjusted sparsity target; wherein, the preset pruning adjustment strategy is used to limit the sparsity reduced or increased by the corresponding model according to the degree of improvement and decline of the model performance; the sparse training unit controls inputting the training image features into the diffusion model to be trained after re-pruning to train the diffusion model to be trained after re-pruning until a preset end condition is reached; the model evaluation unit performs performance evaluation on the diffusion model after re-training to re-obtain a performance evaluation result, and inputs the re-obtained performance evaluation result into the pattern pruning unit.

[0049] In an alternative embodiment, the pattern pruning training module further includes: a first dynamic detection unit that, when training the diffusion model to be trained after pattern pruning, detects the attention layer based on a preset attention detection mechanism; wherein the preset attention mechanism is to detect the attention matrix and the value vector of the attention mechanism of the attention layer; group the attention matrix based on the rows of the attention matrix; for each group, obtain the corresponding position integration element according to the elements in the group and their corresponding positions, and combine a preset selection strategy to obtain the position integration elements of each group; wherein the attention matrix is generated according to a preset algorithm based on the query vector and the key vector of the attention mechanism of the attention layer; select a preset number of position integration elements from the group, and combine the value vector of the attention mechanism to update the attention mechanism of the model. Specifically, reference can be made to the principle of the detection module in the previous text to facilitate the effective utilization of the detection module during the inference application process, ensuring that the way the model processes data during training and inference application processes remains consistent, and no further elaboration will be made here. In an alternative embodiment, before batch operations are performed using the pixel element array control module, the diffusion model needs to be pre-trained. Correspondingly, the system further includes a quantization-aware training module, wherein: the quantization-aware training module performs quantization-aware training on the diffusion model to be trained according to the previously obtained training image features until a preset end condition is reached, at which point the training ends and the trained diffusion model is obtained.

[0050] It should be noted that the data of the diffusion model is in the floating-point formats FLOAT32 and FLOAT16. In the process of multiply-accumulate calculations, the hardware design is more complex, increasing power consumption, area, and other overheads. Through quantization-aware training, the number of model parameters can be greatly reduced while maintaining good generation effects of the model. Additionally, during the actual experiment process, quantization-dequantization operators are added during training to achieve the quantization of parameters from floating-point numbers FLOAT32 and FLOAT16 to integers INT10, improving the computing resource utilization rate of the accelerator, thereby enhancing the overall performance of the accelerator, and the value of FID increases by 0.09.

[0051] Specifically, the quantization-aware training module includes: a quantization processing unit that inserts quantization operators at preset positions of the diffusion model to be trained; a perception training unit that inputs the previously obtained training image features into the diffusion model to be trained to train the diffusion model to be trained.

[0052] It should be noted that the quantization operator is a key component in neural network quantization, which refers to the process of converting data from floating-point (FLOAT) precision to another lower integer precision (such as INT) in a neural network model. This conversion can make the model run more efficiently on hardware, reduce the consumption of computing resources and accelerate the inference process, while also reducing the storage requirements of the model. Additionally, during the quantization-aware training process, quantization operators need to be inserted at specified positions (such as before and after weights and activations) in the diffusion model to be trained to simulate the quantization process at the software level. The specified positions can be set according to weights and activations. For example, for weights, the specified positions can be before each convolutional layer, fully connected layer, or any layer with learnable weights, and after weight updates (gradient application). For activations, the specified positions can be after each activation function, i.e., on the output of the neuron, or between the layer output and layer input to simulate the impact of quantization on the data flow between layers. Specifically, it can be set according to actual design requirements and will not be further limited here.

[0053] In addition, during the process of training the diffusion model to be trained using training image features, during forward propagation, the floating-point weights and activations are quantized based on the previously inserted quantization operators, and during backward propagation, the quantized weights and activations are dequantized to generate a denoised image and determine the gradient of the model parameters, so as to update the weights of the diffusion model to be trained using the gradient of the model parameters and iterate until the loss function constructed based on the denoised image and the label corresponding to the training image features converges, ending the training and obtaining the trained diffusion model.

[0054] In an alternative embodiment, the quantization-aware training module further includes: a second dynamic detection unit that, when performing quantization-aware training on the diffusion model to be trained, detects the attention layer based on a preset attention detection mechanism; wherein the preset attention mechanism is to detect the attention matrix and the value vector of the attention mechanism of the attention layer; group the attention matrix based on the rows of the attention matrix; for each group, obtain the corresponding position integration element according to the elements in the group and their corresponding positions, and combine a preset selection strategy to obtain the position integration elements of each group; wherein the attention matrix is generated according to a preset algorithm based on the query vector and key vector of the attention mechanism of the attention layer; select a preset number of position integration elements from the groups and combine them with the value vector of the attention mechanism to update the attention mechanism of the model. Specifically, reference can be made to the first dynamic detection unit in the previous text and will not be further elaborated here.

[0055] In an alternative embodiment, the output memory is further used to store the model parameters of the trained diffusion model, and the control module controls to store the model parameters of the trained diffusion model in the output memory to an off-chip memory.

[0056] During the actual experiment process, tests were conducted on the DDPM-IP network model (Mang Ning, Enver Sangineto, Angelo Porrello, Simone Calderara, and Rita Cucchiara. 2023. Input perturbation reduces exposure bias in diffusion models. In Proceedings of the 40th International Conference on Machine Learning (ICML'23), Vol. 202. JMLR.org, Article 1093, 26245–26265.) on the Cifar-10 dataset (Krizhevsky A, Hinton G. Learning multiple layers of features from tiny images[J]. 2009.). The experiments show that when the model sparsity reaches 89%, the data is quantized to INT10, and the Top-4 of the attention mechanism layer is sparse, the metric for the quality of the generated images increases by 0.51. At the same time, on the designed processor, the overall peak energy efficiency can reach 12.88 TOPS / W and the overall peak area efficiency can reach 0.389 TOPS / mm2, which are 1.1 - 6.5 times and 3.7 - 26.5 times higher than the previous related work respectively.

[0057] In summary, in the embodiment of the present invention, the memory buffer module caches the model parameters and the noise image features to avoid reloading these data every time an image is generated, thereby improving the calculation efficiency. And the pixel element array control module processes the data through batch operations to better utilize the computing resources and improve the speed of data processing. Further, the accumulation module synthesizes multiple intermediate results to facilitate improving the quality and stability of image generation. Furthermore, the quantization module quantizes the accumulated results to reduce the error of image generation and improve the clarity and visual effect of the images, thereby effectively utilizing the hardware resources, improving the utilization rate of computing resources, and reducing resource waste.

[0058] The following describes the multi-operator isomorphic diffusion model acceleration method provided by the present invention. The multi-operator isomorphic diffusion model acceleration method described below can be correspondingly referred to the multi-operator isomorphic diffusion model acceleration system described above.

[0059] Figure 7 A flowchart of a multi-operator isomorphic diffusion model acceleration method is shown, which is applied to any one of the above-mentioned multi-operator isomorphic diffusion model acceleration systems. The method includes: S71, Obtain the noise image features and the model parameters of the diffusion model after training is completed. The noise image features are extracted based on the previously obtained noise image; S72, Generate input data according to the model parameters and the noise image features, and perform batch operations in combination with the diffusion model algorithm to obtain the intermediate results corresponding to the operation batches; S73, Obtain the accumulated result according to all the input intermediate results; S74, Perform quantization processing on the accumulated result to obtain the image generation result after quantization processing.

[0060] It should be noted that the specific principle of the embodiment of the present invention is the same as that of the above system embodiment. For details, please refer to the previous system embodiment, and no more detailed explanations will be given here.

[0061] Figure 8 An example of the physical structure diagram of an electronic device is shown as Figure 8 shown. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete communication with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the acceleration method based on the multi-operator isomorphic diffusion model. The method includes: obtaining the noise image features and the model parameters of the diffusion model after training is completed. The noise image features are extracted based on the previously obtained noise image; generating input data according to the model parameters and the noise image features, and performing batch operations in combination with the diffusion model algorithm to obtain the intermediate results corresponding to the operation batches; obtaining the accumulated result according to all the input intermediate results; performing quantization processing on the accumulated result to obtain the image generation result after quantization processing.

[0062] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0063] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the steps of obtaining the noise image features and the model parameters of the diffusion model after training provided by the above-mentioned various methods. The noise image features are extracted based on the previously obtained noise images; generate input data according to the model parameters and the noise image features, and perform batch operations in combination with the diffusion model algorithm to obtain intermediate results corresponding to the operation batches; obtain an accumulation result according to all the input intermediate results; perform quantization processing on the accumulation result to obtain a quantized image generation result.

[0064] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the steps of obtaining the noise image features and the model parameters of the diffusion model after training provided by the above-mentioned various methods. The noise image features are extracted based on the previously obtained noise images; generate input data according to the model parameters and the noise image features, and perform batch operations in combination with the diffusion model algorithm to obtain intermediate results corresponding to the operation batches; obtain an accumulation result according to all the input intermediate results; perform quantization processing on the accumulation result to obtain a quantized image generation result.

[0065] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0066] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. An acceleration system based on a multi-operator isomorphic diffusion model, characterized in that, Including: A memory buffer module for caching the model parameters and noise image features of the diffusion model after training is completed, where the noise image features are extracted based on a previously obtained noise image; A pixel element array control module that generates input data according to the model parameters and the noise image features, performs batch operations in combination with the algorithm of the diffusion model to obtain intermediate results corresponding to the operation batches, and inputs the intermediate results corresponding to the operation batches into the accumulation module; The accumulation module obtains an accumulation result based on all the input intermediate results and inputs the accumulation result into the quantization module; The quantization module performs quantization processing on the accumulation result to obtain a quantized image generation result.

2. The acceleration system based on the multi-operator isomorphic diffusion model according to claim 1, wherein The model parameters include weight parameters and index parameters, and the index parameters are used to represent the image index positions corresponding to the weights; The pixel element array control module includes: An array controller that determines operation batches and the model parameters and noise image features corresponding to each operation batch participating in the operation according to the model parameters and the noise image features, and inputs the model parameters and noise image features corresponding to each operation batch participating in the operation into the corresponding pixel element array; Each processing element block in the pixel element array looks up the corresponding index parameter according to the weight parameter in the input model parameters to determine the image index position, obtains an image reading mode, combines the noise image features to obtain input data, and uses the diffusion model algorithm according to the input data and the input weight parameters to obtain an intermediate result, and inputs the intermediate result into the accumulation module; 3. The acceleration system based on the multi-operator isomorphic diffusion model according to claim 1, wherein The quantization module includes: A quantization unit that obtains a multiplication result according to the accumulation result and a preset scale factor, and performs a preset shift operation on the multiplication result to obtain a quantization processing result; A comparison unit compares each value in the quantization processing result. If there is a value greater than a preset maximum value, the preset maximum value is used as the quantized image generation result and output. If there is a value less than a preset minimum value, the preset minimum value is used as the quantized image generation result and output.

4. The acceleration system based on the multi-operator isomorphic diffusion model according to claim 1, characterized in that, The system further includes a detection module, where: The detection module obtains the attention matrix and the value vector of the attention mechanism in the model parameters, groups the attention matrix based on the rows of the attention matrix, and for each group, obtains the corresponding position integration element according to the elements in the group and their corresponding positions, combines a preset selection strategy to obtain the position integration elements selected by each group and inputs them into the pixel element array control module, and controls the input of the value vector into the pixel element array control module; where the attention matrix is generated according to a preset algorithm based on the query vector and the key vector of the attention mechanism of the attention layer; The pixel element array control module performs batch operations in combination with the algorithm of the diffusion model according to the input data, the input position integration element, and the value vector of the attention mechanism.

5. The acceleration system based on the multi-operator isomorphic diffusion model according to claim 4, wherein The memory buffer module includes an input memory and a vector memory. The input memory is used to store the value vectors of the attention mechanism of the diffusion model attention layer, and the vector memory is used to store the attention matrix; The detection module includes: A control unit that controls the input of value vectors into the pixel element array control module, groups the attention matrix based on the rows of the attention matrix, and for each group, integrates each element in the group with the corresponding position of the element to obtain the integrated elements at the corresponding positions, and controls the input of the integrated elements at the corresponding positions into the element selection unit; The element selection unit, for each group, sequentially selects each position integrated element within the group, compares the selected position integrated element with the remaining position integrated elements one by one to obtain the comparison result corresponding to the position integrated element, and selects a preset number of position integrated elements from the corresponding group according to the comparison results of the position integrated elements, and inputs the selected position integrated elements into the pixel element array control module.

6. The accelerated system based on the multi-operator isomorphic diffusion model according to claim 1, wherein The system further includes a mode pruning training module, where: The mode pruning training module performs mode pruning on the basic operators in the previously obtained diffusion model to be trained according to a preset sparsity, and uses the previously obtained training image features to train the diffusion model to be trained after mode pruning until a preset end condition is reached, and then ends the training to obtain the trained diffusion model; where the basic operators include at least one of a one-dimensional convolutional layer, a two-dimensional convolutional layer, and a fully connected layer.

7. The accelerated system based on the multi-operator isomorphic diffusion model according to claim 6, characterized in that, The mode pruning training module includes: A mode pruning unit that performs mode pruning on the basic operators in the previously obtained diffusion model to be trained according to a preset sparsity to obtain the parameters of the diffusion model to be trained after mode pruning; A sparse training unit that controls the input of the previously obtained training image features into the diffusion model to be trained after mode pruning to train the diffusion model to be trained after mode pruning until a preset end condition is reached; or, A mode pruning unit that determines a sparsity target according to the preset sparsity, and performs mode pruning on the basic operators of the previously obtained diffusion model to be trained according to the sparsity target to obtain the diffusion model to be trained after mode pruning; A sparse training unit that controls the input of the previously obtained training image features into the diffusion model to be trained after mode pruning to train the diffusion model to be trained after mode pruning until a preset end condition is reached; A model evaluation unit that performs performance evaluation on the diffusion model after preliminary training to obtain a performance evaluation result, and inputs the performance evaluation result into the mode pruning unit; The pattern pruning unit determines whether the model performance meets the preset fluctuation range and whether the model sparsity meets the preset sparsity according to the performance evaluation result. If the model performance meets the preset fluctuation range and the model sparsity meets the preset sparsity, the corresponding trained diffusion model is used as the trained diffusion model after training; otherwise, according to the performance evaluation result, combined with the preset pruning adjustment strategy, the sparsity target is adjusted, and the model is re-pruned according to the adjusted sparsity target; wherein, the preset pruning adjustment strategy is used to limit the reduced or increased sparsity of the corresponding model according to the degree of improvement and decline of the model performance. The sparse training unit controls to input the training image features into the to-be-trained diffusion model after re-pruning to train the to-be-trained diffusion model after re-pruning until the preset end condition is reached. The model evaluation unit performs a performance evaluation on the re-trained diffusion model, obtains a performance evaluation result again, and inputs the obtained performance evaluation result again into the pattern pruning unit. The pattern pruning training module further includes: The first dynamic detection unit, when training the to-be-trained diffusion model after pattern pruning, detects the attention layer based on a preset attention detection mechanism; wherein, the preset attention mechanism is to detect the attention matrix and the value vector of the attention mechanism of the attention layer; group the attention matrix based on the rows of the attention matrix; for each group, obtain the corresponding position integration element according to the elements in the group and their corresponding positions, and combine the preset selection strategy to obtain the position integration elements of each group; wherein, the attention matrix is generated according to a preset algorithm based on the query vector and the key vector of the attention mechanism of the attention layer; select a preset number of position integration elements from the group and update the attention mechanism of the model in combination with the value vector of the attention mechanism.

8. The acceleration system based on the multi-operator isomorphic diffusion model according to claim 1, wherein The system further includes a quantization-aware training module, wherein: The quantization-aware training module performs quantization-aware training on the to-be-trained diffusion model according to the previously obtained training image features until the preset end condition is reached, ends the training, and obtains the trained diffusion model after training.

9. The acceleration system based on the multi-operator isomorphic diffusion model according to claim 8, characterized in that, The quantization-aware training module includes: The quantization processing unit inserts quantization operators at preset positions of the to-be-trained diffusion model. The perception training unit inputs the previously obtained training image features into the to-be-trained diffusion model to train the to-be-trained diffusion model. The quantization-aware training module further includes: The second dynamic detection unit, when performing quantization-aware training on the diffusion model to be trained, detects the attention layer based on a preset attention detection mechanism; wherein, the preset attention mechanism is to detect the attention matrix and the value vector of the attention mechanism of the attention layer; group the attention matrix based on the rows of the attention matrix; for each group, obtain the corresponding position integration element according to the elements in the group and their corresponding positions, and combine a preset selection strategy to obtain the position integration elements of each group; wherein, the attention matrix is generated by a preset algorithm based on the query vector and the key vector of the attention mechanism of the attention layer; select a preset number of position integration elements from the groups, and combine the value vector of the attention mechanism to update the attention mechanism of the model.

10. A multi-operator isomorphic diffusion model acceleration method, which is applied to the multi-operator isomorphic diffusion model acceleration system according to any one of claims 1-9, and is characterized in that, The method includes: Obtaining the noise image features and the model parameters of the diffusion model after training is completed, where the noise image features are extracted based on the previously obtained noise image; Generating input data according to the model parameters and the noise image features, and performing batch operations in combination with the algorithm of the diffusion model to obtain the intermediate results of the corresponding operation batches; Obtaining an accumulation result according to all the input intermediate results; Performing quantization processing on the accumulation result to obtain the image generation result after quantization processing.