Accelerator of video generation model and computing system

By designing sparse and matrix processing units in the accelerator of the video generation model, the problem of redundancy overhead in VGM calculation is solved, and efficient video generation and excellent performance and energy efficiency ratio are achieved.

CN120075550APending Publication Date: 2025-05-30SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510224062.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing video generation models (VGMs) have redundant computing overhead in computing, making it difficult to fully utilize the performance of GPU and FPGA, resulting in low throughput and high energy consumption.

Method used

An accelerator for video generation model is designed, including a sparse unit, a matrix processing engine and a recovery unit. Through inter-frame and intra-frame sparse, similar marks are recorded, and calculations are performed in the matrix processing engine to generate output activation, realizing online sparse and mixed precision calculations.

Benefits of technology

It effectively reduces the computational overhead of video generation model, improves performance and energy efficiency ratio, and realizes efficient VGM inference without reducing video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075550A_ABST
    Figure CN120075550A_ABST
Patent Text Reader

Abstract

The present invention discloses an accelerator and a computing system for a video generation model, the accelerator comprising: a sparsification unit, the input activation of which comprises a plurality of consecutive frames, each frame being segmented into a plurality of marks; the rarefaction unit is used for performing inter-frame rarefaction and intra-frame rarefaction on input activation containing a plurality of frames so as to screen and record similar marks; the matrix processing engine is used for calculating the marks to obtain output marks; wherein if the two marks are recorded to be similar, an output mark of one mark is calculated, and the output mark serves as the output mark of the two marks; and a recovery unit which generates an output activation according to the calculation result of the matrix processing engine and the record of the similar marks. According to the accelerator, the problem of redundant calculation overhead in an existing video generation model is solved by utilizing the similarity of video frames in time and space dimensions and designing a customized sparse and recovery unit, hardware resources can be fully utilized, and the accelerator has excellent performance and energy efficiency ratio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to an accelerator and a computing system for a video generation model. Background Art

[0002] With the emergence of large language models (LLMs), video generation models (VGMs) are regarded as the second milestone on the road to general artificial intelligence. As a representative of multi-modal large models, video generation models have revolutionized the production mode of video content creation. They can provide users with creative videos that meet their requirements under appropriate prompts. Video generation models (VGMs) can provide users with creative videos that meet their needs under appropriate prompts. The mainstream video generation models are developed based on the diffusion Transformer (i.e., Diffusion Transformer, DiT) structure, and the DiT structure follows the law of expansion and shows stronger capabilities as the training data increases.

[0003] Due to the inherent similarity of the videos generated in the time and space dimensions, VGMs exhibit significant redundancy in computing. Sparsification methods are an important way to reduce latency and energy consumption and are widely used to accelerate computationally intensive deep learning models. However, it is difficult for the sparsified VGMs to fully utilize the performance of existing general-purpose graphics processing units (GPUs). This is because the sparsity in activation introduces irregular memory access and synchronization processing, resulting in non-negligible overhead on the GPU. There is a huge gap of up to 10 times between the actual and theoretical acceleration of the NVIDIA 3090 GPU under tests of different video sizes. Field-programmable gate arrays (FPGAs) are one of the good candidates for accelerating sparsified deep learning models. However, due to the significant difference in peak computing performance, there is still a large gap between the throughput rates of FPGAs and GPUs. Taking the latest AMD V80 FPGA as an example, it only provides a peak computing performance of about 6.5 TOPS FP16, while the NVIDIA 3090 GPU provides a peak computing performance of 142 TOPS FP16.

[0004] Sparsification is a common method for accelerating computationally intensive models, but it is difficult for sparse VGMs to fully utilize the effective throughput rate of GPUs. FPGAs are good candidates for accelerating sparsified deep learning models, but due to the significant gap in peak computing performance with GPUs (>21 times), existing FPGA accelerators still face the problem of low throughput rate (<2 TOPS) on VGMs. To achieve a higher throughput rate than GPUs, the sparse VGM acceleration based on FPGAs still faces the following challenges:

[0005] 1) There is a large amount of underutilized redundant computation in VGM. The activated similarities exist simultaneously in the temporal and spatial dimensions. However, the current state-of-the-art DiT accelerators only utilize a part of the similarities (i.e., temporal or spatial) in VGM, resulting in a sparsity lower than 46.44%. When running at such a low sparsity, the throughput of sparse VGM on the GPU is even worse than that of dense VGM.

[0006] 2) The digital signal processors (DSPs) on FPGAs cannot provide full computational performance in mixed precision. Different layers of the video generation model (VGM) have different sensitivities to computational precision. For example, VGM uses INT8 precision in linear computations and FP16 precision in attention computations. However, the DSP58 on the AMD V80 FPGA can only be configured offline in FP16 or INT8 mode, resulting in a serious underutilization of computing power in the fixed-point number - floating-point number separation processing engine design.

[0007] 3) Existing scheduling methods face the problem of low computational utilization in online sparsification. Existing scheduling methods include static scheduling and dynamic scheduling. Static scheduling methods rely on offline latency prediction, which is not available in online sparsification. Dynamic scheduling methods result in unacceptable additional overhead due to the limited performance of the embedded CPUs on FPGAs and the complex sparse DiT structure. Since the redundancy of VGM exists in continuously changing videos, existing scheduling methods cannot effectively handle the sparsely activated regions that change dynamically during runtime, leading to low computational utilization of sparse VGM. Summary of the Invention

[0008] In view of the above defects of the prior art, the present invention provides an accelerator and a computing system for a video generation model; it supports online sparsification and mixed precision, and can efficiently perform VGM inference without degrading video quality. To achieve the above technical objectives, the present invention provides:

[0009] An accelerator for a video generation model, comprising:

[0010] A sparsification unit, whose input activation includes a continuous plurality of frames, and each frame is segmented into a plurality of tokens; the sparsification unit performs inter-frame sparsification and intra-frame sparsification on the input activation containing a plurality of frames to screen and record similar tokens;

[0011] A matrix processing engine for computing the tokens to obtain output tokens; wherein, if two tokens are recorded as similar, the output token of one of the tokens is computed and used as the output token of these two tokens;

[0012] A recovery unit for generating output activation according to the calculation result of the matrix processing engine and the record of similar tokens.

[0013] A further improvement of the present invention lies in that, during the process of inter-frame sparsification, one frame is selected from a plurality of consecutive input-activated frames as a reference frame, and the similarity between each marker in the reference frame and the corresponding markers in other frames is calculated respectively; when the similarity between two markers is greater than the inter-frame similarity threshold, the two markers are recorded as similar.

[0014] A further improvement of the present invention lies in that, during the process of intra-frame sparsification, one marker is selected from the multiple markers of each frame as a reference marker, and the similarity between each of the other markers in the frame and the reference marker is calculated respectively. When the similarity between two markers is greater than the intra-frame similarity threshold, the two markers are recorded as similar.

[0015] A further improvement of the present invention lies in that if a marker in a frame is recorded as similar to the corresponding marker in the reference frame during the process of inter-frame sparsification, then during the process of intra-frame sparsification, the similarity between this marker and the reference marker is set to 0.

[0016] A further improvement of the present invention is as follows:

[0017] During the process of intra-frame sparsification, an intra-frame index table is used to record whether each marker in each frame is similar to the reference marker of its frame.

[0018] During the process of inter-frame sparsification, an inter-frame index table is used to record whether each marker in each frame is similar to the corresponding marker in the reference frame.

[0019] A further improvement of the present invention lies in that during the process of the restoration unit generating the output activation, the output marker of each marker is determined one by one.

[0020] During the process of determining the output marker of each marker, it is judged whether the marker is similar to the reference marker of its frame according to the intra-frame index table; if they are similar, the output marker of the reference marker is used as the output marker of this marker; and it is judged whether the marker is similar to the corresponding marker in the reference frame according to the inter-frame index table; if they are similar, the output marker of the corresponding marker in the reference frame is used as the output marker of this marker.

[0021] A further improvement of the present invention lies in that the cosine similarity is used to measure the similarity between two markers.

[0022] A further improvement of the present invention lies in that the matrix processing engine obtains the output marker of each marker through linear calculation; the matrix processing engine is arranged in a programmable gate array chip.

[0023] A further improvement of the present invention lies in that the matrix processing engine includes an extended digital signal processing unit.

[0024] The extended digital signal processing unit includes two 8-bit integer multipliers, two integer adders, and a digital signal processing unit; the extended digital signal processing unit responds to a mode control signal and switches between an integer operation state and a floating-point operation state;

[0025] In the integer operation state, the extended digital signal processing unit performs four 8-bit integer multiply-accumulate operations at a time;

[0026] In the floating-point operation state, the extended digital signal processing unit performs two 16-bit floating-point multiply-accumulate operations at a time.

[0027] The present invention also provides a computing system, which includes the accelerator of the above video generation model.

[0028] The technical solution provided by the present invention has the following technical effects: By utilizing the similarity of video frames in the temporal and spatial dimensions and designing customized sparse and recovery units, the problem of redundant computational overhead existing in the existing video generation models is solved, the hardware resources can be fully utilized, and excellent performance and energy efficiency ratio are achieved.

[0029] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the drawings to fully understand the purpose, features and effects of the present invention. Description of the Drawings

[0030] Figure 1 Shows the overall hardware architecture of the accelerator of the video generation model;

[0031] Figure 2 Shows a schematic diagram of spatio-temporal activation sparsification;

[0032] Figure 3 Is a schematic diagram of the extended digital signal processing unit;

[0033] Figure 4 Is a schematic diagram of dynamic-static adaptive scheduling;

[0034] Figure 5 Is a comparison chart of the model accuracy evaluation results;

[0035] Figure 6 Is a comparison chart of the acceleration and energy efficiency evaluation results. Detailed Embodiments

[0036] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0037] As Figure 1 shown, an embodiment of the present invention provides an accelerator for a video generation model, which includes: a sparsification unit (Sparsification Unit, SU), a matrix processing engine (Matrix Processing Engine, MPE), and a recovery unit (Recovery Unit, CU).

[0038] As Figure 2 shown, the sparsification unit can achieve low-overhead online sparsification. The input activations of the sparsification unit include multiple consecutive frames, and each frame is divided into multiple tokens. The sparsification unit performs inter-frame sparsification and intra-frame sparsification on the input activations including multiple frames to screen and record similar tokens. Figure 2 In an example, the input activations include three frames, and each frame is divided into four tokens.

[0039] The essence of activation sparsification is that if there is a high degree of similarity between two tokens, then we can only calculate one of the tokens and share the result (output token) with the other token.

[0040] During the process of performing inter-frame sparsification, multiple consecutive frames of the input activations are taken as a group, and one frame is selected from each group as a reference frame. The similarities between each token in the reference frame and the corresponding tokens in other frames are calculated respectively; when the similarity between two tokens is greater than the inter-frame similarity threshold, the two tokens are recorded as similar.

[0041] During the process of performing intra-frame sparsification, one token is selected from multiple tokens in each frame as a reference token, and the similarities between other tokens in this frame and the reference token are calculated respectively. When the similarity between two tokens is greater than the intra-frame similarity threshold, the two tokens are recorded as similar.

[0042] If a token in a frame is recorded as similar to the corresponding token in the reference frame during the process of inter-frame sparsification, then during the process of intra-frame sparsification, the similarity between this token and the reference token is set to 0.

[0043] In this embodiment, cosine similarity is used to measure the similarity between two tokens. The inter-frame similarity threshold and the intra-frame similarity threshold are relative; during the QKV / O projection process, a threshold of 0.95 / 0.98 is used; during the FFN process, a threshold of 0.92 is used.

[0044] During the intra-frame sparsification process, an intra-frame index table is used to record whether each token in each frame is similar to the reference token in its frame; during the inter-frame sparsification process, an inter-frame index table is used to record whether each token in each frame is similar to the corresponding token in the reference frame.

[0045] Refer to Figure 2 (a) for an example, where similarity table-1 is the similarity record table for inter-frame sparsification and similarity table-2 is the similarity record table for intra-frame sparsification. During the inter-frame sparsification process, frame F.1 is used as the reference frame to calculate the similarity between each token in frames F.2 and F.3 and the corresponding tokens in the reference frame. During the intra-frame sparsification process, the first token in each frame is used as the reference token. The second and third tokens in frame F.2 and the third token in frame F.3 are determined to be similar during the inter-frame sparsification process; therefore, during the intra-frame sparsification process, their similarity is set to 0.

[0046] The matrix processing engine is used to calculate the tokens to obtain output tokens. Among them, if two tokens are recorded as similar, only the output token of one of the tokens is calculated, and the other is pruned during the sparsification process, and this output token is used as the output token of these two tokens, thereby saving the calculation amount and reducing the calculation time. In a specific embodiment, the matrix processing engine is used to perform a linear operation on the tokens.

[0047] The recovery unit generates output activations based on the calculation results of the matrix processing engine and the records of similar tokens. During the process of the recovery unit generating output activations, the output tokens of each token are determined one by one.

[0048] During the process of determining the output tokens of each token, it is judged whether the token is similar to the reference token in its frame according to the intra-frame index table; if similar, the output token of the reference token is used as the output token of this token; and it is judged whether the token is similar to the corresponding token in the reference frame according to the inter-frame index table; if similar, the output token of the corresponding token in the reference frame is used as the output token of this token.

[0049] Refer to Figure 2Example of (b), where Index Table-2 is the intra-frame index table and Index Table-1 is the inter-frame index table. In both index tables, the ones marked as 1 indicate the existence of similar markers. During the specific implementation process, the address relationship of the output markers of each marker can be found according to the content of the index table. For each pair of similar markers, the corresponding output markers can be copied according to the address relationship to obtain the complete output activation. The Special Function Unit (SFU) is used to perform non-linear operations on the output activation.

[0050] In this embodiment, the matrix processing engine, the sparsification unit, and the restoration unit are arranged in a Field-Programmable Gate Array (FPGA) chip. The matrix processing engine needs to perform a large number of numerical operations. In a specific embodiment, the FPGA chip is AMD V80. The AMD V80 FPGA is equipped with the hardware IP DSP58, which can be configured in multiple computing modes. However, runtime reconfiguration between these configurations is not achievable.

[0051] Therefore, in this embodiment, DSP58 is expanded using logic resources to form the Figure 3 expanded digital signal processing unit (DSP-Expansion, DSP-E) as shown.

[0052] The expanded digital signal processing unit includes two 8-bit integer multipliers (INT8 MUL-0, INT8 MUL-1), two integer adders (INT ADD-0, INT ADD-1), and a digital signal processing unit (DSP58); the expanded digital signal processing unit responds to the mode control signal ( Figure 3 the FP16 signal in it), and switches between the integer operation state and the floating-point operation state.

[0053] In the integer operation state, the expanded digital signal processing unit performs four 8-bit integer multiply-accumulate operations (INT8 MAC) at a time; in the floating-point operation state, the expanded digital signal processing unit performs two 16-bit floating-point multiply-accumulate operations (FP16 MAC) at a time.

[0054] In a specific embodiment, the extended digital signal processing unit includes a scalar-configured DSP58 and some additional configurable circuit designs to provide higher computing performance. In addition to two 8-bit integer multipliers and an integer adder, the extended digital signal processing unit also includes circuits for processing the sign of floating-point numbers (Sign logic, Sign Adjust) and circuits for processing the exponent part of floating-point numbers (Exponent Alignment, Exponent Adjustment). The extended digital signal processing unit can be configured to calculate FP16 multiply-accumulate or INT8 multiply-accumulate at runtime, improving the equivalent computing power of the FPGA under mixed-precision computing requirements.

[0055] The accelerator of the video generation model also includes an embedded processor (Embedded CPU) for performing two-stage sparse-aware scheduling ( Figure 1 the blue identification in), where static compilation realizes operator fusion and weight quantization, and dynamic scheduling adjusts the execution order between operators to achieve maximum utilization through sparsification. Specifically:

[0056] During model optimization and static compilation, the high-level IR (intermediate representation) is converted into an executable IR. Model optimization mainly involves two steps: operator fusion and quantization. We perform operator fusion between composable layer pairs and obtain an IR with quantization parameters. The present invention introduces dynamic adaptive scheduling at runtime, which does not require recompiling instructions. By introducing weight preloading and a priority-based sparse scheduling mechanism, it only adjusts the execution order of operators according to the actual workload to improve computing utilization.

[0057] Taking the original dense model with FP16 precision as the baseline, the accelerator of the video generation model in this embodiment includes mixed-precision quantization and activation sparsification. For the mixed-precision method, we quantize all the linear layers in the DiT block and keep the attention map calculation and the modules outside the DiT at FP16 precision. We prune the similar activation vectors between frames and within frames, and adopt thresholds of 0.95 / 0.98 for QKV / O projection and 0.92 for FFN. As Figure 5 shown, compared with the baseline, the model accuracy of the accelerator of the video generation model in this embodiment is almost the same (only the average loss is 0.008), while the loss of INT8 quantization is 0.042. This embodiment tests the actual video generation for the above different compression methods. By comparison, it can be seen that the video generated by the INT8 quantization method is completely chaotic. While the video quality generated by the accelerator of the video generation model in this embodiment is almost the same as the original method.

[0058] As Figure 6As shown, in the Latte-1 and Open-Sora1.2 models, the present invention compares the acceleration and energy efficiency of F-VGM with those of the NVIDIA 3090 GPU and the current state-of-the-art accelerators. Generally speaking, compared with the NVIDIA 3090 GPU with a peak performance gap of more than 21 times, the present invention achieves an average performance improvement of 1.30 times and an energy efficiency improvement of 4.49 times on the AMD V80 FPGA. Compared with the SOTA accelerators, the F-VGM of the present invention achieves an average performance improvement of 2.84 times and an energy efficiency improvement of 1.71 times.

[0059] The above embodiments are only illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. All equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. An accelerator for a video generation model, characterized in that: include: A sparse unit whose input activations consist of multiple consecutive frames, each of which is split into multiple tokens; The sparsification unit performs inter-frame sparsification and intra-frame sparsification on input activations containing multiple frames to filter and record similar tags; A matrix processing engine is used to calculate the tags to obtain output tags; wherein, if two tags are recorded as similar, the output tag of one of the tags is calculated and used as the output tag of the two tags; The recovery unit generates output activations based on the calculation results of the matrix processing engine and the records of similar labels.

2. The accelerator of a video generation model according to claim 1, characterized in that: In the process of inter-frame sparseness, a frame is selected from multiple consecutive frames of input activation as a reference frame, and the similarity between each mark in the reference frame and the corresponding mark in other frames is calculated respectively; When the similarity between two tags is greater than the inter-frame similarity threshold, the two tags are recorded as similar.

3. The accelerator of a video generation model according to claim 2, characterized in that: During the intra-frame thinning process, one of the multiple markers in each frame is selected as a reference marker, and the similarities between the other markers in the frame and the reference marker are calculated respectively. When the similarity between two markers is greater than the intra-frame similarity threshold, the two markers are recorded as similar.

4. The accelerator of a video generation model according to claim 3, characterized in that: If a marker in a frame is recorded as being similar to a corresponding marker in a reference frame during inter-frame thinning, the similarity between the marker and the reference marker is set to 0 during intra-frame thinning.

5. The accelerator for video generation model according to any one of claims 2 to 4, characterized in that: During the intra-frame thinning process, an intra-frame index table is used to record whether each marker in each frame is similar to the reference marker in the frame where it is located; During the inter-frame thinning process, an inter-frame index table is used to record whether each marker in each frame is similar to the corresponding marker in the reference frame.

6. The accelerator of a video generation model according to claim 5, characterized in that: In the process of generating output activations by the recovery unit, the output tags of each tag are determined one by one; In the process of determining the output mark of each mark, judging whether the mark is similar to the reference mark of the frame in which it is located according to the intra-frame index table; If they are similar, the output mark of the reference mark is used as the output mark of the mark; and judging whether the mark is similar to the corresponding mark in the reference frame according to the inter-frame index table; If they are similar, the output label of the corresponding label in the reference frame is used as the output label of the label.

7. An accelerator for video generation model according to any one of claims 2 to 4, characterized in that: Cosine similarity is used to measure the similarity between two tags.

8. The accelerator for video generation model according to claim 1, characterized in that: The matrix processing engine obtains the output mark of each mark through linear calculation; the matrix processing engine is arranged in a programmable gate array chip.

9. The accelerator for video generation model according to claim 1, characterized in that: The matrix processing engine includes an extended digital signal processing unit; The extended digital signal processing unit includes two 8-bit shaping multipliers, two shaping adders and a digital signal processing unit; the extended digital signal processing unit switches between a shaping operation state and a floating-point operation state in response to a mode control signal; In the shaping operation state, the extended digital signal processing unit performs four 8-bit shaping multiplication and accumulation operations at a time; In the floating point operation state, the extended digital signal processing unit performs two 16-bit floating point multiplication and accumulation operations at a time.

10. A computing system, characterized in that: An accelerator comprising the video generation model described in any one of claims 1 to 9.