Video diffusion transformer acceleration method, apparatus, device, medium and product

CN122534294APending Publication Date: 2026-08-07CHINA MOBILE COMM LTD RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE COMM LTD RES INST
Filing Date
2026-04-30
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]本申请提供一种视频扩散变换器加速方法、装置、设备、介质及产品,用以解决传统缓存策略依赖像素级或特征级相似性计算,无法有效区分噪声模式与语义内容变化,难以保证生成视频的视觉保真度的缺陷

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122534294A_ABST
    Figure CN122534294A_ABST
Patent Text Reader

Abstract

The application provides a video diffusion transformer acceleration method, device, equipment, medium and product, relates to the technical field of video diffusion transformer acceleration, and the method comprises the following steps: acquiring a video latent feature representation of a current diffusion time step; determining a similarity measurement standard based on the frequency domain features of a plurality of frequency bands of the video latent feature representation, and the similarity measurement standard is used for judging whether to cache the video latent feature representation, so as to realize video diffusion transformer acceleration. The application converts the video latent feature representation to the frequency domain for analysis. The similarity measurement standard can judge the video latent feature representation from different frequency bands. The application can effectively distinguish noise patterns and semantic content changes, thereby ensuring the visual fidelity of the generated video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video diffusion converter acceleration technology, and in particular to a video diffusion converter acceleration method, apparatus, device, medium and product. Background Technology

[0002] Existing video generation technologies are mainly based on the theoretical foundation of diffusion models, achieving video synthesis by progressively denoising data through Markov chains. In the field of video diffusion transformer acceleration, acceleration methods include traditional caching strategies, attention sparsity methods, knowledge distillation methods, quantization techniques, etc. However, traditional caching strategies rely on pixel-level or feature-level similarity calculations, which cannot effectively distinguish between noise patterns and semantic content changes. Specifically, although high-frequency components such as random noise patterns or fine textures have small visual impacts, their numerical differences are significant, while low-frequency structural changes such as key object shapes or scene layout changes are easily ignored. The lack of frequency perception causes caching decisions to deviate from the perceptual characteristics of the human visual system, making it difficult to guarantee the visual fidelity of the generated video. Summary of the Invention

[0003] This application provides a video diffusion converter acceleration method, apparatus, device, medium, and product to address the shortcomings of traditional caching strategies that rely on pixel-level or feature-level similarity calculations, which cannot effectively distinguish noise patterns from semantic content changes and thus cannot guarantee the visual fidelity of the generated video.

[0004] This application provides a video diffusion converter acceleration method, which includes the following steps.

[0005] Obtain the latent feature representation of the video at the current diffusion time step; Based on the frequency domain features of the video latent feature representation in multiple frequency bands, a similarity metric is determined. The similarity metric is used to judge whether to cache the video latent feature representation in order to accelerate the video diffusion transformer.

[0006] As one embodiment, determining the similarity metric based on the frequency domain features of the video latent feature representation in multiple frequency bands includes: A three-dimensional Fourier transform is performed on the video latent feature representation to determine the frequency domain features of the video latent feature representation in multiple frequency bands; The video latent features are filtered in the frequency domain of each frequency band using a multi-pass filter based on multiple frequency bands to determine the key frequency components of the video latent features in each frequency band. Based on the key frequency components in each frequency band represented by the latent features of the video, a similarity metric is determined.

[0007] As one embodiment, determining the similarity metric based on the key frequency components of the video latent feature representation in each frequency band includes: The relative distance between the key frequency components of the video latent feature representation in each frequency band and the key frequency components of the cached video latent feature representation in each frequency band is determined, and the maximum value of the relative distance is used as the similarity metric.

[0008] As one embodiment, after determining the relative distance, the method further includes: Based on the similarity metric, the buffer rate during the video diffusion converter acceleration process is determined, and the buffer rate is negatively correlated with the similarity metric.

[0009] As an example, the multiple frequency bands include a first frequency band, a second frequency band, and a third frequency band with sequentially increasing frequency ranges. The frequency domain features of the first frequency band are used to characterize the frequency domain features of video structural information, the frequency domain features of the second frequency band are used to characterize the frequency domain features of video texture details, and the frequency domain features of the third frequency band are used to characterize the frequency domain features of video edge information.

[0010] As one embodiment, after obtaining the video latent feature representation at the current diffusion time step, the method further includes: Based on the video latent feature representation, determine the different frequency spectral energies of the video latent feature representation; Based on the different frequency spectral energies of the video latent feature representation, determine the motion intensity metric information of the video latent feature representation; Based on the motion intensity metric information, the buffer step size during the video diffusion converter acceleration process is determined, and the buffer step size is negatively correlated with the motion intensity metric information.

[0011] This application also provides a video diffusion converter acceleration device, comprising: The acquisition module is used to acquire the video latent feature representation at the current diffusion time step; An acceleration module is used to determine a similarity metric based on the frequency domain features of the video latent feature representation in multiple frequency bands. The similarity metric is used to judge whether to cache the video latent feature representation in order to accelerate the video diffusion transformer.

[0012] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the video diffusion converter acceleration method as described above.

[0013] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video diffusion converter acceleration method as described above.

[0014] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the video diffusion converter acceleration method as described above.

[0015] This application provides a video diffusion transformer acceleration method, apparatus, device, medium, and product. The method includes: acquiring a video latent feature representation at the current diffusion time step; determining a similarity metric based on the frequency domain features of the video latent feature representation in multiple frequency bands, wherein the similarity metric is used to determine whether to cache the video latent feature representation to achieve video diffusion transformer acceleration. This application converts the video latent feature representation to the frequency domain for analysis. The similarity metric allows for evaluation of the video latent feature representation from different frequency bands, determining whether to cache it. Unlike traditional caching strategies that rely on pixel-level or feature-level similarity calculations, this application can effectively distinguish between noise patterns and semantic content changes, thereby ensuring the visual fidelity of the generated video. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a comparative diagram of existing video generation acceleration methods based on caching mechanisms.

[0018] Figure 2 This is a flowchart illustrating the video diffusion converter acceleration method provided in this application.

[0019] Figure 3 This is a schematic diagram of the video diffusion converter provided in this application.

[0020] Figure 4 This is one of the structural schematic diagrams of the video diffusion converter acceleration device provided in this application.

[0021] Figure 5 This is the second schematic diagram of the video diffusion converter acceleration device provided in this application.

[0022] Figure 6 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] During their research, the inventors discovered that with the introduction of a self-attention mechanism, the Video Diffusion Transformer (DiT) architecture has significant advantages over the traditional convolutional deep learning network (U-Net) framework, enabling more effective modeling of global dependencies. The DiT architecture is a hierarchical architecture composed of L cascaded blocks, each containing self-attention, cross-attention, and multilayer perceptron operations. The attention and multilayer perceptron operations dynamically adjust parameters to handle different noise levels during the diffusion process. Pyramid Attention Broadcasting (PAB) identifies significant temporal redundancy by recognizing the U-shaped pattern of attention changes in the diffusion steps and selectively applies broadcasting strategies based on the variance of each attention layer. An adaptive caching strategy dynamically identifies and caches redundant computations, determining whether to reuse features by analyzing the similarity of latent features between consecutive time steps. The Temporal Step Embedding Perceptual Caching (TeaCache) method utilizes the similarity between noisy inputs modulated by temporal step embeddings as an indicator of output similarity. Addressing the high computational complexity of the 3D full attention mechanism in DiT, an attention sparsity method leverages the inherent spatial and temporal redundancy of video data. Attention-based sparse methods (AdaSpa and sparse video generation) introduce dynamic block-level sparsity and online precision-aware search, combined with log-sum-exp caching techniques, significantly improving efficiency while maintaining generation quality. Sparse video generation methods perform adaptive sparsity selection based on frame similarity, leveraging spatial and temporal sparsity and optimizing the sparse kernel, reducing inference time by 50% without sacrificing visual quality. Phased consistency models utilize distillation-based acceleration methods, while Q-DM employs efficient low-bit quantization techniques to reduce computational costs by decreasing the numerical precision of model parameters.

[0025] like Figure 1As shown, in terms of video diffusion transformer acceleration techniques, traditional caching strategies (such as Pyramid Attention Broadcasting (PAB), adaptive caching strategies, and temporal step embedding perceptual caching (adaCache and TeaCache) rely on pixel-level or feature-level similarity calculations, none of which can effectively distinguish between noise patterns and semantic content changes. While attention-sparse methods utilize spatiotemporal redundancy, they lack dynamic adaptation mechanisms for fast-moving scenes, exacerbating motion blur and temporal inconsistency issues. Knowledge distillation methods (such as Phased Consistency Model (PCM)) and quantization techniques (such as Q-DM) can reduce computational costs, but they require additional training stages and rely on external data sources, leading to a significant decrease in generation quality when inference steps are reduced.

[0026] To address this, this application provides a video diffusion converter acceleration scheme that overcomes the shortcomings of traditional caching strategies and, in a preferred embodiment, overcomes the shortcomings of attention sparsity methods, knowledge distillation methods, and quantization techniques.

[0027] Figure 2 This is a flowchart illustrating the video diffusion converter acceleration method provided in this application, as shown below. Figure 2 As shown, this application provides a video diffusion converter acceleration method, including steps S210-S220.

[0028] Step S210: Obtain the video latent feature representation at the current diffusion time step.

[0029] The principle behind accelerating video diffusion transformers is to reduce the effective number of diffusion steps and increase the single-step span. Before acceleration, the video diffusion transformer needs to start from the maximum step number T and strictly follow the order t=T,T-1,…,1 for denoising. After acceleration, the video diffusion transformer performs skip-step inference on sparse subsequences (such as t=1000,800,600,…,0). By utilizing high-order numerical solvers or uniform distillation techniques, it can accurately predict the denoising direction even with large step sizes (i.e., the noise span between adjacent steps t→tk becomes larger), thereby compressing the number of iterations from hundreds of steps to tens or even a few steps, achieving a significant speedup in video generation.

[0030] The diffusion time step refers to the discrete stage of noise addition or removal. The preset noise level index can be represented as t. The larger the step number t, the more noise the current video latent variables contain, and the more chaotic the picture is. The smaller the step number t, the closer the picture is to the clear final result. The video generation process is the reverse process of iterative evolution from high step number to low step number.

[0031] Video latent feature representation refers to the low-dimensional tensor obtained after the original high-dimensional video pixels are compressed by the encoder. Essentially, it maps the massive pixel data of the video to a smaller feature map, shifting the computational load from the original pixel domain to the lightweight latent domain while preserving the spatiotemporal semantics.

[0032] Step S220: Based on the frequency domain features of the video latent feature representation in multiple frequency bands, a similarity metric is determined. The similarity metric is used to judge whether to cache the video latent feature representation in order to accelerate the video diffusion transformer.

[0033] To address the limitations of spatial domain metrics in capturing meaningful changes in perception, this application introduces a frequency-aware buffer. It performs diffusion transformer inference by analyzing the similarity of latent video features in the spatial frequency domain. Specifically, it determines the frequency domain features of the latent video feature representation in multiple frequency bands by performing a Fourier transform on the latent video feature representation.

[0034] Optionally, multiple frequency bands refer to multiple different frequency ranges, such as high frequency, mid frequency and low frequency. The frequency domain features of multiple frequency bands are used to characterize different types of information represented by the latent features of the video, such as high frequency components, mid frequency components and low frequency components.

[0035] This application establishes a frequency-specific similarity metric (FreqDiff) that can independently evaluate the latent feature representation of video within each frequency band and differentiate the changes in the latent feature representation of video across different frequency ranges. This enables the video diffusion transformer to prioritize perceptual key features that affect the quality of video generation while adaptively filtering out less relevant variations in other frequency bands.

[0036] Optionally, the similarity metric is used to evaluate the rate of change of the video latent feature representation in different frequency ranges. If the similarity metric exceeds the metric threshold, it is determined that the rate of change of the video latent feature representation in different frequency ranges is high, indicating that the video content is evolving rapidly, and it is determined that the video latent feature representation needs to be cached. Otherwise, it is determined that the video content has temporal redundancy, and the cached video latent feature representation can be directly reused without caching the video latent feature representation of the current diffusion time step.

[0037] Understandably, this application transforms the video latent feature representation into the frequency domain for analysis. Through similarity measurement standards, the video latent feature representation can be evaluated from different frequency bands to determine whether to cache the video latent feature representation. Unlike traditional caching strategies that rely on pixel-level or feature-level similarity calculations, this application can effectively distinguish noise patterns from semantic content changes, thereby ensuring the visual fidelity of the generated video.

[0038] As one embodiment, determining the similarity metric based on the frequency domain features of the video latent feature representation in multiple frequency bands includes: A three-dimensional Fourier transform is performed on the video latent feature representation to determine the frequency domain features of the video latent feature representation in multiple frequency bands; The video latent features are filtered in the frequency domain of each frequency band using a multi-pass filter based on multiple frequency bands to determine the key frequency components of the video latent features in each frequency band. Based on the key frequency components in each frequency band represented by the latent features of the video, a similarity metric is determined.

[0039] Optionally, let the latent features of the video at the current diffusion time step t be represented as: , Represents the set of real numbers. Denotes the diffusion time step dimension, and denotes the total. One diffusion time step, Height refers to the number of pixels or spatial grid points in the vertical direction of the feature map of each frame of the video. Width refers to the number of pixels or spatial grid points in the horizontal direction of the feature map of each frame of the video. The number of channels refers to the number of feature maps in each frame of the video at each spatial location. The dimension of the feature vectors on the ).

[0040] A three-dimensional Fourier transform is performed on the video latent feature representation to decompose its spatiotemporal characteristics into the spatiotemporal frequency domain. The expression for the three-dimensional Fourier transform of the video latent feature representation is as follows: (1) (2) in, The latent features of the video are represented in the frequency domain across multiple frequency bands. It is a three-dimensional Fourier transform function. For the Fast Fourier Transform along the time dimension, For the Fast Fourier Transform along the vertical spatial dimension, For the Fast Fourier Transform along the horizontal spatial dimension, This indicates element-wise multiplication.

[0041] Optionally, a three-dimensional Fourier transform is performed on the video latent feature representation to determine the frequency domain features of the video latent feature representation at high, mid, and low frequencies, respectively. The frequency domain features of the video latent feature representation in each frequency band are filtered using a multi-pass filter based on multiple frequency bands, i.e., a multi-pass filter based on high, mid, and low frequencies (…). , , The video latent features are represented in the frequency domain of each frequency band and filtered. Through the filtering process, key feature components are further separated in the three-dimensional spectral domain, thereby more effectively handling diverse visual and dynamic patterns in the video generation process.

[0042] The main structural changes in the video are concentrated in the low-frequency temporal and spatial domains. The low-frequency temporal domain can effectively capture near-static background areas, slow motion, and gradual lighting changes, while the low-frequency spatial domain mainly represents the outlines of objects and global geometry. In contrast, during video generation, when the changes in low-frequency spatial features are minimal at a certain stage, more attention should be paid to feature updates driven by high-frequency components. High-frequency components typically correspond to sudden motion, rapidly changing boundaries, and fine spatial details such as textures and sharp edges, playing a crucial role in determining the quality of the generated output. This application's embodiments use multi-pass filters across multiple frequency bands to filter the video's latent feature representations in the frequency domain of each band, better capturing and utilizing these changes, and further separating key feature components in the three-dimensional spectral domain, thereby more effectively handling diverse visual and dynamic patterns during video generation.

[0043] Optionally, the video latent features represented in each frequency band are filtered using a multi-pass filter based on multiple frequency bands. The calculation formula for determining the key frequency components of the video latent features represented in each frequency band is as follows: (3) (4) in, Indicates the frequency band type; The spectral range of each frequency band is defined: , , ; three-dimensional frequency coordinates Centered on, frequency coordinates These correspond to time frequency, vertical spatial frequency, and horizontal spatial frequency, respectively. This represents the Hadamard product (element-by-element product). It is an indicator function used to select frequency components within a specified range for each frequency band.

[0044] Optionally, a similarity metric is determined based on the key frequency components of the video latent feature representation in each frequency band, including: determining the similarity metric based on the amount of change between the key frequency components of the video latent feature representation and the cached video latent feature representation in each frequency band.

[0045] Understandably, this application performs a three-dimensional Fourier transform on the video latent feature representation to determine the frequency domain features of the video latent feature representation in multiple frequency bands; filters the frequency domain features of the video latent feature representation in each frequency band based on a multi-pass filter of multiple frequency bands to determine the key frequency components of the video latent feature representation in each frequency band; and determines a similarity metric based on the key frequency components of the video latent feature representation in each frequency band. This allows for independent evaluation of the video latent feature representation within each frequency band, differentiated processing of the changes in the video latent feature representation in different frequency ranges, enabling the video diffusion transformer to prioritize perceptual key features that affect the quality of video generation, while adaptively filtering out less relevant variations in other frequency bands.

[0046] As one embodiment, determining the similarity metric based on the key frequency components of the video latent feature representation in each frequency band includes: The relative distance between the key frequency components of the video latent feature representation in each frequency band and the key frequency components of the cached video latent feature representation in each frequency band is determined, and the maximum value of the relative distance is used as the similarity metric.

[0047] Optionally, the relative distance includes the relative L1 distance, which is a relative error measure after normalizing the ordinary L1 distance (sum of absolute errors), converting the absolute error into a proportional error relative to the magnitude of the data itself.

[0048] Optionally, the formula for calculating the similarity metric is as follows: (5) in, As a similarity metric, (.) indicates extracting the real part of the complex number's spectrum. Let the real part of the key frequency component in each frequency band be represented by the latent features of the video. The cached video latent features are represented by the real parts of the key frequency components in each frequency band, and the cached video latent features are represented by the spread time step of the previous distance calculation. The corresponding latent feature representation of the video, To ensure that the denominator is a very small number, we must prevent it from being zero.

[0049] Similarity metrics, as a normalization measure, are used to measure the rate of change of the latent representation of a video in the spatiotemporal domain. A higher normalization metric value indicates that the video content is evolving rapidly, and the corresponding... A cached state is one rich in information; conversely, a lower value indicates temporal redundancy, allowing for safe reuse of previously computed features (the previous diffusion time step involved in distance calculation). (Features in the corresponding video latent feature representation), without needing to be recalculated.

[0050] It is understood that, by calculating the relative distance between the key frequency components of the video latent feature representation in each frequency band and the key frequency components of the cached video latent feature representation in each frequency band, and using the maximum value of the relative distance as a similarity metric, this embodiment can adaptively identify the most informative potential changes in different frequency bands based on the dynamic characteristics of the diffusion process. This ensures that important latent features are preserved throughout the entire generation process, thereby improving inference efficiency while also guaranteeing the visual fidelity of the generated results.

[0051] As one embodiment, after determining the relative distance, the method further includes: Based on the similarity metric, the buffer rate during the video diffusion converter acceleration process is determined, and the buffer rate is negatively correlated with the similarity metric.

[0052] Optionally, based on the similarity metric, the cache rate during the video diffusion transformer acceleration process is determined, including: mapping the similarity metric to the cache rate according to the mapping rules. The cache rate is used to determine the number of diffusion time steps that reuse the same cached data after the current diffusion time step. The higher the similarity metric, the lower the cache rate, and the more frequent the distance calculation is required.

[0053] Optionally, to determine whether to reuse cached data at the current diffusion time step, this application defines the above mapping rules, which incorporate similarity metrics. Mapping to cache rate The mapping rule consists of a set of buffer rates derived from the original denoising schedule (i.e., the total number of diffusion steps) and corresponding metric thresholds, used to select an appropriate buffer rate, as shown in the following expression: (6) in, This indicates that the distance metric is mapped to the corresponding cache rate. (.), ∈ Represents the non-negative distance metric at the current diffusion time step t; Γ This represents the buffer rate, and N is the total number of steps in the original diffusion process.

[0054] For the denoising step accelerated by the video diffusion transformer, when the number of steps skipped... Less than cache rate At this time, a reuse strategy is activated, which achieves a balance between maintaining generation quality and dynamically controlling computation frequency and cache reuse. Specifically, the video latent feature representation is obtained by using the latent feature representation of the previous video. Latent feature representation of each preceding adjacent video residual Add them together to get the result; otherwise, when At that time, the video potential representation will be recalculated based on the following cache scheduling: (7) in, This represents the l-th processing block in the video spread converter.

[0055] It is understood that, in accordance with the mapping rules, the similarity metric is mapped to the cache rate. The cache rate is used to determine the number of diffusion time steps after the current diffusion time step that reuse the same cached data. The higher the similarity metric, the lower the cache rate, and the more frequent the distance calculation is required. This can accelerate the video diffusion transformer while avoiding the loss of video changes and ensuring the visual fidelity of the generated video.

[0056] As an example, the multiple frequency bands include a first frequency band, a second frequency band, and a third frequency band with sequentially increasing frequency ranges. The frequency domain features of the first frequency band are used to characterize the frequency domain features of video structural information, the frequency domain features of the second frequency band are used to characterize the frequency domain features of video texture details, and the frequency domain features of the third frequency band are used to characterize the frequency domain features of video edge information.

[0057] Optionally, the first frequency band is low frequency, the second frequency band is mid frequency, and the third frequency band is high frequency.

[0058] Understandably, this application decomposes the latent feature representation of video into frequency domain features of low, medium and high frequency sub-bands, establishes a frequency band-specific similarity metric, accurately distinguishes important perceptual features from irrelevant changes, solves the problem of noise and semantic changes being confused in spatial domain metrics, and is conducive to achieving accurate frequency band separation and processing through bandpass filters.

[0059] As one embodiment, after obtaining the video latent feature representation at the current diffusion time step, the method further includes: Based on the video latent feature representation, determine the different frequency spectral energies of the video latent feature representation; Based on the different frequency spectral energies of the video latent feature representation, determine the motion intensity metric information of the video latent feature representation; Based on the motion intensity metric information, the buffer step size during the video diffusion converter acceleration process is determined, and the buffer step size is negatively correlated with the motion intensity metric information.

[0060] The steps provided in this application embodiment can be executed synchronously with step S220. Step S220 analyzes the changes in the video latent feature representation from the spatial frequency domain dimension, while this application embodiment analyzes the changes in the video latent feature representation from the temporal frequency domain dimension. Through the collaborative work of the two paths, refined analysis and intelligent cache management of the video latent feature representation are achieved.

[0061] To ensure the generation of high-quality videos in scenes with rapid motion, this application proposes a motion energy regularization strategy that adaptively adjusts the buffer step size based on motion intensity. Since video sequences with intense motion require smaller buffer steps to maintain temporal coherence, while static or stable scenes allow for longer buffer steps to improve computational efficiency, this application transforms latent video features into the time-frequency domain using Discrete Fourier Transform to calculate the High-to-Low Frequency Energy Ratio (HLER). HLER quantifies motion intensity by comparing the energy distribution between high-frequency and low-frequency bands.

[0062] Specifically, firstly, a Fast Fourier Transform is performed on the latent feature representation of the video along the time dimension T to transform it into the frequency domain: (8) The spectral energy distribution is obtained by calculating the sum of squares of the real and imaginary parts in the frequency domain representation: (9) in, This represents the element-wise square, used to capture the energy (i.e., power) of signals in different frequency bands. To characterize motion intensity, embodiments of this application use a cutoff index. The frequency spectrum is divided into low-frequency and high-frequency parts. Therefore, low-frequency energy... and high frequency energy The definition is as follows: (10) in, Indicates extracting the first Frequency spectral energy data at each time step The energy ratio of high frequency to low frequency is used as the main measure of motion intensity, and the calculation is as follows: (11) A higher HLER value indicates more intense motion in the video sequence.

[0063] This proposal implements a dynamic caching strategy that adaptively adjusts the time cache step size based on the scenario. : (12) in, This indicates the adjusted cache step size. The baseline cache step size, Used to control the sensitivity of adaptation. Dynamic caching strategies can ensure shorter cache steps in fast-moving scenarios (i.e., high HLER values) to maintain time fidelity, while extending cache intervals in static or less dynamic scenarios (i.e., low HLER values) to improve computational efficiency.

[0064] It is understood that this application determines the different frequency spectrum energies of the video latent feature representation based on the video latent feature representation; determines the motion intensity metric information of the video latent feature representation based on the different frequency spectrum energies of the video latent feature representation; and determines the buffer step size in the video diffusion transformer acceleration process based on the motion intensity metric information. The buffer step size is negatively correlated with the motion intensity metric information, which can automatically improve the calculation accuracy and ensure timing consistency in high-speed motion scenarios, optimize resource allocation and improve processing efficiency in static scenarios, and achieve an intelligent balance between quality and speed. This solves the defect that although the attention sparsity method utilizes spatiotemporal redundancy, it lacks a dynamic adaptation mechanism for fast motion scenarios, which leads to the aggravation of motion blur and timing inconsistency problems.

[0065] Figure 3 This is a schematic diagram of the video diffusion converter provided in this application, as shown below. Figure 3 As shown, the video diffusion transformer is composed of multiple basic transformation units (Transformer Blocks) stacked together. Each basic transformation unit is equipped with a cache acceleration module (AdaFreq Cache). The cache acceleration module is used to execute the video diffusion transformer acceleration method provided in this application. The cache acceleration module does not require training and is plug-and-play. It solves the problem that knowledge distillation methods and quantization techniques require additional training stages and rely on external data sources, which leads to a significant decrease in generation quality when reducing inference steps.

[0066] Existing basic transform units analyze the spatial correlation between input and output embeddings during the diffusion process, relying on pixel-level or feature-level similarity calculations. This fails to effectively distinguish between noise patterns and semantic content changes, limiting their ability to capture meaningful perceptual changes. This application's embodiment adds a buffer acceleration module to the basic transform unit. By decomposing the video's latent feature representation into low, medium, and high-frequency domain features and establishing a frequency band-specific similarity metric, it can accurately distinguish between perceptually important features and irrelevant changes. Through three-dimensional Fourier transform and bandpass filtering techniques, it achieves precise modeling and control of the spatial and temporal frequency domains, overcoming the shortcomings of traditional methods in handling frequency components. By quantifying motion intensity using the temporal high-low frequency energy ratio (HLER), it constructs and develops an adaptive buffering strategy for motion intensity, solving the temporal consistency problem in fast-moving scenes, dynamically adjusting the buffer interval, and avoiding motion blur.

[0067] In summary, this application involves technological innovation and system integration in four core sub-fields: computer vision and video generation technology, artificial intelligence and deep learning architecture, high-performance computing and hardware co-optimization, and digital signal processing and frequency domain analysis. In computer vision and video generation technology, this application is deeply applied to a new generation video generation framework based on a diffusion model, focusing on solving key technical challenges in high-resolution (720p and above) and multi-frame-rate (45-129fps) video synthesis, covering latent diffusion model optimization, accurate modeling of the denoising process, and spatiotemporal consistency preservation and enhancement. In artificial intelligence and deep learning architecture, this application optimizes the core bottlenecks of Transformer-based generation models, particularly accelerating the inference process of the diffusion transformer, covering self-attention mechanisms, cross-attention modules, and computational redundancy elimination and feature reuse strategies for multilayer perceptrons. In digital signal processing and frequency domain analysis, this application applies digital signal processing technology to video latent feature analysis, achieving accurate modeling and control of the spatial and temporal frequency domains through methods such as three-dimensional Fourier transform, bandpass filtering, and energy spectrum analysis. In the area of ​​low-latency video generation applications, this application provides a low-latency, high-efficiency video generation solution for application scenarios with extremely high real-time requirements, including interactive media creation, virtual reality, autonomous driving simulation, and mobile and edge computing. Through deep integration and collaborative innovation across these multiple technology fields, this application constructs a complete and efficient video generation acceleration solution that significantly reduces the generation latency of 720p video while maintaining visual quality.

[0068] The video diffusion converter acceleration device provided in this application is described below. The video diffusion converter acceleration device described below can be referred to in correspondence with the video diffusion converter acceleration method described above.

[0069] Figure 4This is one of the structural schematic diagrams of the video diffusion converter acceleration device provided in this application, such as... Figure 4 As shown, this application also provides a video diffusion converter acceleration device, comprising: The acquisition module 410 is used to acquire the video latent feature representation at the current diffusion time step; The acceleration module 420 is used to determine a similarity metric based on the frequency domain features of the video latent feature representation in multiple frequency bands. The similarity metric is used to judge whether to cache the video latent feature representation in order to accelerate the video diffusion transformer.

[0070] As one embodiment, the acceleration module 420 is used for: A three-dimensional Fourier transform is performed on the video latent feature representation to determine the frequency domain features of the video latent feature representation in multiple frequency bands; The video latent features are filtered in the frequency domain of each frequency band using a multi-pass filter based on multiple frequency bands to determine the key frequency components of the video latent features in each frequency band. Based on the key frequency components in each frequency band represented by the latent features of the video, a similarity metric is determined.

[0071] As one embodiment, the acceleration module 420 is used for: The relative distance between the key frequency components of the video latent feature representation in each frequency band and the key frequency components of the cached video latent feature representation in each frequency band is determined, and the maximum value of the relative distance is used as the similarity metric.

[0072] As one embodiment, the acceleration module 420 is used for: Based on the similarity metric, the buffer rate during the video diffusion converter acceleration process is determined, and the buffer rate is negatively correlated with the similarity metric.

[0073] As an example, the multiple frequency bands include a first frequency band, a second frequency band, and a third frequency band with sequentially increasing frequency ranges. The frequency domain features of the first frequency band are used to characterize the frequency domain features of video structural information, the frequency domain features of the second frequency band are used to characterize the frequency domain features of video texture details, and the frequency domain features of the third frequency band are used to characterize the frequency domain features of video edge information.

[0074] As one embodiment, the acceleration module 420 is used for: Based on the video latent feature representation, determine the different frequency spectral energies of the video latent feature representation; Based on the different frequency spectral energies of the video latent feature representation, determine the motion intensity metric information of the video latent feature representation; Based on the motion intensity metric information, the buffer step size during the video diffusion converter acceleration process is determined, and the buffer step size is negatively correlated with the motion intensity metric information.

[0075] Figure 5 This is the second structural schematic diagram of the video diffusion converter acceleration device provided in this application, as shown below. Figure 5 As shown, this application also provides a video diffusion converter acceleration device, including a cache acceleration module. The cache acceleration module is used to obtain the video latent feature representation of the current diffusion time step; and to determine a similarity metric based on the frequency domain features of the video latent feature representation in multiple frequency bands. The similarity metric is used to judge whether to cache the video latent feature representation in order to achieve video diffusion converter acceleration.

[0076] As one embodiment, the cache acceleration module is further configured to: A three-dimensional Fourier transform is performed on the video latent feature representation to determine the frequency domain features of the video latent feature representation in multiple frequency bands; The video latent features are filtered in the frequency domain of each frequency band using a multi-pass filter based on multiple frequency bands to determine the key frequency components of the video latent features in each frequency band. Based on the key frequency components in each frequency band represented by the latent features of the video, a similarity metric is determined.

[0077] As one embodiment, the cache acceleration module is further configured to: The relative distance between the key frequency components of the video latent feature representation in each frequency band and the key frequency components of the cached video latent feature representation in each frequency band is determined, and the maximum value of the relative distance is used as the similarity metric.

[0078] As one embodiment, the cache acceleration module is further configured to: Based on the similarity metric, the buffer rate during the video diffusion converter acceleration process is determined, and the buffer rate is negatively correlated with the similarity metric.

[0079] As an example, the multiple frequency bands include a first frequency band, a second frequency band, and a third frequency band with sequentially increasing frequency ranges. The frequency domain features of the first frequency band are used to characterize the frequency domain features of video structural information, the frequency domain features of the second frequency band are used to characterize the frequency domain features of video texture details, and the frequency domain features of the third frequency band are used to characterize the frequency domain features of video edge information.

[0080] As one embodiment, the cache acceleration module is further configured to: Based on the video latent feature representation, determine the different frequency spectral energies of the video latent feature representation; Based on the different frequency spectral energies of the video latent feature representation, determine the motion intensity metric information of the video latent feature representation; Based on the motion intensity metric information, the buffer step size during the video diffusion converter acceleration process is determined, and the buffer step size is negatively correlated with the motion intensity metric information.

[0081] It should be noted that the video diffusion converter acceleration device provided in this application has the same technical effects as the video diffusion converter acceleration method, which will not be elaborated further.

[0082] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a video diffusion converter acceleration method, which includes: Obtain the latent feature representation of the video at the current diffusion time step; Based on the frequency domain features of the video latent feature representation in multiple frequency bands, a similarity metric is determined. The similarity metric is used to judge whether to cache the video latent feature representation in order to accelerate the video diffusion transformer.

[0083] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0084] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the video diffusion converter acceleration method provided by the above methods, the method including: Obtain the latent feature representation of the video at the current diffusion time step; Based on the frequency domain features of the video latent feature representation in multiple frequency bands, a similarity metric is determined. The similarity metric is used to judge whether to cache the video latent feature representation in order to accelerate the video diffusion transformer.

[0085] In another aspect, this application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the video diffusion converter acceleration method provided by the methods described above, the method comprising: Obtain the latent feature representation of the video at the current diffusion time step; Based on the frequency domain features of the video latent feature representation in multiple frequency bands, a similarity metric is determined. The similarity metric is used to judge whether to cache the video latent feature representation in order to accelerate the video diffusion transformer.

[0086] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0087] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A video diffusion converter acceleration method, characterized in that, include: Obtain the latent feature representation of the video at the current diffusion time step; Based on the frequency domain features of the video latent feature representation in multiple frequency bands, a similarity metric is determined. The similarity metric is used to judge whether to cache the video latent feature representation in order to accelerate the video diffusion transformer.

2. The video diffusion converter acceleration method according to claim 1, characterized in that, The determination of a similarity metric based on the frequency domain features of the video latent features represented in multiple frequency bands includes: A three-dimensional Fourier transform is performed on the video latent feature representation to determine the frequency domain features of the video latent feature representation in multiple frequency bands; The video latent features are filtered in the frequency domain of each frequency band using a multi-pass filter based on multiple frequency bands to determine the key frequency components of the video latent features in each frequency band. Based on the key frequency components in each frequency band represented by the latent features of the video, a similarity metric is determined.

3. The video diffusion converter acceleration method according to claim 2, characterized in that, The determination of similarity metrics based on the key frequency components in each frequency band represented by the latent features of the video includes: The relative distance between the key frequency components of the video latent feature representation in each frequency band and the key frequency components of the cached video latent feature representation in each frequency band is determined, and the maximum value of the relative distance is used as the similarity metric.

4. The video diffusion converter acceleration method according to claim 3, characterized in that, After determining the relative distance, the process also includes: Based on the similarity metric, the buffer rate during the video diffusion converter acceleration process is determined, and the buffer rate is negatively correlated with the similarity metric.

5. The video diffusion converter acceleration method according to any one of claims 1 to 4, characterized in that, Multiple frequency bands include a first frequency band, a second frequency band, and a third frequency band with sequentially increasing frequency ranges. The frequency domain features of the first frequency band are used to characterize the frequency domain features of video structural information, the frequency domain features of the second frequency band are used to characterize the frequency domain features of video texture details, and the frequency domain features of the third frequency band are used to characterize the frequency domain features of video edge information.

6. The video diffusion converter acceleration method according to claim 1, characterized in that, After obtaining the video latent feature representation at the current diffusion time step, the method further includes: Based on the video latent feature representation, determine the different frequency spectral energies of the video latent feature representation; Based on the different frequency spectral energies of the video latent feature representation, determine the motion intensity metric information of the video latent feature representation; Based on the motion intensity metric information, the buffer step size during the video diffusion converter acceleration process is determined, and the buffer step size is negatively correlated with the motion intensity metric information.

7. A video diffusion converter acceleration device, characterized in that, include: The acquisition module is used to acquire the video latent feature representation at the current diffusion time step; An acceleration module is used to determine a similarity metric based on the frequency domain features of the video latent feature representation in multiple frequency bands. The similarity metric is used to judge whether to cache the video latent feature representation in order to accelerate the video diffusion transformer.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the video diffusion converter acceleration method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the video diffusion converter acceleration method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video diffusion converter acceleration method as described in any one of claims 1 to 6.