Model reasoning acceleration method and system, electronic equipment, storage medium and product
By obtaining and using intermediate variables with similarity not lower than the preset threshold in the calculation module, the problem of slow model inference speed caused by fast cache expansion is solved, and the model inference speed is improved.
Patent Information
- Application Number
- CN202510724702.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-30
AI Technical Summary
In the prior art, the rapid expansion of caches of computing devices leads to the problem of slow model inference speed.
By obtaining the intermediate variables of the serialized model in the first computing module, including a key-value copy, an intermediate layer latent feature and a deep output feature, it is determined that the feature whose similarity is not lower than the preset similarity threshold is the input of the second computing module to reduce redundant calculations.
The redundant calculation of the second computing module is reduced, thereby speeding up the model inference.
Smart Images

Figure CN120258152A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, system, electronic device, storage medium and product for accelerating model inference. Background Art
[0002] In related model inference acceleration solutions, sufficient space is usually reserved in the memory of a computing device to cache data during the model inference calculation process. During the inference calculation process, the cache of the computing device expands rapidly, resulting in a slow model inference speed. Summary of the Invention
[0003] This application provides a method, system, electronic device, storage medium and product for accelerating model inference, so as to at least solve the problem in related technologies that the cache of a computing device expands rapidly, resulting in a slow model inference speed.
[0004] This application provides a method for accelerating model inference, including: Obtaining intermediate variables of a serialized model in a first computing module, where the intermediate variables include at least one of a key-value copy, an intermediate-layer latent feature, and a deep output feature. The intermediate-layer latent feature is a feature with a similarity not lower than a preset similarity threshold between the first computing module and a second computing module. The intermediate-layer latent feature is determined by a shallow computing block in the first computing module, and the deep output feature is determined by a deep computing block in the first computing module; Determining that a feature with a similarity not lower than a preset similarity threshold is an input to a deep computing block in the second computing module, so that the second computing module obtains a prediction result of the serialized model.
[0005] This application provides a system for accelerating model inference. The system includes: a first computing module, a cache module, and at least one second computing module. The cache module is communicatively connected to the first computing module and the second computing module respectively; The first computing module is used for performing inference calculation on a serialized model to obtain intermediate variables, where the intermediate variables include at least one of a key-value copy, an intermediate-layer latent feature, and a deep output feature; The cache module is used for caching the intermediate variables obtained by the inference calculation of the first computing module; The second computing module is used for performing inference calculation on the serialized model according to the intermediate variables in the cache module.
[0006] This application also provides a device for accelerating model inference, including: An acquisition unit for acquiring intermediate variables of a serialized model in a first computing module, where the intermediate variables include at least one of a key-value copy, an intermediate-layer latent feature, and a deep output feature. The intermediate-layer latent feature is a feature with a similarity not lower than a preset similarity threshold in the first computing module and the second computing module, the intermediate-layer latent feature is determined by a shallow computing block in the first computing module, and the deep output feature is determined by a deep computing block in the first computing module; A determination unit for determining that a feature with a similarity not lower than a preset similarity threshold is an input to a deep computing block in the second computing module, so that the second computing module obtains a prediction result of the serialized model.
[0007] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above model inference acceleration methods when executing the computer program.
[0008] This application also provides a computer-readable storage medium storing a computer program, where the computer program implements the steps of any of the above XX methods when executed by a processor.
[0009] This application also provides a computer program product including a computer program that implements the steps of any of the above model inference acceleration methods when executed by a processor.
[0010] Through this application, it includes acquiring intermediate variables of a serialized model in a first computing module, where the intermediate variables include at least one of a key-value copy, an intermediate-layer latent feature, and a deep output feature. The intermediate-layer latent feature is a feature with a similarity not lower than a preset similarity threshold in the first computing module and the second computing module, the intermediate-layer latent feature is determined by a shallow computing block in the first computing module, and the deep output feature is determined by a deep computing block in the first computing module; determining that a feature with a similarity not lower than a preset similarity threshold is an input to a deep computing block in the second computing module, so that the second computing module obtains a prediction result of the serialized model, which solves the technical problem that the cache of the second computing module expands rapidly in the related solution, resulting in a slow model inference speed, and achieves the technical effect of reducing redundant calculations of the second computing module and thus accelerating the model inference speed. Description of the Drawings
[0011] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1Schematic flowchart of a model inference acceleration method provided by an embodiment of the present application; Figure 2 Schematic structural diagram of a model inference acceleration system provided by an embodiment of the present application; Figure 3 Schematic structural diagram of a decision control module provided by an embodiment of the present application; Figure 4 Schematic structural diagram of a data frame format provided by an embodiment of the present application; Figure 5 Schematic structural diagram of a model inference acceleration device provided by an embodiment of the present application. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.
[0014] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and not to describe a specific order or sequence.
[0015] To facilitate those skilled in the art to better understand the technical solutions described in the embodiments of the present disclosure, before introducing the embodiments of the present disclosure, the technical terms in the embodiments of the present disclosure are explained as follows.
[0016] Graphics Processing Unit (GPU): A processor used to quickly process image and video data.
[0017] CXL: An open industry standard for high-bandwidth and low-latency device interconnection, allowing fast and reliable data transmission between different components within a computer system, aiming to solve bottleneck problems in high-performance computing, including memory capacity, memory bandwidth, and input / output latency, etc. CXL can also achieve memory expansion and memory sharing, and can communicate with computing accelerators such as GPUs and Field-Programmable Gate Arrays (FPGAs) and other peripherals, providing a faster and more flexible data exchange and processing method.
[0018] Type 2 Devices: Type 2 devices support both CXL.mem and CXL.cache protocols.
[0019] Device Memory: Refers to the memory visible to the device, such as the High Bandwidth Memory (HBM) carried by the GPU.
[0020] Denoising Diffusion Probabilistic Models (DDPM): A generative model based on the diffusion process. By gradually adding Gaussian noise to data (such as images) and then training a neural network to gradually remove the noise, it can recover the original data from pure random noise. It has good generation quality but usually has a slow inference speed because multiple denoising iterations are required.
[0021] Denoising Diffusion Implicit Models (DDIM): DDIM is an improved version of DDPM. It introduces a non - Markovian denoising process, enabling image generation to be completed with fewer steps during the inference stage, thus significantly accelerating the inference speed.
[0022] Spatial - Temporal Similarity Selection (STSS): Aims to improve cache utilization efficiency and accelerate model inference speed by selecting and reusing features that are highly similar both spatially and temporally. However, for input scenarios with different complexities, a fixed cache reuse strategy may lead to unreasonable resource allocation, thereby affecting the overall performance of the model.
[0023] Extended Memory: Through providing high - bandwidth and low - latency connections, it enables the host processor to communicate more efficiently with devices such as accelerators, memory buffers, and intelligent input / output devices. CXL extended memory addresses the limitations of traditional memory expansion methods by providing high - bandwidth, low - latency connections and efficient memory management.
[0024] (Diffusion Transformer, DiT): A generative model that combines the Diffusion Model and the transformer architecture, mainly applied to high - resolution image and video generation tasks.
[0025] The core idea of DiT is to replace the U-Net structure in traditional diffusion models (such as DDPM and DDIM) with the self-attention mechanism of the transformer to improve the generation quality and diversity. However, the DiT model has the following problems and challenges. First, low cache management efficiency leads to resource waste: In existing cache reuse technologies (such as STSS feature cache), the cached data expands rapidly during the inference process, occupying a large amount of video memory resources, which limits the real-time processing ability of the model for high-resolution images. Second, it is difficult to adapt the dynamic strategy for specific instances: For different input instances (such as complex scenes vs. simple objects), it is difficult to adaptively adjust the cache reuse strategy, resulting in unreasonable allocation of computing resources in some scenarios (such as excessive reuse of low-frequency features in simple scenes and lack of high-frequency features in complex scenes). Experiments show that the computational redundancy rate in simple scenes under the fixed strategy is as high as 40%, while the key feature omission rate in complex scenes exceeds 15%. Finally, the adaptability of the hardware communication protocol to cache expansion is insufficient: The data frame format of the traditional communication protocol is not optimized for the DiT cache expansion requirements, resulting in an increase in communication latency during high-frequency cache updates, which restricts the improvement of the end-to-end inference speed. There is bandwidth competition in the existing protocol during multi-device collaboration (such as computing devices and cache devices), and it is difficult to meet the demand for real-time feature synchronization in the later stage of DiT denoising.
[0026] In the related model inference acceleration scheme, DiT is accelerated by exploring the feature similarity of adjacent time steps. Specifically, it first identifies the most similar features in structure, then caches and reuses these highly similar features to reduce redundant calculations, thereby accelerating DiT while ensuring the consistency of the generation results with the original model to the greatest extent. At the same time, a lightweight decision network for specific instances is designed, which can dynamically allocate resources and provide better content quality. However, there is no independent decision control in the related model inference scheme, and the cache reuse strategy is fixed. During the calculation process, the cache expands rapidly, resulting in slow model inference speed.
[0027] Through this application, including obtaining the intermediate variables of the serialized model in the first computing module, the intermediate variables include at least one of key-value copies, intermediate-layer latent features, and deep output features. The intermediate-layer latent features are the features with a similarity not lower than a preset similarity threshold in the first computing module and the second computing module. The intermediate-layer latent features are determined by the shallow computing block in the first computing module, and the deep output features are determined by the deep computing block in the first computing module. It is determined that the features with a similarity not lower than the preset similarity threshold are the inputs of the deep computing block in the second computing module, so that the second computing module obtains the prediction result of the serialized model, which solves the technical problem of rapid cache expansion in the second computing module in the related scheme, resulting in slow model inference speed, and achieves the technical effect of reducing the redundant calculations of the second computing module, thereby accelerating the model inference speed.
[0028] A model inference acceleration method provided by an embodiment of the present disclosure. The execution entity of model inference acceleration can be a Compute Express Link (CXL) device. The model inference acceleration method can be applied to artificial intelligence models with serialized serial output, including tasks such as image generation, video processing, speech generation, content parsing, and translation.
[0029] To enable those skilled in the art of the present technology to better understand the solution of this application, the following further details this application in conjunction with the accompanying drawings and specific embodiments.
[0030] Figure 1 It is a schematic flowchart of a model inference acceleration method provided by an embodiment of the present disclosure.
[0031] As Figure 1 shown, the method includes the following steps: Step 101, obtain intermediate variables of the serialized model in the first computing module. The intermediate variables include at least one of key-value copies, intermediate-layer latent features, and deep output features. The intermediate-layer latent features are features with a similarity not lower than a preset similarity threshold in the first computing module and the second computing module. The intermediate-layer latent features are determined by the shallow computing block in the first computing module, and the deep output features are determined by the deep computing block in the first computing module; In some embodiments, the first computing module refers to a computing unit or computing device that executes the inference process of the serialized model, which can be a Central Processing Unit (CPU), GPU, or other hardware accelerators.
[0032] In some embodiments, the serialized model refers to a model that processes input data sequentially, such as the DiT model, which gradually generates an output through a series of shallow and deep computing blocks.
[0033] In some embodiments, the key-value copy refers to a copy of the key and value matrices generated in the self-attention mechanism. The key-value copy can be represented by KV-cache, where K is the key and V is the value, and is used for reuse in subsequent calculation steps.
[0034] In some embodiments, the intermediate-layer latent features refer to the feature representations generated by the shallow computing block. These features are usually features that can be reused by the computing module to reduce the repeated calculations of the computing module.
[0035] In some embodiments, the deep output features refer to the final or near-final feature representations generated by the deep computing block, which are usually used for prediction results.
[0036] In some embodiments, the features with a similarity not lower than a preset similarity threshold refer to the features whose similarity degree between two computing blocks determined by a similarity measurement method reaches a preset standard, and are used to determine the features that can be reused between different computing modules. Among them, the similarity measurement method can be feature divergence, cosine similarity, etc.
[0037] In some embodiments, the shallow computing block is a computing unit located in the early stage of the model, responsible for preliminary feature extraction; the deep computing block is a computing unit located in the later stage of the model, responsible for more complex processing based on the early features and the generation of the final output.
[0038] In some embodiments, during the model inference process, intermediate variables are obtained from the first computing module, and these intermediate variables can be cached by the CXL device for subsequent use. Among them, the CXL device can be a CXL-Type2 device.
[0039] Step 102: Determine the features with a similarity not lower than the preset similarity threshold as the input of the deep computing block in the second computing module, so that the second computing module obtains the prediction result of the serialized model.
[0040] In some embodiments, the deep computing block in the second computing module refers to a computing unit located in a relatively deep layer of the model architecture, responsible for further processing based on the features provided by the shallow computing block and finally outputting the prediction result. Here, the shallow computing block can be the shallow computing block of the second computing module or the shallow computing block of the aforementioned first computing module. To reduce the computational redundancy of the model and accelerate the computational speed of the model, in this application, the deep computing block in the second computing module further processes based on the features provided by the shallow computing block of the aforementioned first computing module. In addition, the number of the second computing modules in this application is not limited, that is, the second computing module can have one, two or more, and the total number of the deep computing block and the shallow computing block in each second computing module (usually 10) is the same.
[0041] In some embodiments, by determining the features with a similarity not lower than the preset similarity threshold (i.e., the intermediate layer latent features in the intermediate variables) as the input of the deep computing block in the second computing module, some shallow computations are skipped, rather than recalculating all features from scratch. The existing features can be directly used for subsequent computations, saving a large amount of computational resources and time, thereby accelerating the entire inference process.
[0042] Through this application, it includes obtaining intermediate variables of the serialization model in the first computing module. The intermediate variables include at least one of key-value copies, intermediate-layer latent features, and deep output features. The intermediate-layer latent features are features with a similarity not lower than a preset similarity threshold in the first computing module and the second computing module. The intermediate-layer latent features are determined by the shallow computing blocks in the first computing module, and the deep output features are determined by the deep computing blocks in the first computing module. It is determined that the features with a similarity not lower than the preset similarity threshold are the inputs of the deep computing blocks in the second computing module, so that the second computing module obtains the prediction result of the serialization model, solving the technical problem that the cache of the second computing module expands rapidly in the related solutions, resulting in slow model inference speed, and achieving the technical effect of reducing the redundant calculations of the second computing module and thus accelerating the speed of model inference.
[0043] In some embodiments, obtaining the intermediate variables of the serialization model in the first computing module includes: Obtaining the key-value copies of two adjacent shallow computing blocks in the first computing module. The key-value copies are copies of the key-value matrices generated by the original images in the self-attention layer; In some embodiments, before obtaining the key-value copies of two adjacent shallow computing blocks in the first computing module, the key-value copies corresponding to each shallow computing block are cached in the CXL device instead of the reserved space of the first computing module.
[0044] In some embodiments, two adjacent shallow computing blocks refer to two shallow computing blocks that are next to each other in the network structure.
[0045] Based on the key-value copies of two adjacent shallow computing blocks, determine the intermediate variables of the serialization model.
[0046] In some embodiments, based on the key-value copies of two adjacent shallow computing blocks, according to the requirements of the application scenario, similarity measurement methods such as cosine similarity, Euclidean distance, or feature divergence can be used to compare the two key-value copies to determine the intermediate variables of the serialization model.
[0047] In some embodiments, by obtaining the key-value copies of two adjacent shallow computing blocks in the first computing module and based on the key-value copies of two adjacent shallow computing blocks, determining the intermediate variables of the serialization model not only reduces repeated calculations and improves inference efficiency, but also ensures that the model can flexibly adjust the cache policy according to the actual data features, so as to achieve performance optimization while maintaining high-quality output.
[0048] In some embodiments, due to the characteristics of the diffusion model itself, similar spatio-temporal features exist between adjacent network layers. The feature data can be distinguished according to this characteristic, so as to determine the model layer at which the spatio-temporal features start to change significantly. In this application, hierarchical attention divergence is used as the decision basis.
[0049] In some embodiments, determining the intermediate variable of the serialization model based on the key-value copies of two adjacent shallow computing blocks includes: Based on the key-value copies of two adjacent shallow computing blocks, determining the attention divergence corresponding to the two adjacent shallow computing blocks, where the attention divergence is used to determine the similarity of the spatio-temporal features output by the two adjacent shallow computing blocks; In some embodiments, the attention divergence can be defined by comparing the differences between two sets of key-value copies. Specifically, the cosine similarity can be used to measure the direction consistency of the two sets of key-value copy vectors by calculating the cosine similarity between them; or the Euclidean distance can be used to calculate the Euclidean distance between the two sets of key-value copies to evaluate their spatial differences.
[0050] In response to the attention divergence being not lower than a preset divergence threshold, determining the intermediate layer latent feature corresponding to the latter shallow computing block among the two adjacent shallow computing blocks as the intermediate variable of the serialization model.
[0051] In some embodiments, the preset divergence threshold refers to a preset value used to determine the intermediate variable. If the calculated attention divergence is not lower than the preset divergence threshold, it is considered that there are significant differences in the output features between the two shallow computing blocks. The former shallow computing block can be used as a reference, and the latter shallow computing block determines the parameter for the depth boundary.
[0052] In some embodiments, if there are significant differences in the output features between two shallow computing blocks, determining the intermediate layer latent feature corresponding to the latter shallow computing block as the intermediate variable of the serialization model means that these features can be directly utilized in subsequent inference processes instead of being recalculated.
[0053] In some embodiments, by determining the intermediate layer latent feature corresponding to the latter shallow computing block among two adjacent shallow computing blocks as the intermediate variable of the serialization model in response to the attention divergence being not lower than the preset divergence threshold, it is possible to dynamically select which features are suitable for caching and reuse, thereby optimizing the efficiency of the entire inference process.
[0054] In some embodiments, determining the attention divergence corresponding to two adjacent shallow computing blocks based on the key-value copies of the two adjacent shallow computing blocks includes: Based on the key-value copies of two adjacent shallow computing blocks, determining the feature map of each shallow computing block; In some embodiments, in the self-attention mechanism, the query, key, and value matrices calculate the attention scores through operations such as dot product, and then perform weighted summation to obtain the feature map of each shallow computing block.
[0055] Based on the feature map and a preset perturbation factor, determining the attention divergence corresponding to the two adjacent shallow computing blocks.
[0056] In some embodiments, based on the feature map and a preset perturbation factor, the mathematical expression for determining the attention divergence corresponding to two adjacent shallow computing blocks is as follows:
[0057] where i and j respectively represent the i-th shallow computing block and the j-th shallow computing block, i and j are adjacent, and typically j = i + 1. represents the i-th feature map and the j-th feature map. represents the attention divergence corresponding to the i-th feature map and the j-th feature map, c represents an element in the feature vector, a certain channel in the feature map; C represents the channel dimension, i.e., the number of channels. represents the feature map normalized by Softmax (the channel dimension is normalized to a probability distribution such that the values on each channel are in the range [0, 1], and the sum of all channel values is equal to 1). represents the preset perturbation factor, and the preset perturbation factor is a small constant such as is 10 to the power of negative 7, which is used to avoid the denominator in the logarithmic function being zero.
[0058] In some embodiments, after obtaining the intermediate variable of the serialized model in the first computing module, the model inference acceleration method further includes: Based on the intermediate variable, determining the deep output feature of the serialized model in the first computing module; In some embodiments, based on the intermediate variable, using the intermediate variable as the input of the deep computing block in the first computing module, the deep output feature of the serialized model in the first computing module is obtained.
[0059] Based on the deep output feature, determining the prediction result of the serialized model in the first computing module.
[0060] In some embodiments, the prediction result of the serialized model in the first computing module is the result obtained through full-scale calculation, and full-scale calculation refers to calculation through all shallow computing blocks and deep computing blocks.
[0061] In some embodiments, after determining the deep output feature of the serialized model in the first computing module based on the intermediate variable, the model inference acceleration method further includes: Determining the deep output feature as the input of the shallow computing block in the third computing module, so that the third computing module obtains the prediction result of the serialized model.
[0062] In some embodiments, when the next time step of the diffusion model does not require cache reuse, the deep output feature needs to be directly used as the input of the next shallow computing block in the next time step.
[0063] In some embodiments, the third computing module refers to a computing module other than the first computing module and the second computing module, and the full calculation is performed in the third computing module. This application does not limit the number of third computing modules, and the number of third computing modules can be determined according to user needs. If the speed of model reasoning is taken into account, the fewer the number of third computing modules, the better. If the accuracy of model reasoning is taken into account, the more the number of third computing modules, the better.
[0064] In some embodiments, the model reasoning acceleration method further includes: Obtain the high-frequency ratio of the serialized model in the second computing module; In some embodiments, according to the difference characteristics of information generation at different time steps of the diffusion model, generally speaking, the image information generated at the early moment is mainly concentrated in the low frequency, and the high-frequency information begins to be gradually enriched in the subsequent steps. Therefore, according to this characteristic, the cache can be reused as much as possible in the early stage, and the full amount of computing blocks can be used for calculation when key information is generated in the later stage.
[0065] In some embodiments, the high-frequency ratio is used to determine at which diffusion step to reuse the intermediate layer potential features. High-frequency information represents the details, edges, textures and other parts of the image / signal that change dramatically. In a neural network, the more high-frequency information, the more details the currently processed data or features contain, which are not suitable for direct reuse.
[0066] In response to the high-frequency ratio of the serialized model being not less than a preset high-frequency ratio, a target time step in the serialized model is determined. The target time step is used to distinguish the time step of intermediate layer potential feature reuse and the time step of full calculation. The full calculation includes shallow calculation block calculation and deep calculation block calculation in the second calculation module.
[0067] In some embodiments, the preset high-frequency ratio is used to determine whether to enter the full calculation process, and the preset high-frequency ratio can be determined based on historical experience.
[0068] In some embodiments, if the high-frequency ratio of the serialized model is not lower than the preset high-frequency ratio, it means that it currently includes more detailed information and is no longer suitable for feature reuse; if the high-frequency ratio of the serialized model is lower than the preset high-frequency ratio, it means that the intermediate layer potential feature reuse strategy can be adopted to reduce redundant calculations.
[0069] In some embodiments, by determining the target time step in the serialization model in response to the high-frequency ratio of the serialization model being not less than a preset high-frequency ratio, in the diffusion model, the intermediate layer potential features can be reused in the early time steps (with larger noise); while in the later time steps, i.e., the image detail recovery stage, the high-frequency information increases and the full amount is calculated to ensure the generation quality.
[0070] In some embodiments, obtaining the high-frequency ratio of the serialized model in the second computing module includes: Obtain an attention weight matrix and input features, where the attention weight matrix is used to indicate the degree of spatial attention to the input features; In some embodiments, the input features refer to the original image received by the model or the feature representation after preliminary transformation.
[0071] Based on the attention weight matrix, input features, and a preset high-frequency cut-off frequency threshold, determine the high-frequency ratio of the serialized model in the second computing module.
[0072] In some embodiments, the preset high-frequency cut-off frequency threshold is used to distinguish high-frequency components from low-frequency components.
[0073] In some embodiments, the mathematical expression for determining the high-frequency ratio of the serialized model in the second computing module is as follows:
[0074] where fx and fy are the horizontal and vertical components in the frequency domain coordinates; F is the feature after Fourier transform; x, y, and c represent dimensions of H×W×C (height, width, number of channels); A is the attention weight matrix, with dimensions of H×W, indicating the degree of spatial attention to the input features, represents the preset high-frequency cut-off frequency threshold.
[0075] In some embodiments, finally, set the high-frequency ratio through a threshold to determine that reuse will no longer be performed at a later time step. Specifically, in the later stage of the diffusion model (when t gradually decreases to 1), the high-frequency ratio will be very high. The later time step can be determined by historical experience. When t approaches 1, usually full-scale calculation is performed, that is, both the shallow calculation block and the deep calculation block participate in the calculation process.
[0076] Through the present application, including obtaining intermediate variables of the serialized model in the first computing module, where the intermediate variables include at least one of key-value copies, intermediate-layer latent features, and deep output features. The intermediate-layer latent features are features with a similarity not lower than a preset similarity threshold in the first computing module and the second computing module. The intermediate-layer latent features are determined by the shallow calculation block in the first computing module, and the deep output features are determined by the deep calculation block in the first computing module; determining the features with a similarity not lower than the preset similarity threshold as the input of the deep calculation block in the second computing module, so that the second computing module obtains the prediction result of the serialized model, solves the technical problem that the cache of the second computing module expands rapidly in the related solution, resulting in slow model inference speed, and achieves the technical effect of reducing redundant calculations of the second computing module and thus accelerating the model inference speed.
[0077] Embodiments of the present application also provide a model inference acceleration system, which includes: a first computing module, a cache module, and at least one second computing module. The cache module is communicatively connected to the first computing module and the second computing module respectively; The first computing module is used to serialize the inference calculation of the model to obtain intermediate variables, which include at least one of key-value copies, intermediate-layer latent features, and deep output features; The cache module is used to cache the intermediate variables obtained by the inference calculation of the first computing module; The second computing module is used to perform the inference calculation of the serialized model according to the intermediate variables in the cache module.
[0078] In some embodiments, the first computing module includes at least one shallow computing block and at least one deep computing block; The shallow computing block is used to determine the intermediate-layer latent features; The deep computing block is used to determine the deep output features based on the intermediate-layer latent features.
[0079] In some embodiments, as Figure 2 shown, Figure 2 is a schematic structural diagram of a model inference acceleration system provided by an embodiment of the present application. In the inference stage of the diffusion model, it is in reverse order of time t. The first computing module is Figure 2 the computing device corresponding to the feature caching step (step t), and the second computing module is Figure 2 the computing device corresponding to the feature reuse step (step t-1). Both the first computing module and the second computing module include a shallow computing block (shallow block) and a deep computing block (deep block). Noised Image is the input image, and Predicted Noiset and Predicted Noiset-1 are the predicted outputs corresponding to steps t and t-1 in the diffusion model. Among them, ① The model input performs forward calculations of multiple computing blocks (blocks) in multiple computing devices. Here, the block is determined according to the specific architecture in the model algorithm, generally about dozens. ② The key-value copies (K, V matrices) calculated by each shallow computing block in the first computing module are transferred to the cache module, that is Figure 2In the CXL - type2 device, ③ according to the calculation of the feature divergence of the key - value copies of two adjacent shallow - layer calculation blocks, it is determined whether to cache the intermediate - layer latent features at this shallow - layer calculation block. ④ Continue the forward calculation of the remaining blocks and output the deep - layer output features. ⑤ According to the deep - layer output features, output the predicted output (the predicted noise) of this diffusion step. ⑥ After determining the shallow - deep boundary parameter, directly use the intermediate - layer latent features cached in step ③ as the input of the deep - layer block to continue the calculation, thus skipping the calculation process of the shallow - layer block. ⑦ Reuse the key - value copy in step ② for this deep - layer block. ⑧ Continue to calculate the deep - layer block until the output of this step is output.
[0080] In some embodiments, the cache module includes a decision - control module, and the decision - control module at least includes a cache - feature divergence calculation module and a high - low frequency calculation module for deep - layer output features; The cache - feature divergence calculation module is used to determine a first parameter, and the first parameter is used to indicate the boundary value between the deep - layer calculation block and the shallow - layer calculation block; The high - low frequency calculation module for deep - layer output features is used to determine a second parameter, and the second parameter is used to indicate the time step of feature reuse in the serialization model.
[0081] In some embodiments, the decision - control module is the aforementioned Figure 2 decision controller. Further, as Figure 3 shown, Figure 3 is a schematic structural diagram of a decision - control module provided by an embodiment of the present application. The cache - feature divergence calculation module determines the first parameter according to the calculated value and the threshold. The first parameter, that is, the shallow - deep boundary parameter, is used to determine which layer of intermediate latent features to cache. If there is no need to cache, direct full - volume calculation is performed; the high - low frequency calculation module for deep - layer output features is used to determine at which diffusion step to reuse the intermediate - layer latent features.
[0082] In some embodiments, the system further includes a cache - reuse determination module; The cache - reuse determination module is used to determine whether to cache the intermediate - layer latent features and whether to reuse the intermediate - layer latent features according to the value at the target bit. The target bit is a reserved bit in the original data - frame format.
[0083] In some embodiments, taking the standard CXL flit data packet as an example, based on the H3 - S2M DRS header + S2MNDR header information, the reserved bits are the 7th and 6th bits of the 12th byte in the original data - frame format.
[0084] In some embodiments, as Figure 4 shown, Figure 4A structural schematic diagram of a data frame format provided by an embodiment of the present application. In the original data frame format, the 7th and 6th bits of the 12th byte are used as flag bits for whether to perform intermediate layer potential feature caching and whether to perform multiplexing; in the 13th byte, the first parameter (the depth boundary parameter i, where i is usually less than 256) is stored and passed to the CXL device. This data packet will be read by the CXL device and control the data flow direction; in the 14th byte, the cumulative value of the boundary parameter is stored. The role of the cumulative value is to control the total number of blocks that need to be cached as a whole, so as to be able to estimate the total number of skipped blocks as a whole, and be used to control extreme situations; in the 15th byte, the cumulative value of the multiplexing step is stored. This cumulative value is used to control the number of steps of the multiplexing cache to avoid excessive multiplexing times, thereby losing important high-frequency information. Specifically, if the multiplexing times are excessive, the full-scale calculation of the later-stage blocks will be less, and the later stage is the moment to reflect high-frequency information.
[0085] In some embodiments, by customizing the CXL protocol packet transmission design, high-speed communication is ensured for the CXL device and the computing device to transmit key control information.
[0086] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.
[0087] An embodiment of the present application also provides a model inference acceleration device 500. Figure 5 A structural schematic diagram of a model inference acceleration device provided by an embodiment of the present disclosure, as Figure 5 shown, including: An acquisition unit 501, configured to acquire intermediate variables of the serialized model in the first computing module, where the intermediate variables include at least one of a key-value copy, an intermediate layer potential feature, and a deep output feature. The intermediate layer potential feature is a feature with a similarity not lower than a preset similarity threshold in the first computing module and the second computing module. The intermediate layer potential feature is determined by a shallow computing block in the first computing module, and the deep output feature is determined by a deep computing block in the first computing module; A determination unit 502, configured to determine that a feature with a similarity not lower than a preset similarity threshold is an input to a deep computing block in the second computing module, so that the second computing module obtains a prediction result of the serialized model.
[0088] Further, in a possible implementation manner of an embodiment of the present disclosure, the acquisition unit 501 is configured to: Acquire key-value copies of two adjacent shallow computing blocks in the first computing module. The key-value copy is a copy of the key-value matrix generated by the original image in the self-attention layer; Determine the intermediate variable of the serialization model based on the key-value copies of two adjacent shallow computing blocks.
[0089] Further, in a possible implementation manner of the embodiments of the present disclosure, the obtaining unit 501 is configured to: Based on the key-value copies of two adjacent shallow computing blocks, determine the attention divergence corresponding to the two adjacent shallow computing blocks, where the attention divergence is used to determine the similarity of the spatio-temporal features output by the two adjacent shallow computing blocks; In response to the attention divergence being not lower than a preset divergence threshold, determine the intermediate-layer latent feature corresponding to the latter shallow computing block among the two adjacent shallow computing blocks as the intermediate variable of the serialization model.
[0090] Further, in a possible implementation manner of the embodiments of the present disclosure, the obtaining unit 501 is configured to: Based on the key-value copies of two adjacent shallow computing blocks, determine the feature map of each shallow computing block; Based on the feature map and a preset perturbation factor, determine the attention divergence corresponding to the two adjacent shallow computing blocks.
[0091] Further, in a possible implementation manner of the embodiments of the present disclosure, the model inference acceleration device 500 further includes a prediction result determination unit, and the prediction result determination unit is configured to: Based on the intermediate variable, determine the deep output feature of the serialization model in the first computing module; Based on the deep output feature, determine the prediction result of the serialization model in the first computing module.
[0092] Further, in a possible implementation manner of the embodiments of the present disclosure, the prediction result determination unit is further configured to: Determine the deep output feature as the input of the shallow computing block in the third computing module, so that the third computing module obtains the prediction result of the serialization model.
[0093] Further, in a possible implementation manner of the embodiments of the present disclosure, the model inference acceleration device 500 further includes a time step determination unit, and the time step determination unit is configured to: Obtain the high-frequency ratio of the serialization model in the second computing module; In response to the high-frequency ratio of the serialization model being not lower than a preset high-frequency ratio, determine the target time step in the serialization model, where the target time step is used to distinguish the time step of reusing the intermediate-layer latent feature and the time step of full-scale calculation, and the full-scale calculation includes the calculation of the shallow computing block and the deep computing block in the second computing module.
[0094] Further, in a possible implementation manner of the embodiments of the present disclosure, the time step determination unit is further configured to: Obtain an attention weight matrix and input features, where the attention weight matrix is used to indicate the degree of spatial attention to the input features; Based on the attention weight matrix, the input features, and a preset high-frequency cut-off frequency threshold, determine the high-frequency proportion of the serialization model in the second calculation module.
[0095] Through this application, it includes obtaining intermediate variables of the serialization model in the first calculation module, where the intermediate variables include at least one of key-value copies, intermediate-layer latent features, and deep output features. The intermediate-layer latent features are features with a similarity not lower than a preset similarity threshold in the first calculation module and the second calculation module. The intermediate-layer latent features are determined by the shallow calculation block in the first calculation module, and the deep output features are determined by the deep calculation block in the first calculation module; determine that the features with a similarity not lower than the preset similarity threshold are the input of the deep calculation block in the second calculation module, so that the second calculation module obtains the prediction result of the serialization model, solving the technical problem that the cache of the second calculation module expands rapidly in the related solutions, resulting in slow model inference speed, and achieving the technical effect of reducing the redundant calculation of the second calculation module and thus accelerating the model inference speed.
[0096] For the description of the features in the corresponding embodiments of the model inference acceleration device, reference can be made to the relevant descriptions in the corresponding embodiments of the model inference acceleration method, which will not be elaborated here one by one.
[0097] An embodiment of this application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-mentioned model inference acceleration method embodiments.
[0098] An embodiment of this application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above-mentioned model inference acceleration method embodiments when running.
[0099] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs, and other media that can store computer programs.
[0100] An embodiment of this application also provides a computer program product. The above-mentioned computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above-mentioned model inference acceleration method embodiments.
[0101] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps in any of the above-described embodiments of the model inference acceleration method.
[0102] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0103] The above has introduced in detail a model inference acceleration method, system, electronic device, storage medium, and product provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for accelerating model inference, characterized in that, Including: Obtain intermediate variables of the serialized model in the first computing module, where the intermediate variables include at least one of a key-value copy, intermediate-layer latent features, and deep output features. The intermediate-layer latent features are features with a similarity not lower than a preset similarity threshold in the first computing module and the second computing module. The intermediate-layer latent features are determined by a shallow computing block in the first computing module, and the deep output features are determined by a deep computing block in the first computing module. Determine that the features with a similarity not lower than the preset similarity threshold are inputs to the deep computing block in the second computing module, so that the second computing module obtains the prediction result of the serialized model.
2. The model inference acceleration method according to claim 1, wherein The obtaining of the intermediate variables of the serialized model in the first computing module includes: Obtain the key-value copies of two adjacent shallow computing blocks in the first computing module. The key-value copies are copies of the key-value matrix generated by the original image in the self-attention layer. Based on the key-value copies of the two adjacent shallow computing blocks, determine the intermediate variables of the serialized model.
3. The model inference acceleration method according to claim 2, wherein The determining of the intermediate variables of the serialized model based on the key-value copies of the two adjacent shallow computing blocks includes: Based on the key-value copies of the two adjacent shallow computing blocks, determine the attention divergence corresponding to the two adjacent shallow computing blocks. The attention divergence is used to determine the similarity of the spatio-temporal features output by the two adjacent shallow computing blocks. In response to the attention divergence not being lower than a preset divergence threshold, determine that the intermediate-layer latent features corresponding to the latter shallow computing block among the two adjacent shallow computing blocks are the intermediate variables of the serialized model.
4. The model inference acceleration method according to claim 3, wherein The determining of the attention divergence corresponding to the two adjacent shallow computing blocks based on the key-value copies of the two adjacent shallow computing blocks includes: Based on the key-value copies of the two adjacent shallow computing blocks, determine the feature maps of each shallow computing block. Based on the feature maps and a preset perturbation factor, determine the attention divergence corresponding to the two adjacent shallow computing blocks.
5. The model inference acceleration method according to claim 1, wherein After obtaining the intermediate variables of the serialized model in the first computing module, the method further includes: Based on the intermediate variables, determine the deep output features of the serialized model in the first computing module. Based on the deep output features, determine the prediction result of the serialized model in the first computing module.
6. The model inference acceleration method according to claim 5, wherein After determining the deep output features of the serialized model in the first computing module based on the intermediate variables, the method further includes: Determine that the deep output features are inputs to the shallow computing block in the third computing module, so that the third computing module obtains the prediction result of the serialized model.
7. The model inference acceleration method according to claim 1, wherein The method further includes: Obtain the high-frequency ratio of the serialized model in the second computing module. In response to the high-frequency ratio of the serialized model not being lower than a preset high-frequency ratio, determine the target time step in the serialized model. The target time step is used to distinguish the time step for reusing intermediate-layer latent features and the time step for full-scale calculation. The full-scale calculation includes the calculation of the shallow computing block and the deep computing block in the second computing module.
8. The model inference acceleration method according to claim 7, wherein The obtaining of the high-frequency ratio of the serialized model in the second computing module includes: Obtain an attention weight matrix and input features, where the attention weight matrix is used to indicate the degree of spatial attention to the input features; Based on the attention weight matrix, the input features, and a preset high-frequency cut-off frequency threshold, determine the high-frequency proportion of the serialization model in the second calculation module.
9. A model inference acceleration system, characterized in that, The system includes: a first calculation module, a cache module, and at least one second calculation module, and the cache module is communicatively connected to the first calculation module and the second calculation module respectively; The first calculation module is used for the inference calculation of the serialization model to obtain intermediate variables, where the intermediate variables include at least one of key-value copies, intermediate-layer latent features, and deep output features; The cache module is used to cache the intermediate variables obtained by the inference calculation of the first calculation module; The second calculation module is used for the inference calculation of the serialization model according to the intermediate variables in the cache module.
10. The model inference acceleration system according to claim 9, wherein The first calculation module includes at least one shallow calculation block and at least one deep calculation block; The shallow calculation block is used to determine intermediate-layer latent features; The deep calculation block is used to determine deep output features based on the intermediate-layer latent features.
11. The model inference acceleration system according to claim 9, wherein The cache module includes a decision control module, and the decision control module at least includes a cache feature divergence calculation module and a high-low frequency calculation module for deep output features; The cache feature divergence calculation module is used to determine a first parameter, and the first parameter is used to indicate the demarcation value for the deep calculation block and the shallow calculation block; The high-low frequency calculation module for deep output features is used to determine a second parameter, and the second parameter is used to indicate the time step of feature reuse in the serialization model.
12. The model inference acceleration system according to claim 9, wherein The system further includes a cache reuse determination module; The cache reuse determination module is used to determine whether to cache the intermediate-layer latent features and whether to reuse the intermediate-layer latent features according to the value at the target bit, where the target bit is a reserved bit in the original data frame format.
13. An electronic device, characterized in that, Comprising: A memory for storing a computer program; A processor for implementing the steps of the model inference acceleration method according to any one of claims 1 to 8 when executing the computer program.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the model inference acceleration method according to any one of claims 1 to 8 when executed by a processor.
15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the model inference acceleration method according to any one of claims 1 to 8 when executed by a processor.
Citation Information
Patent Citations
Image generation method and device, equipment and medium
CN116701692A
Video description generation method and device based on deep learning model
CN117292293A
Model double-level decoding method and device
CN118114655A
Model reasoning method, electronic equipment, storage medium and computer program product
CN118350470A
Transform-based method and device for constructing sample efficient world model
CN119271974A
Cited By
Diffusion Transform reasoning acceleration method and device based on distributed perception sparsity
CN121787559A