Model reasoning acceleration method, system, electronic device, storage medium and product

By obtaining and using the intermediate variables of the first computing module during the model inference process, especially the feature with a similarity not lower than the threshold as the input of the second computing module, the problem of rapid cache expansion is solved, and the model inference speed is improved.

CN120258152BActive Publication Date: 2025-08-12INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510724702.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-12
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

In the prior art, the rapid expansion of cache during model inference process leads to excessive resource use of computing equipment, affecting the speed of model inference.

Method used

By obtaining the intermediate variables of the serialized model in the first computing module, including key-value copies, intermediate layer potential features and deep output features, it is determined that the feature whose similarity is not lower than the preset similarity threshold is the input of the second computing module, and redundant calculations are reduced.

Benefits of technology

The redundant calculation of the second computing module is reduced, thereby speeding up the model inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258152B_ABST
    Figure CN120258152B_ABST
Patent Text Reader

Abstract

The present application discloses a model reasoning acceleration method, system, electronic device, storage medium and product, which relate to the field of artificial intelligence technology, including obtaining intermediate variables of a serialized model in a first computing module, the intermediate variables including at least one of a key-value copy, an intermediate-layer potential feature and a deep-layer output feature, the intermediate-layer potential feature being a feature whose similarity in the first computing module and the second computing module is not less than a preset similarity threshold, the intermediate-layer potential feature being determined by a shallow computing block in the first computing module, and the deep-layer output feature being determined by a deep computing block in the first computing module; determining a feature whose similarity is not less than a preset similarity threshold as an input to the deep computing block in the second computing module, so that the second computing module obtains a prediction result of the serialized model, solving the technical problem of rapid cache expansion in related solutions, which leads to slow model reasoning speed, and achieving the technical effect of reducing redundant calculations and thereby speeding up the model reasoning speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to model reasoning acceleration methods, systems, electronic devices, storage media, and products. Background Art

[0002] In related model inference acceleration solutions, sufficient space is usually reserved in the memory of the computing device to cache data during the model inference calculation process. During the inference calculation process, the cache of the computing device expands rapidly, resulting in slow model inference speed. Summary of the Invention

[0003] This application provides a model reasoning acceleration method, system, electronic device, storage medium and product to at least solve the problem in related technologies of rapid expansion of computing device cache, resulting in slow model reasoning speed.

[0004] This application provides a model reasoning acceleration method, including:

[0005] Obtaining intermediate variables of the serialized model in the first computing module, the intermediate variables including at least one of a key-value copy, an intermediate-layer potential feature, and a deep-layer output feature, wherein the intermediate-layer potential feature is a feature in the first computing module and the second computing module whose similarity is not less than a preset similarity threshold, the intermediate-layer potential feature is determined by a shallow computing block in the first computing module, and the deep-layer output feature is determined by a deep computing block in the first computing module;

[0006] Determine that the features with similarity not lower than a preset similarity threshold are input to the deep computing block in the second computing module, so that the second computing module obtains the prediction result of the serialization model.

[0007] The present application provides a model reasoning acceleration system, the system comprising: a first computing module, a cache module and at least one second computing module, the cache module being communicatively connected to the first computing module and the second computing module respectively;

[0008] The first computing module is used for inference calculation of the serialized model to obtain intermediate variables, which include at least one of the key-value copy, the intermediate layer potential feature, and the deep output feature;

[0009] The cache module is used to cache the intermediate variables obtained by the inference calculation of the first calculation module;

[0010] The second computing module is used to perform inference calculations of the serialized model based on the intermediate variables in the cache module.

[0011] This application also provides a model reasoning acceleration device, including:

[0012] an acquisition unit, configured to acquire intermediate variables of the serialized model in the first computing module, the intermediate variables including at least one of a key-value copy, an intermediate-layer potential feature, and a deep-layer output feature, the intermediate-layer potential feature being a feature in the first computing module and the second computing module whose similarity is not less than a preset similarity threshold, the intermediate-layer potential feature being determined by a shallow computing block in the first computing module, and the deep-layer output feature being determined by a deep computing block in the first computing module;

[0013] The determining unit is used to determine the features whose similarity is not less than a preset similarity threshold as the input of the deep computing block in the second computing module, so that the second computing module obtains the prediction result of the serialization model.

[0014] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned model inference acceleration methods when executing the computer program.

[0015] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned XX methods are implemented.

[0016] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned model reasoning acceleration methods when executed by a processor.

[0017] Through this application, it includes obtaining the intermediate variables of the serialized model in the first computing module, the intermediate variables including at least one of a key-value copy, an intermediate-layer potential feature and a deep-layer output feature, the intermediate-layer potential feature is a feature in the first computing module and the second computing module whose similarity is not less than a preset similarity threshold, the intermediate-layer potential feature is determined by a shallow computing block in the first computing module, and the deep-layer output feature is determined by a deep computing block in the first computing module; the feature whose similarity is not less than a preset similarity threshold is determined as the input of the deep computing block in the second computing module, so that the second computing module obtains the prediction result of the serialized model, solves the technical problem of rapid expansion of the cache of the second computing module in the related scheme, resulting in slow model inference speed, and achieves the technical effect of reducing redundant calculations of the second computing module and thereby speeding up the speed of model inference. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1A flowchart of a model reasoning acceleration method provided in an embodiment of the present application;

[0020] Figure 2 A schematic diagram of the structure of a model reasoning acceleration system provided in an embodiment of the present application;

[0021] Figure 3 A schematic diagram of the structure of a decision control module provided in an embodiment of the present application;

[0022] Figure 4 A schematic diagram of the structure of a data frame format provided in an embodiment of the present application;

[0023] Figure 5 A schematic diagram of the structure of a model reasoning acceleration device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0026] In order to facilitate those skilled in the art to better understand the technical solutions described in the embodiments of the present disclosure, the technical terms in the embodiments of the present disclosure are explained as follows before introducing the embodiments of the present disclosure.

[0027] Graphics Processing Unit (GPU): A processor used to quickly process image and video data.

[0028] CXL: An open industry standard for high-bandwidth, low-latency device interconnection, it allows fast, reliable data transfer between different components within a computer system. It aims to address bottlenecks in high-performance computing, including memory capacity, memory bandwidth, and input / output latency. CXL also enables memory expansion and sharing, and can communicate with peripherals such as computing accelerators like GPUs and field-programmable gate arrays (FPGAs), providing faster and more flexible data exchange and processing.

[0029] Type 2 devices: Type 2 devices support both CXL.mem and CXL.cache protocols.

[0030] Device memory: refers to the memory visible to the device, such as the High Bandwidth Memory (HBM) carried by the GPU.

[0031] Denoising Diffusion Probabilistic Models (DDPMs): These are generative models based on a diffusion process. They gradually add Gaussian noise to data (such as an image) and then train a neural network to gradually remove the noise, thereby recovering the original data from pure random noise. They have good generation quality but are generally slow to infer because they require multiple denoising iterations.

[0032] Denoising Diffusion Implicit Models (DDIM): DDIM is an improved version of DDPM that introduces a non-Markov denoising process, which allows image generation to be completed with fewer steps during the inference phase, significantly speeding up inference.

[0033] Spatial-Temporal Similarity Selection (STSS): This approach aims to improve cache efficiency and accelerate model inference by selecting and reusing features that are highly similar in both space and time. However, for input scenarios of varying complexity, a fixed cache reuse strategy can lead to irrational resource allocation, impacting overall model performance.

[0034] Extended Memory: This provides high-bandwidth, low-latency connectivity, enabling more efficient communication between the host processor and devices such as accelerators, memory buffers, and intelligent input / output devices. CXL extended memory addresses the limitations of traditional memory expansion methods by providing high-bandwidth, low-latency connectivity and efficient memory management.

[0035] Diffusion Transformer (DiT): is a generative model that combines the diffusion model and transformer architecture, mainly used in high-resolution image and video generation tasks.

[0036] The core idea of DiT is to replace the U-Net structure used in traditional diffusion models (such as DDPM and DDIM) with the transformer's self-attention mechanism to improve generation quality and diversity. However, the DiT model faces the following issues and challenges. First, inefficient cache management leads to resource waste: Existing cache reuse techniques (such as STSS feature cache) rapidly expand cache data during inference, occupying a large amount of video memory resources and limiting the model's ability to process high-resolution images in real time. Second, dynamic policy adaptation for specific instances is difficult: Cache reuse policies are difficult to adapt to different input instances (e.g., complex scenes vs. simple objects), resulting in irrational allocation of computing resources in some scenarios (e.g., excessive reuse of low-frequency features in simple scenes and omission of high-frequency features in complex scenes). Experiments show that under a fixed policy, the computational redundancy rate for simple scenes is as high as 40%, while the omission rate of key features for complex scenes exceeds 15%. Finally, the hardware communication protocol is not compatible with cache expansion: The data frame format of traditional communication protocols is not optimized for DiT cache expansion requirements, resulting in increased communication latency for high-frequency cache updates, which limits the improvement of end-to-end inference speed. Existing protocols face bandwidth competition when multiple devices collaborate (such as computing devices and cache devices), making it difficult to meet the requirements for real-time feature synchronization in the later stages of DiT denoising.

[0037] In the related model inference acceleration scheme, DiT is accelerated by exploring the similarity of features between adjacent time steps. Specifically, features with the most similar structures are prioritized for identification, and these highly similar features are cached and reused to reduce redundant computations, thereby accelerating DiT while maximizing consistency with the original model generation results. At the same time, a lightweight decision network for specific instances is designed that can dynamically allocate resources and provide better content quality. However, in the related model inference scheme, there is no independent decision control, and the cache reuse strategy is fixed. During the calculation process, the cache expands rapidly, resulting in slow model inference speed.

[0038] Through this application, it includes obtaining the intermediate variables of the serialized model in the first computing module, the intermediate variables including at least one of a key-value copy, an intermediate-layer potential feature and a deep-layer output feature, the intermediate-layer potential feature is a feature in the first computing module and the second computing module whose similarity is not less than a preset similarity threshold, the intermediate-layer potential feature is determined by a shallow computing block in the first computing module, and the deep-layer output feature is determined by a deep computing block in the first computing module; the feature whose similarity is not less than a preset similarity threshold is determined as the input of the deep computing block in the second computing module, so that the second computing module obtains the prediction result of the serialized model, solves the technical problem of rapid expansion of the cache of the second computing module in the related scheme, resulting in slow model inference speed, and achieves the technical effect of reducing redundant calculations of the second computing module and thereby speeding up the speed of model inference.

[0039] The embodiments of the present disclosure provide a model inference acceleration method, in which the execution subject of the model inference acceleration can be a Compute Express Link (CXL) device. The model inference acceleration method can be applied to artificial intelligence models with serialized serial output, including tasks such as image generation, video processing, speech generation, content analysis, and translation.

[0040] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0041] Figure 1 A flowchart of a model reasoning acceleration method provided in an embodiment of the present disclosure.

[0042] like Figure 1 As shown, the method comprises the following steps:

[0043] Step 101: Obtain intermediate variables of the serialized model in the first computing module. The intermediate variables include at least one of a key-value copy, an intermediate-layer potential feature, and a deep-layer output feature. The intermediate-layer potential feature is a feature whose similarity between the first computing module and the second computing module is not less than a preset similarity threshold. The intermediate-layer potential feature is determined by a shallow computing block in the first computing module, and the deep-layer output feature is determined by a deep computing block in the first computing module.

[0044] In some embodiments, the first computing module refers to a computing unit or computing device that performs the serialized model inference process, which may be a central processing unit (CPU), a GPU, or other hardware accelerator.

[0045] In some embodiments, a sequential model refers to a model that processes input data sequentially, such as a DiT model, and gradually generates output through a series of shallow and deep computation blocks.

[0046] In some embodiments, the key-value copy refers to a copy of the key and value matrix generated in the self-attention mechanism. The key-value copy can be represented by a KV-cache, where K is the key and V is the value, for reuse in subsequent calculation steps.

[0047] In some embodiments, the intermediate layer potential features refer to feature representations generated by shallow layer computing blocks. These features are generally features that can be reused by computing modules to reduce repeated computations of the computing modules.

[0048] In some embodiments, deep layer output features refer to the final or near-final feature representations generated by the deep layer computation blocks, which are typically used to predict results.

[0049] In some embodiments, a feature whose similarity is not less than a preset similarity threshold refers to a feature whose feature similarity between two computing blocks determined by a similarity measurement method reaches a preset standard, and is used to determine features that can be reused between different computing modules, wherein the similarity measurement method can be feature divergence, cosine similarity, etc.

[0050] In some embodiments, the shallow computing block is a computing unit located in the early stage of the model, responsible for preliminary feature extraction; the deep computing block is a computing unit located in the later stage of the model, responsible for more complex processing based on early features and generating final output.

[0051] In some embodiments, during the model inference process, intermediate variables are obtained from the first computing module. These intermediate variables may be cached by a CXL device for subsequent use, wherein the CXL device may be a CXL-Type 2 device.

[0052] Step 102 : Determine the features whose similarity is not less than a preset similarity threshold as inputs to the deep computing block in the second computing module, so that the second computing module obtains the prediction results of the serialization model.

[0053] In some embodiments, the deep computing block in the second computing module refers to a computing unit located at a deeper level of the model architecture, which is responsible for further processing based on the features provided by the shallow computing block and finally outputting the prediction results. The shallow computing block here can be a shallow computing block of the second computing module or a shallow computing block of the aforementioned first computing module. In order to reduce the computational redundancy of the model and speed up the calculation speed of the model, the deep computing block in the second computing module in this application is further processed based on the features provided by the shallow computing block of the aforementioned first computing module. In addition, this application does not limit the number of second computing modules, that is, the second computing module can have one, two or more, and the total number of deep computing blocks and shallow computing blocks in each second computing module (usually 10) is the same.

[0054] In some embodiments, by determining features whose similarity is not less than a preset similarity threshold (i.e., intermediate layer potential features in the intermediate variables) as the input of the deep calculation block in the second calculation module, some shallow calculations are skipped, instead of recalculating all features from scratch. Existing features can be directly used for subsequent calculations, saving a lot of computing resources and time, thereby speeding up the entire reasoning process.

[0055] Through this application, it includes obtaining the intermediate variables of the serialized model in the first computing module, the intermediate variables including at least one of a key-value copy, an intermediate-layer potential feature and a deep-layer output feature, the intermediate-layer potential feature is a feature in the first computing module and the second computing module whose similarity is not less than a preset similarity threshold, the intermediate-layer potential feature is determined by a shallow computing block in the first computing module, and the deep-layer output feature is determined by a deep computing block in the first computing module; the feature whose similarity is not less than a preset similarity threshold is determined as the input of the deep computing block in the second computing module, so that the second computing module obtains the prediction result of the serialized model, solves the technical problem of rapid expansion of the cache of the second computing module in the related scheme, resulting in slow model inference speed, and achieves the technical effect of reducing redundant calculations of the second computing module and thereby speeding up the speed of model inference.

[0056] In some embodiments, obtaining the intermediate variables of the serialized model in the first computing module includes:

[0057] Obtain key-value copies of two adjacent shallow computation blocks in the first computation module. The key-value copies are copies of the key-value matrix generated by the original image in the self-attention layer.

[0058] In some embodiments, before obtaining key-value copies of two adjacent shallow computation blocks in the first computation module, the key-value copies corresponding to each shallow computation block are cached in the CXL device instead of the reserved space of the first computation module.

[0059] In some embodiments, two adjacent shallow computing blocks refer to two shallow computing blocks that are immediately adjacent to each other in the network structure.

[0060] Based on the key-value copies of two adjacent shallow calculation blocks, the intermediate variables of the serialization model are determined.

[0061] In some embodiments, based on the key-value copies of two adjacent shallow calculation blocks, similarity measurement methods such as cosine similarity, Euclidean distance or feature divergence can be used to compare the two key-value copies according to the requirements of the application scenario to determine the intermediate variables of the serialization model.

[0062] In some embodiments, by obtaining key-value copies of two adjacent shallow computing blocks in the first computing module, and determining the intermediate variables of the serialized model based on the key-value copies of the two adjacent shallow computing blocks, not only repeated calculations are reduced and reasoning efficiency is improved, but also it is ensured that the model can flexibly adjust the cache strategy according to actual data characteristics, thereby achieving performance optimization while maintaining high-quality output.

[0063] In some embodiments, due to the characteristics of the diffusion model itself, similar spatiotemporal features exist between adjacent network layers. Feature data can be distinguished based on this characteristic to determine the model layer number where the spatiotemporal features begin to change significantly. In this application, hierarchical attention divergence is used as the decision basis.

[0064] In some embodiments, determining the intermediate variables of the serialization model based on the key-value copies of two adjacent shallow computation blocks includes:

[0065] Based on the key-value copies of two adjacent shallow computing blocks, the attention divergence corresponding to the two adjacent shallow computing blocks is determined. The attention divergence is used to determine the similarity of the spatiotemporal features output by the two adjacent shallow computing blocks.

[0066] In some embodiments, attention divergence can be defined by comparing the differences between two sets of key-value copies. Specifically, cosine similarity can be used to measure the directional consistency of the two sets of key-value copy vectors by calculating the cosine similarity between the two sets of key-value copy vectors; Euclidean distance can also be used to calculate the Euclidean distance between the two sets of key-value copies to evaluate their spatial differences.

[0067] In response to the attention divergence being not lower than a preset divergence threshold, an intermediate layer potential feature corresponding to a latter shallow layer calculation block of two adjacent shallow layer calculation blocks is determined as an intermediate variable of the serialization model.

[0068] In some embodiments, the preset divergence threshold refers to a preset value used to determine the intermediate variable. If the calculated attention divergence is not lower than the preset divergence threshold, it is considered that there is a significant difference in the output features between the two shallow calculation blocks. The former shallow calculation block can be used as a reference, and the latter shallow calculation block can determine the parameters of the deep and shallow boundary.

[0069] In some embodiments, if there are significant differences in the output features between two shallow computation blocks, the intermediate-layer potential features corresponding to the latter shallow computation block are determined as intermediate variables of the serialization model, which means that these features can be directly utilized in subsequent reasoning processes instead of being recalculated.

[0070] In some embodiments, in response to the attention divergence being not lower than a preset divergence threshold, the intermediate layer potential features corresponding to the latter shallow computing block of two adjacent shallow computing blocks are determined as intermediate variables of the serialization model, and it is possible to dynamically select which features are suitable for caching and reuse, thereby optimizing the efficiency of the entire reasoning process.

[0071] In some embodiments, determining the attention divergence corresponding to the two adjacent shallow computation blocks based on the key-value copies of the two adjacent shallow computation blocks includes:

[0072] Determine the feature graph of each shallow computation block based on the key-value copies of two adjacent shallow computation blocks;

[0073] In some embodiments, in the self-attention mechanism, the query, key, and value matrices are operated through dot product and other operations to calculate the attention score, and then the weighted sum is performed to obtain the feature map of each shallow calculation block.

[0074] Based on the feature map and the preset perturbation factor, the attention divergence corresponding to two adjacent shallow computation blocks is determined.

[0075] In some embodiments, based on the feature map and the preset perturbation factor, the mathematical expression for determining the attention divergence corresponding to two adjacent shallow computation blocks is as follows:

[0076]

[0077] Among them, i and j represent the i-th shallow computing block and the j-th shallow computing block respectively. i and j are adjacent, usually j=i+1. represents the i-th feature map and the j-th feature map, represents the attention divergence corresponding to the i-th feature map and the j-th feature map, c represents the element in the feature vector, a channel in the feature map; C represents the channel dimension, that is, the number of channels; represents the feature map normalized by Softmax (the channel dimension is normalized to a probability distribution so that the value on each channel is in the range [0,1] and the sum of all channel values is equal to 1), Indicates the preset disturbance factor, which is a small constant. (like is 10 to the power of -7), used to avoid the denominator being zero in the logarithmic function.

[0078] In some embodiments, after obtaining the intermediate variables of the serialized model in the first computing module, the model inference acceleration method further includes:

[0079] Determine the deep output features of the sequenced model in the first computing module based on the intermediate variables;

[0080] In some embodiments, based on the intermediate variables, the intermediate variables are used as inputs of the deep computing blocks in the first computing module to obtain the deep output features of the serialization model in the first computing module.

[0081] Based on the deep output features, the prediction results of the serialization model in the first computing module are determined.

[0082] In some embodiments, the prediction result of the serialized model in the first computing module is a result obtained through full calculation, and the full calculation refers to calculation through all shallow calculation blocks and deep calculation blocks.

[0083] In some embodiments, after determining the deep output features of the serialized model in the first computing module based on the intermediate variables, the model inference acceleration method further includes:

[0084] The deep output features are determined as inputs to the shallow computing blocks in the third computing module, so that the third computing module obtains the prediction results of the serialization model.

[0085] In some embodiments, when the next time step of the diffusion model does not require cache reuse, the deep output features need to be directly used as the input of the next shallow calculation block in the next time step.

[0086] In some embodiments, the third computing module refers to a computing module other than the first computing module and the second computing module, and full calculation is performed in the third computing module. This application does not limit the number of third computing modules, and the number of third computing modules can be determined according to user needs. If the speed of model reasoning is taken into account, the fewer the number of third computing modules, the better. If the accuracy of model reasoning is taken into account, the more the number of third computing modules, the better.

[0087] In some embodiments, the model inference acceleration method further includes:

[0088] Obtain the high-frequency ratio of the serialized model in the second computing module;

[0089] In some embodiments, according to the diffusion model, there are different characteristics of information generation at different time steps. Generally speaking, the image information generated in the early moments is mainly concentrated in low frequencies, and high-frequency information begins to be gradually enriched in subsequent steps. Therefore, based on this characteristic, the cache can be reused as much as possible in the early stages, and the full amount of computing blocks can be used for calculation when key information is generated in the later stages.

[0090] In some embodiments, the high-frequency ratio is used to determine at which diffusion step to reuse the potential features of the intermediate layer. High-frequency information represents the details, edges, textures and other parts of the image / signal that change dramatically. In a neural network, the more high-frequency information, the more details the currently processed data or features contain and are not suitable for direct reuse.

[0091] In response to the high-frequency ratio of the serialized model being not less than the preset high-frequency ratio, the target time step in the serialized model is determined. The target time step is used to distinguish the time step of the intermediate layer potential feature reuse and the time step of the full calculation. The full calculation includes the shallow calculation block calculation and the deep calculation block calculation in the second calculation module.

[0092] In some embodiments, the preset high-frequency ratio is used to determine whether to enter the full calculation process, and the preset high-frequency ratio can be determined based on historical experience.

[0093] In some embodiments, if the high-frequency ratio of the serialized model is not lower than the preset high-frequency ratio, it means that it currently includes more detailed information and is no longer suitable for feature reuse; if the high-frequency ratio of the serialized model is lower than the preset high-frequency ratio, it means that the intermediate layer potential feature reuse strategy can be adopted to reduce redundant calculations.

[0094] In some embodiments, by determining the target time step in the serialization model in response to the high-frequency ratio of the serialization model being not less than a preset high-frequency ratio, the intermediate layer potential features can be reused in the early time steps (with greater noise) of the diffusion model; while in the later time steps, i.e., the image detail recovery stage, the high-frequency information increases and the full amount is calculated to ensure the generation quality.

[0095] In some embodiments, obtaining the high-frequency ratio of the serialized model in the second computing module includes:

[0096] Get the attention weight matrix and input features. The attention weight matrix is used to indicate the degree of spatial attention to the input features.

[0097] In some embodiments, the input feature refers to the original image received by the model or the feature representation after preliminary transformation.

[0098] Based on the attention weight matrix, input features and the preset high-frequency cutoff frequency threshold, the high-frequency proportion of the serialization model in the second calculation module is determined.

[0099] In some embodiments, a preset high-frequency cutoff frequency threshold is used to distinguish high-frequency components from low-frequency components.

[0100] In some embodiments, the mathematical expression for determining the high-frequency ratio of the serialized model in the second computing module is as follows:

[0101]

[0102] Among them, fx, fy are the horizontal and vertical components in the frequency domain coordinates; F is the feature after Fourier transform; x, y, c represent the dimensions of H×W×C (height, width, number of channels); A is the attention weight matrix, with dimensions of H×W, which represents the degree of spatial attention to the input features. Indicates the preset high frequency cutoff frequency threshold.

[0103] In some embodiments, a high-frequency ratio is finally set by a threshold, thereby determining that reuse will no longer occur at later time steps. Specifically, in the later stages of the diffusion model (when t gradually decreases to 1), the high-frequency ratio will be very high. The later time steps can be determined based on historical experience. When t approaches 1, full calculation is usually performed, i.e., both shallow and deep calculation blocks participate in the calculation process.

[0104] Through this application, it includes obtaining the intermediate variables of the serialized model in the first computing module, the intermediate variables including at least one of a key-value copy, an intermediate-layer potential feature and a deep-layer output feature, the intermediate-layer potential feature is a feature in the first computing module and the second computing module whose similarity is not less than a preset similarity threshold, the intermediate-layer potential feature is determined by a shallow computing block in the first computing module, and the deep-layer output feature is determined by a deep computing block in the first computing module; the feature whose similarity is not less than a preset similarity threshold is determined as the input of the deep computing block in the second computing module, so that the second computing module obtains the prediction result of the serialized model, solves the technical problem of rapid expansion of the cache of the second computing module in the related scheme, resulting in slow model inference speed, and achieves the technical effect of reducing redundant calculations of the second computing module and thereby speeding up the speed of model inference.

[0105] An embodiment of the present application further provides a model inference acceleration system, the system comprising: a first computing module, a cache module, and at least one second computing module, wherein the cache module is communicatively connected to the first computing module and the second computing module respectively;

[0106] The first computing module is used for inference calculation of the serialized model to obtain intermediate variables, which include at least one of the key-value copy, the intermediate layer potential feature, and the deep output feature;

[0107] The cache module is used to cache the intermediate variables obtained by the inference calculation of the first calculation module;

[0108] The second computing module is used to perform inference calculations of the serialized model based on the intermediate variables in the cache module.

[0109] In some embodiments, the first computing module includes at least one shallow computing block and at least one deep computing block;

[0110] The shallow computation block is used to determine the latent features of the middle layer;

[0111] The deep computation block is used to determine the deep output features based on the intermediate layer potential features.

[0112] In some embodiments, as Figure 2 As shown, Figure 2This is a structural diagram of a model reasoning acceleration system provided by an embodiment of the present application. The reasoning phase of the diffusion model is counted down by time t. The first calculation module is Figure 2 The computing device corresponding to the feature cache step (step t) in the second computing module is Figure 2 The computing device corresponding to the feature reuse step (step t-1) in the first computing module and the second computing module includes a shallow computing block (shallow block) and a deep computing block (deep block). Noised Image is the input image, Predicted Noiset and Predicted Noiset-1 are the predicted outputs corresponding to steps t and t-1 in the diffusion model. ① The model input is forward-calculated in multiple computing devices using multiple computing blocks. The number of blocks here is determined by the specific architecture of the model algorithm and is generally around dozens. ② The key-value copies (K, V matrices) calculated by each shallow computing block in the first computing module are transferred to the cache module, that is, Figure 2 In a CXL-type2 device, ③ based on the feature divergence calculation of the key-value copies of two adjacent shallow computation blocks, decide whether to cache the intermediate-layer latent features in the shallow computation block. ④ Continue forward computing the remaining blocks and output the deep-layer output features. ⑤ Based on the deep-layer output features, output the predicted output (predicted noise) for this diffusion step. ⑥ After determining the shallow-to-deep boundary parameter, directly use the intermediate-layer latent features cached in step ③ as the input to the deep block for continued computation, skipping the computation of the shallow block. ⑦ Reuse the key-value copies from step ② in the deep block. ⑧ Continue computing the deep block until the output of this step is output.

[0113] In some embodiments, the cache module includes a decision control module, which includes at least a cache feature divergence calculation module and a deep output feature high and low frequency calculation module;

[0114] The cache feature divergence calculation module is used to determine a first parameter, and the first parameter is used to indicate a boundary value between a deep calculation block and a shallow calculation block;

[0115] The deep output feature high and low frequency calculation module is used to determine the second parameter, which is used to indicate the time step of feature reuse in the serialization model.

[0116] In some embodiments, the decision control module is the aforementioned Figure 2 The decision controller in , further, such as Figure 3 As shown, Figure 3This is a structural diagram of a decision control module provided in an embodiment of the present application. The cache feature divergence calculation module determines the first parameter based on the calculated value and the threshold. The first parameter is the deep-shallow boundary parameter, which is used to determine which layer of intermediate potential features to cache. If no caching is required, the full calculation is performed directly; the deep output feature high and low frequency calculation module is used to determine at which diffusion step the intermediate layer potential features are reused.

[0117] In some embodiments, the system further includes a cache reuse determination module;

[0118] The cache reuse determination module is used to determine whether to cache the intermediate layer potential features and whether to reuse the intermediate layer potential features according to the value of the target bit, and the target bit is a reserved bit in the original data frame format.

[0119] In some embodiments, taking a standard CXL flit data packet as an example, the H3-S2M DRS header + S2MNDR header information is used as a basis, and the reserved bits are the 7th and 6th bits of the 12th byte in the original data frame format.

[0120] In some embodiments, as Figure 4 As shown, Figure 4 This is a schematic diagram of the structure of a data frame format provided by an embodiment of the present application. In the original data frame format, bits 7 and 6 of the 12th byte are used as flags to indicate whether intermediate-layer latent feature caching and reuse are to be performed. The 13th byte stores a first parameter (the shallow-deep boundary parameter i, typically less than 256) and transmits it to the CXL device. This data packet is read by the CXL device and controls data flow. The 14th byte stores the cumulative value of the boundary parameter, which controls the number of blocks to be cached, thereby estimating the number of blocks to be skipped and providing control for extreme situations. The 15th byte stores the cumulative value of the reuse steps, which controls the number of reuse steps to prevent excessive reuse and the loss of important high-frequency information. Specifically, excessive reuse results in fewer full block calculations in the later stages of diffusion, when high-frequency information is most readily available.

[0121] In some embodiments, a customized CXL protocol packet transmission design is used to ensure high-speed communication between CXL devices and computing devices for transmitting critical control information.

[0122] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0123] The embodiment of the present application further provides a model reasoning acceleration device 500, Figure 5 A schematic diagram of the structure of a model reasoning acceleration device provided by an embodiment of the present disclosure is shown in FIG. Figure 5 Shown, including:

[0124] An acquisition unit 501 is configured to acquire intermediate variables of the serialized model in the first computing module, the intermediate variables including at least one of a key-value copy, an intermediate-layer potential feature, and a deep-layer output feature. The intermediate-layer potential feature is a feature whose similarity between the first computing module and the second computing module is not less than a preset similarity threshold. The intermediate-layer potential feature is determined by a shallow computing block in the first computing module, and the deep-layer output feature is determined by a deep computing block in the first computing module.

[0125] The determining unit 502 is configured to determine features whose similarity is not less than a preset similarity threshold as inputs to the deep computing block in the second computing module, so that the second computing module obtains a prediction result of the serialization model.

[0126] Furthermore, in a possible implementation of the embodiment of the present disclosure, the acquiring unit 501 is configured to:

[0127] Obtain key-value copies of two adjacent shallow computation blocks in the first computation module. The key-value copies are copies of the key-value matrix generated by the original image in the self-attention layer.

[0128] Based on the key-value copies of two adjacent shallow calculation blocks, the intermediate variables of the serialization model are determined.

[0129] Furthermore, in a possible implementation of the embodiment of the present disclosure, the acquiring unit 501 is configured to:

[0130] Based on the key-value copies of two adjacent shallow computing blocks, the attention divergence corresponding to the two adjacent shallow computing blocks is determined. The attention divergence is used to determine the similarity of the spatiotemporal features output by the two adjacent shallow computing blocks.

[0131] In response to the attention divergence being not lower than a preset divergence threshold, an intermediate layer potential feature corresponding to a latter shallow layer calculation block of two adjacent shallow layer calculation blocks is determined as an intermediate variable of the serialization model.

[0132] Furthermore, in a possible implementation of the embodiment of the present disclosure, the acquiring unit 501 is configured to:

[0133] Determine the feature graph of each shallow computation block based on the key-value copies of two adjacent shallow computation blocks;

[0134] Based on the feature map and the preset perturbation factor, the attention divergence corresponding to two adjacent shallow computation blocks is determined.

[0135] Furthermore, in a possible implementation of the embodiment of the present disclosure, the model reasoning acceleration device 500 further includes a prediction result determination unit, which is configured to:

[0136] Determine the deep output features of the sequenced model in the first computing module based on the intermediate variables;

[0137] Based on the deep output features, the prediction results of the serialization model in the first computing module are determined.

[0138] Furthermore, in a possible implementation of the embodiment of the present disclosure, the prediction result determination unit is further configured to:

[0139] The deep output features are determined as inputs to the shallow computing blocks in the third computing module, so that the third computing module obtains the prediction results of the serialization model.

[0140] Furthermore, in a possible implementation of the embodiment of the present disclosure, the model reasoning acceleration device 500 further includes a time step determination unit, which is configured to:

[0141] Obtain the high-frequency ratio of the serialized model in the second computing module;

[0142] In response to the high-frequency ratio of the serialized model being not less than the preset high-frequency ratio, the target time step in the serialized model is determined. The target time step is used to distinguish the time step of the intermediate layer potential feature reuse and the time step of the full calculation. The full calculation includes the shallow calculation block calculation and the deep calculation block calculation in the second calculation module.

[0143] Furthermore, in a possible implementation of the embodiment of the present disclosure, the time step determination unit is further configured to:

[0144] Get the attention weight matrix and input features. The attention weight matrix is used to indicate the degree of spatial attention to the input features.

[0145] Based on the attention weight matrix, input features and the preset high-frequency cutoff frequency threshold, the high-frequency proportion of the serialization model in the second calculation module is determined.

[0146] Through this application, it includes obtaining the intermediate variables of the serialized model in the first computing module, the intermediate variables including at least one of a key-value copy, an intermediate-layer potential feature and a deep-layer output feature, the intermediate-layer potential feature is a feature in the first computing module and the second computing module whose similarity is not less than a preset similarity threshold, the intermediate-layer potential feature is determined by a shallow computing block in the first computing module, and the deep-layer output feature is determined by a deep computing block in the first computing module; the feature whose similarity is not less than a preset similarity threshold is determined as the input of the deep computing block in the second computing module, so that the second computing module obtains the prediction result of the serialized model, solves the technical problem of rapid expansion of the cache of the second computing module in the related scheme, resulting in slow model inference speed, and achieves the technical effect of reducing redundant calculations of the second computing module and thereby speeding up the speed of model inference.

[0147] For the description of the features in the embodiment corresponding to the model reasoning acceleration device, please refer to the relevant description of the embodiment corresponding to the model reasoning acceleration method, and no further details will be given here.

[0148] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned model reasoning acceleration method embodiments.

[0149] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned model reasoning acceleration method embodiments when running.

[0150] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0151] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above-mentioned model reasoning acceleration method embodiments are implemented.

[0152] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned model reasoning acceleration method embodiments.

[0153] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0154] The above is a detailed introduction to a model reasoning acceleration method, system, electronic device, storage medium and product provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

Claims

1. A model reasoning acceleration method, characterized in that: include: Obtaining intermediate variables of the serialized model in the first computing module, the intermediate variables including at least one of a key-value copy, an intermediate-layer potential feature, and a deep-layer output feature, the intermediate-layer potential feature being a feature whose similarity between the first computing module and the second computing module is not less than a preset similarity threshold, the features whose similarity is not less than the preset similarity threshold being used to determine features that can be reused between different computing modules, the intermediate-layer potential feature being determined by a shallow computing block in the first computing module, and the deep-layer output feature being determined by a deep computing block in the first computing module; Determining the features whose similarity is not less than a preset similarity threshold as inputs to a deep computing block in the second computing module, so that the second computing module obtains a prediction result of the serialization model; Obtaining a high-frequency ratio of the serialized model in the second computing module; In response to the high-frequency ratio of the serialized model being not less than the preset high-frequency ratio, the target time step in the serialized model is determined, and the target time step is used to distinguish the time step of the intermediate layer potential feature reuse and the time step of the full calculation, and the full calculation includes the shallow calculation block calculation and the deep calculation block calculation in the second calculation module.

2. The model reasoning acceleration method according to claim 1, characterized in that: The obtaining of the intermediate variables of the serialized model in the first computing module includes: Obtain key-value copies of two adjacent shallow computation blocks in the first computation module, where the key-value copies are copies of the key-value matrix generated by the original image in the self-attention layer; Based on the key-value copies of the two adjacent shallow computation blocks, intermediate variables of the serialization model are determined.

3. The model reasoning acceleration method according to claim 2, characterized in that: The determining of the intermediate variables of the serialization model based on the key-value copies of the two adjacent shallow calculation blocks includes: Determining, based on the key-value copies of the two adjacent shallow computing blocks, attention divergence corresponding to the two adjacent shallow computing blocks, wherein the attention divergence is used to determine the similarity of spatiotemporal features output by the two adjacent shallow computing blocks; In response to the attention divergence being not lower than a preset divergence threshold, determining the intermediate layer potential feature corresponding to the latter shallow layer calculation block of the two adjacent shallow layer calculation blocks as the intermediate variable of the serialization model.

4. The model reasoning acceleration method according to claim 3, characterized in that: The determining, based on the key-value copies of the two adjacent shallow computation blocks, the attention divergence corresponding to the two adjacent shallow computation blocks comprises: Determining a feature graph of each shallow computing block based on the key value copies of the two adjacent shallow computing blocks; Based on the feature map and a preset disturbance factor, the attention divergence corresponding to the two adjacent shallow calculation blocks is determined.

5. The model reasoning acceleration method according to claim 1, characterized in that: After obtaining the intermediate variables of the serialized model in the first computing module, the method further includes: Determining, based on the intermediate variables, deep output features of the sequenced model in the first computing module; Based on the deep output features, a prediction result of the serialization model in the first computing module is determined.

6. The model reasoning acceleration method according to claim 5, characterized in that: After determining the deep output features of the serialization model in the first computing module based on the intermediate variables, the method further includes: Determine the deep output feature as the input of the shallow calculation block in the third calculation module, so that the third calculation module obtains the prediction result of the serialization model.

7. The model reasoning acceleration method according to claim 1, characterized in that: Obtaining the high-frequency ratio of the serialized model in the second computing module includes: Obtaining an attention weight matrix and input features, wherein the attention weight matrix is used to indicate a degree of spatial attention to the input features; Based on the attention weight matrix, the input features and the preset high-frequency cutoff frequency threshold, the high-frequency proportion of the serialization model in the second computing module is determined.

8. A model reasoning acceleration system, characterized in that: The system includes: a first computing module, a cache module and at least one second computing module, wherein the cache module is communicatively connected to the first computing module and the second computing module respectively; The first computing module is used for inference calculation of the serialized model to obtain intermediate variables, wherein the intermediate variables include at least one of a key-value copy, an intermediate layer potential feature, and a deep layer output feature; The cache module is used to cache the intermediate variables obtained by the inference calculation of the first calculation module; The second computing module is used to perform inference calculations of the serialization model based on the intermediate variables in the cache module.

9. The model reasoning acceleration system according to claim 8, characterized in that: The first computing module includes at least one shallow computing block and at least one deep computing block; The shallow layer calculation block is used to determine the intermediate layer potential features; The deep layer calculation block is used to determine deep layer output features based on the intermediate layer potential features.

10. The model reasoning acceleration system according to claim 8, characterized in that: The cache module includes a decision control module, which includes at least a cache feature divergence calculation module and a deep output feature high and low frequency calculation module; The cache feature divergence calculation module is used to determine a first parameter, where the first parameter is used to indicate a boundary value between a deep calculation block and a shallow calculation block; The deep output feature high and low frequency calculation module is used to determine a second parameter, which is used to indicate the time step of feature reuse in the serialization model.

11. The model reasoning acceleration system according to claim 8, characterized in that: The system also includes a cache reuse determination module; The cache multiplexing determination module is used to determine whether to cache the intermediate layer potential features and whether to multiplex the intermediate layer potential features according to the value of the target bit, and the target bit is a reserved bit in the original data frame format.

12. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the model reasoning acceleration method as claimed in any one of claims 1 to 7 when executing the computer program.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the model inference acceleration method according to any one of claims 1 to 7 are implemented.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the model reasoning acceleration method as claimed in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Model double-level decoding method and device

    CN118114655A

  • Inference method, device, equipment, cluster, product and medium

    CN119558410A