Deep learning inference platform and deep learning inference engine operation method and system

By using global or local tensor cache managers and hybrid attention computing in the deep learning inference engine, the problem of insufficient utilization of hardware characteristics in the existing technology is solved, more efficient memory management and multi-scene adaptability are achieved, and the stability and universality of the system are improved.

CN120297344APending Publication Date: 2025-07-11HUA DATA TECH (SHANGHAI) CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510424545.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing deep learning inference engines cannot fully utilize the hardware features, lack special optimization mechanisms for large-scale cache scenarios, cannot efficiently update and reusable caches, and lack flexible adjustment mechanisms for multi-scene adaptation.

Method used

By installing a deep learning inference engine in the processor, tensor reshaping and management is performed using global or local tensor cache managers, combining mixed attention calculations, output tensors are pre-allocated, and dynamically adapted to multi-step decoding or high-concurrency situations through a unified wrapper selection and memory management mechanism.

Benefits of technology

It improves the reliability and manageability of the deep learning inference engine, reduces memory fragmentation and synchronization conflicts, adapts to diverse text, audio or other sequence data scenarios, and improves versatility and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297344A_ABST
    Figure CN120297344A_ABST
Patent Text Reader

Abstract

The invention provides a deep learning inference platform and an operation method and system of a deep learning inference engine. The deep learning inference engine is carried in a processor, and the operation method comprises the following steps: loading a trained inference model, establishing a tensor cache manager, and pre-distributing an output tensor; receiving input data, obtaining an initial tensor corresponding to the input data, performing tensor remodeling on the initial tensor to obtain a target tensor of a target dimension, and storing the target tensor into a first cache variable of a tensor cache manager; calling a target tensor in the first cache variable and performing mixed attention calculation on the target tensor to obtain a mixed attention calculation result; and writing the mixed attention calculation result into an output tensor for outputting. The method can dynamically adapt to multi-step decoding or high concurrency situations, can adapt to diversified texts, audios or other sequence data and other scenes, enables a deep learning inference engine to have higher universality and expandability, and improves the system performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of deep learning technologies, and in particular, to a deep learning inference platform and a method and system for operating a deep learning inference engine. Background Art

[0002] With the continuous expansion of the scale of deep learning model development, the computational workload and memory requirements in the inference stage are also increasing rapidly. Existing deep learning inference engines or high-performance computing modules generally have limitations in that they cannot fully utilize hardware characteristics, lack a dedicated mechanism for large-scale caching scenarios, have scattered calls, redundant checks, and unstable memory occupancy for context caching, and also lack flexible strategies or a unified abstraction mechanism to handle high-concurrency small requests and long-sequence requests with low concurrency but large contexts, resulting in the need for repeated development or highly customized solutions. Summary of the Invention

[0003] The technical problem to be solved by the present disclosure is to overcome the defects in the prior art that hardware characteristics cannot be fully utilized, there is a lack of a dedicated optimization mechanism for large-scale caching scenarios, the cache cannot be efficiently updated and reused, and there is a lack of a flexible adjustment mechanism for multi-scenario adaptation, and to provide a deep learning inference platform and a method and system for operating a deep learning inference engine.

[0004] The present disclosure solves the above technical problem through the following technical solutions:

[0005] According to a first aspect of the present disclosure, there is provided a method for operating a deep learning inference engine, where the deep learning inference engine is mounted in a processor, and the operating method includes:

[0006] Loading a trained inference model, establishing a tensor cache manager, and pre-allocating output tensors;

[0007] Receiving input data and obtaining an initial tensor corresponding to the input data, performing tensor reshaping on the initial tensor to obtain a target tensor with a target dimension, and storing the target tensor in a first cache variable of the tensor cache manager;

[0008] Invoking the target tensor in the first cache variable and performing hybrid attention calculation on the target tensor to obtain a hybrid attention calculation result;

[0009] Writing the hybrid attention calculation result into the output tensor for output.

[0010] Optionally, the hybrid attention includes global attention and local attention;

[0011] The step of calling the target tensor in the first cache variable and performing hybrid attention calculation on the target tensor to obtain a hybrid attention calculation result includes:

[0012] Call the target tensor in the first cache variable;

[0013] In response to the sequence length of the target tensor being greater than a preset threshold, perform the global attention calculation on the elements in the target tensor that satisfy a first preset spacing to obtain a global attention calculation result, and perform the local attention calculation on the elements in the target tensor that satisfy a second preset spacing to obtain a local attention calculation result;

[0014] Wherein, the first preset spacing is greater than the second preset spacing;

[0015] Perform a merging process on the global attention calculation result and the local attention calculation result to obtain the hybrid attention calculation result.

[0016] Optionally, the step of calling the target tensor in the first cache variable and performing hybrid attention calculation on the target tensor to obtain a hybrid attention calculation result further includes:

[0017] Slice the target tensor into several blocks, and perform attention calculation on each block using a scaled dot-product attention mechanism to obtain several block calculation results;

[0018] Perform a concatenation process on the several block calculation results to obtain the hybrid attention calculation result.

[0019] Optionally, the step of performing attention calculation on each block using a scaled dot-product attention mechanism to obtain several block calculation results includes:

[0020] Perform linear transformation on each block using the matrix dimension parameters of query, key, and value in each head, and perform scaled dot-product attention and softmax (normalized exponential function) operation calculations to obtain the feature matrix output by each head;

[0021] In response to the end of forward propagation, store the feature matrix in the second cache variable of the tensor cache manager;

[0022] Call the feature matrix in the second cache variable for concatenation processing to obtain the block calculation result corresponding to each block.

[0023] Optionally, the running method further includes:

[0024] In response to the start of forward propagation, use an aggregation selection function to determine the target scenario corresponding to the input data;

[0025] Obtain a corresponding target wrapper based on the target scenario, and map the target scenario into the target wrapper;

[0026] Wherein, scenarios corresponding to different wrappers are pre-configured in the deep learning inference engine;

[0027] And / or,

[0028] The step of writing the hybrid attention calculation result into the output tensor for output includes:

[0029] Check the type and dimension of the hybrid attention calculation result;

[0030] In response to the type and the dimension meeting a first preset condition, write the hybrid attention calculation result into the output tensor for output;

[0031] In response to the type and / or the dimension not meeting the first preset condition, reallocate matching memory to form a new output tensor, and write the hybrid attention calculation result into the new output tensor for output;

[0032] And / or,

[0033] The running method further includes: in response to the end of the execution of the kernel function and the absence of subsequent call dependencies, cancel or postpone the display synchronization call for the execution result of the kernel function;

[0034] And / or,

[0035] The running method further includes: checking the data input into the deep learning inference engine and the device attributes of the processor.

[0036] According to a second aspect of the present disclosure, there is provided a running system of a deep learning inference engine, the deep learning inference engine is carried in a processor, and the running system includes an initialization module, a tensor reshaping module, an attention calculation module and an output module;

[0037] The initialization module is used to load a trained inference model, establish a tensor cache manager, and pre-allocate an output tensor;

[0038] The tensor reshaping module is used to receive input data and obtain an initial tensor corresponding to the input data, perform tensor reshaping on the initial tensor to obtain a target tensor with a target dimension, and store the target tensor in a first cache variable of the tensor cache manager;

[0039] The attention calculation module is used to call the target tensor in the first cache variable and perform hybrid attention calculation on the target tensor to obtain a hybrid attention calculation result;

[0040] The output module is used to write the mixed attention calculation result into the output tensor for output.

[0041] According to a third aspect of the present disclosure, a deep learning inference platform is provided, and the deep learning inference platform includes an operating system of the deep learning inference engine described in the second aspect of the present disclosure.

[0042] According to a fourth aspect of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and configured to run on the processor. When the processor executes the computer program, the operating method of the deep learning inference engine described in the first aspect of the present disclosure is implemented.

[0043] According to a fifth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the operating method of the deep learning inference engine described in the first aspect of the present disclosure is implemented.

[0044] According to a sixth aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the operating method of the deep learning inference engine described in the first aspect of the present disclosure is implemented.

[0045] On the basis of conforming to the common knowledge in the art, the above preferred conditions can be combined arbitrarily to obtain various preferred examples of the present disclosure.

[0046] The positive and progressive effects of the present disclosure are as follows: By using a global or local tensor cache manager, tensors are reshaped and managed uniformly, and a scalable data access layer is stored to dynamically adapt to multi-step decoding or high-concurrency scenarios, reducing unnecessary data movement, memory fragmentation, and synchronization conflicts, making the deep learning inference engine have higher reliability and manageability. And through pre-allocation, memory management is advanced and unified, reducing the impact of repeated memory applications in high-concurrency scenarios on the stability of the processor. At the same time, mixed attention calculation is adaptively performed in combination with the current sequence length and context requirements, which not only solves the redundancy problem of pure global attention calculation in long sequences but also overcomes the bottleneck that pure local attention cannot capture long-range dependencies, and can adapt to diverse scenarios such as text, audio, or other sequence data, making the deep learning inference engine have higher generality and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a flowchart of the operating method of the deep learning inference engine in Embodiment 1;

[0048] Figure 2 It is a flowchart of step S3 in the operating method of the deep learning inference engine in Embodiment 1;

[0049] Figure 3 Schematic diagram of the modules of the operating system of the deep learning inference engine in Embodiment 2;

[0050] Figure 4 Schematic diagram of the structure of the electronic device in Embodiment 4. Detailed implementation manners

[0051] The present disclosure will be further described below by way of embodiments, but the present disclosure is not limited to the scope of the described embodiments.

[0052] In the embodiments of the present disclosure, prefix words such as "first" and "second" are only used to distinguish different described objects, and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of ordinal numbers and other prefix words for distinguishing described objects in the embodiments of the present disclosure does not constitute a limitation on the described objects. The statements of the described objects refer to the descriptions in the claims or the context of the embodiments, and should not constitute unnecessary limitations due to the use of such prefix words. In addition, in the description of the present embodiment, unless otherwise specified, the meaning of "a plurality" is two or more.

[0053] In a deep learning inference engine or a high-performance computing module, multi-head attention needs to frequently perform batch matrix operations on three tensors: query (Q), key (K), and value (V). However, existing general frameworks often cannot fully utilize hardware characteristics for optimization, or lack a dedicated mechanism for large-scale cache scenarios. To avoid repeated calculations on the same context, many inference engines use a KV cache (key-value cache) to store historical states, so as to quickly generate predictions for subsequent moments. However, in terms of how to efficiently update or reuse the cache, existing solutions often have limitations such as scattered calls, redundant checks, and unstable memory occupancy. In addition, deep learning inference has both online real-time scenarios and offline batch scenarios, both high-concurrency small requests and low-concurrency but long-sequence requests with large contexts. Existing model frameworks lack flexible strategies or unified abstraction mechanisms to handle these differences, resulting in the need for repeated development or high customization.

[0054] Specifically, existing deep learning inference engines or high-performance computing modules have the following limitations or deficiencies in large-scale inference scenarios:

[0055] 1) Redundant wrapper selection logic

[0056] Currently, many inference frameworks often use a large number of if-else (conditional judgment) branches to perform wrapper selection when processing different types of inputs or modes, resulting in high code maintenance complexity and insufficient scalability.

[0057] 2) Repeated Tensor Shape Reshaping and Memory Operations

[0058] The current Q, K, and V tensors are repeatedly reshaped when calculating multi-head attention. Due to the lack of a unified caching and management strategy, these redundant operations will accumulate in multiple inference loops, wasting precious hardware resources.

[0059] 3) Frequent Data Copying and Memory Allocation

[0060] In the decoding stage of model inference, especially for auto-regressive or recurrent generation scenarios, if tensors cannot be effectively reused, continuous data copying, memory allocation, and deallocation will occur, affecting system stability.

[0061] 4) Unnecessary Synchronization Operations

[0062] Some explicit synchronizations (e.g., the explicit synchronization call function torch.cuda.synchronize()) are only really needed for debugging or very few special scenarios. However, many current inference frameworks default to performing synchronization after important operators, resulting in potential suppression of GPU (Graphics Processing Unit) parallel performance.

[0063] 5) Improper Invocation Location of the store_kv_cache (a function for storing key-value pair cache)

[0064] KV cache usually lacks unified planning in terms of update timing, data indexing, and memory layout, resulting in scattered calls or repeated checks in multiple places, which is not conducive to simplifying the process and making full use of the cache.

[0065] 6) Redundant Conditional Judgments and Assertions

[0066] When dealing with multiple sizes, types, or hardware backends, repeated and scattered conditional judgments often occur, resulting in reduced system execution efficiency and increased maintenance difficulty.

[0067] 7) Unreasonable Allocation Method of Output Tensors

[0068] The current deep learning inference engines ignore the memory management requirements of high-concurrency or multi-step decoding scenarios and often need to reapply for output tensors in each forward propagation, resulting in serious memory fragmentation.

[0069] 8) Lack of Deep Utilization of Hardware Characteristics

[0070] Large-scale matrix operations or attention calculations fail to optimize in combination with algorithms such as block strategies and hybrid attention and hardware characteristics, making it difficult to meet the requirements of long sequence and large batch inferences.

[0071] In applications such as large-scale language models, conversational robots, automatic question and answer, machine translation, and text summarization, reasoning efficiency and resource management are key factors affecting the overall system quality and user experience.

[0072] In view of this, the present disclosure provides a deep learning reasoning platform and an operating method and system for a deep learning reasoning engine to solve the problems that existing deep learning reasoning engines cannot fully utilize hardware characteristics and lack a special optimization mechanism for large-scale caching scenarios, cannot efficiently update and reuse caches, and lack a flexible adjustment mechanism for multi-scenario adaptation.

[0073] Example 1

[0074] In a specific embodiment of the present disclosure, a method for operating a deep learning inference engine is provided. The deep learning inference engine is mounted in a processor, such as Figure 1 As shown, the operation method includes:

[0075] S1. Load the trained inference model, establish a tensor cache manager, and pre-allocate output tensors.

[0076] S2, receiving input data and obtaining an initial tensor corresponding to the input data, reshaping the initial tensor to obtain a target tensor of a target dimension, and storing the target tensor in a first cache variable of a tensor cache manager;

[0077] S3, calling the target tensor in the first cache variable and performing a hybrid attention calculation on the target tensor to obtain a hybrid attention calculation result;

[0078] S4. Write the mixed attention calculation results into the output tensor for output.

[0079] Specifically, the deep learning inference engine can be installed on a variety of hardware platforms, such as GPU, TPU (tensor processing unit), etc.

[0080] Through step S1, the trained neural network inference model is loaded, and the model weight of the inference model is obtained to make predictions and inferences based on the model weights, and a tensor cache manager is established to use a set of global or local tensor cache managers to uniformly manage tensors in the inference process. Before executing the operator, the tensor attributes are matched and checked. If the type and shape of the tensor match the cache variables in the tensor cache manager, the cache variables are directly reused or in-situ operations are used to reasonably map the output and intermediate results to the same storage area. If they do not match, memory expansion allocation is performed on demand. By adopting the three-step strategy of "reuse-check-allocation", pre-allocation and reuse are combined, supplemented by a fast check mechanism for input tensors and cached tensors, and dynamically adapt to multi-step decoding or high concurrency situations to reduce unnecessary data movement and avoid frequent allocation and release of large blocks of memory during multi-step reasoning, which significantly reduces the problem of memory fragmentation and enhances the predictability of the system.

[0081] At the same time, before multi-step reasoning or high-concurrency reasoning, pre-allocate the output tensor for storing the reasoning results. For example, use torch.empty_like(q) (a function for creating a new tensor) and q.new_empty() (a function for creating a new tensor) to create the output area at one time, and then delay or reallocate it when the situation changes (for example, the sequence length changes), so that multiple calls are minimized and adjusted only when needed.

[0082] In multi-head attention, the initial query (Q), key (K), and value (V) tensors need to be reshaped as multi-head shapes. Step S2 receives input data and obtains the initial tensor corresponding to the input data. At the beginning of forward propagation (e.g., extend_forward or decode_forward), a tensor reshaping operation (e.g., view, contiguous, or similar operations) is performed to reshape the Q, K, and V tensors to the target dimensions required for multi-head attention, and the reshaped target tensors are saved in predefined cache variables. Each attention head only needs to index into the same cache array. If the subsequent calculation process meets the conditions, the reshaped target tensor is directly obtained from the cache variable through a pointer or reference for calling. By performing a tensor reshaping operation only once at the beginning of forward propagation, the tensor reshaping operation is uniformly pre-placed and cached, and all tensor reshaping and data layout are completed at the beginning of forward propagation, forming an extensible data access layer, avoiding repeated calls to view and contiguous, and having higher versatility and scalability compared to the existing decentralized step-by-step reshaping operations.

[0083] Then, in step S3, the target tensor in the cache variable is called for hybrid attention calculation. Through the adaptive switching between global attention and local attention, the hybrid attention calculation result that takes into account both long distances and short distances is obtained.

[0084] In step S4, the hybrid attention calculation result of the multi-head attention calculation is written into the pre-allocated output tensor for output, so as to be returned to the upper-layer application.

[0085] In this specific embodiment, by using a global or local tensor cache manager to uniformly reshape and manage tensors, a scalable data access layer is stored and formed, dynamically adapting to multi-step decoding or high-concurrency scenarios, reducing unnecessary data movement, reducing memory fragmentation and synchronization conflicts, making the deep learning inference engine have higher reliability and manageability, and through pre-allocation, memory management is advanced and unified, reducing the impact of repeated memory applications in high-concurrency scenarios on the processor stability. At the same time, by adaptively performing hybrid attention calculation in combination with the current sequence length and context requirements, not only the redundancy problem of pure global attention calculation in long sequences is solved, but also the bottleneck that pure local attention cannot capture long-range dependencies is overcome, and it can adapt to diverse scenarios such as text, audio, or other sequence data, making the deep learning inference engine have higher generality and scalability.

[0086] In a specific embodiment, the hybrid attention includes global attention and local attention; as Figure 2 shown, step S3 includes:

[0087] S31. Call the target tensor in the first cache variable;

[0088] S32. In response to the sequence length of the target tensor being greater than a preset threshold, perform global attention calculation on the elements in the target tensor that satisfy the first preset spacing to obtain the global attention calculation result, and perform local attention calculation on the elements in the target tensor that satisfy the second preset spacing to obtain the local attention calculation result;

[0089] wherein, the first preset spacing is greater than the second preset spacing;

[0090] S33. Perform a merging process on the global attention calculation result and the local attention calculation result to obtain the hybrid attention calculation result.

[0091] Specifically, when calculating the attention of the target tensor, local attention or global attention is dynamically selected according to the sequence length of the target tensor and the information requirements of the above context. If the context within the window is crucial, local attention is applied to maintain a high-density fine-grained calculation within the local window to capture key neighboring dependencies; if distant information is required, global attention is enabled. For example, when the sequence length exceeds a preset threshold, low-density or sparse global attention is used for the distant dependence region; then the calculation results of local attention and global attention are merged and output to obtain an attention distribution that balances efficiency and context coverage.

[0092] By combining local attention and global attention, the sequence length is dynamically identified and the close-range and long-range dependencies in the sequence are adaptively processed. Fine-grained calculations are performed in most of the local windows that need to be focused on, and sparse or global modes are used in more distant or low-impact regions, solving the redundancy problem of pure global attention calculation in long sequences and overcoming the bottleneck that pure local attention cannot capture long-range dependencies. Compared with global attention, it can significantly reduce the redundant scanning of distal context and adapt to multi-language or cross-domain applications. It can be used for text reasoning and extended to scenarios such as speech and vision that require local attention and global attention.

[0093] In a specific embodiment, step S3 further includes: dividing the target tensor into several blocks, using the scaled dot-product attention mechanism to calculate the attention for each block to obtain several block calculation results; performing a splicing process on the several block calculation results to obtain a mixed attention calculation result;

[0094] Among them, when calculating the attention of each block, the matrix dimension parameters of the query, key, and value in each head perform a linear transformation on each block and perform scaled dot-product attention and softmax operation calculations to obtain the feature matrix output by each head; in response to the end of the forward propagation, the feature matrix is stored in the second cache variable of the tensor cache manager; the feature matrix in the second cache variable is called for splicing processing to obtain the block calculation result corresponding to each block.

[0095] Specifically, the matrix multiplications of multi-head attention (e.g., QKᵀ, KV, etc.) are sliced into several blocks with controlled sizes; local parallelism is achieved using thread blocks or other parallel mechanisms, and operations such as scaled dot-product attention and softmax are performed within each block to obtain the feature matrix output by each head. At the end of the forward propagation, store_kv_cache is uniformly called to store the feature matrix into the cache variable of the tensor cache manager to update the context cache. Then, by calling the feature matrix in the cache variable, the feature matrices are concatenated in a manner matching the block slicing to obtain the block calculation result corresponding to each block. The block calculation results are temporarily stored in the cache (static random access memory SRAM or shared memory), and all block calculation results are concatenated to obtain the complete hybrid attention calculation result.

[0096] Among them, in cloud distributed deployment or on edge devices, the block size can be adaptively adjusted according to resource conditions, so as to achieve a more suitable balance of memory usage and operation efficiency by configuring the block size and hybrid attention parameters in different hardware environments.

[0097] By making full use of shared memory or on-chip cache at the hardware level, reducing the dependence on the bandwidth of the video memory (HBM), a stable and efficient data processing flow can be provided, which is especially effective in the case of long sequences or large batches. For the possible long sequence input scenarios, block matrix operations and hybrid attention mechanisms are adopted, and inter-block parallel computing and asynchronous calls are performed, which not only improve the data processing speed but also reduce the risk probability of "video memory explosion" or access latency such as failures or crashes caused by a large number of memory operations. At the same time, the context cache (KV cache) is updated at the end of each step of inference, without having to be called multiple times in multiple branches. In other scenarios, it is only referenced or read without repeated updates to avoid index misalignment or data overwrite. By uniformly managing the update timing of the KV cache, repeated updates in multiple logical branches are avoided, ensuring the consistency of cache writing timing and index, thus reducing repeated operations and ensuring context consistency, making the existing hardware resources more reasonably arranged.

[0098] In a specific embodiment, the running method further includes:

[0099] In response to the start of forward propagation, an aggregation selection function is used to determine the target scenario corresponding to the input data;

[0100] Based on the target scenario, the corresponding target wrapper is obtained, and the target scenario is mapped into the target wrapper;

[0101] Among them, different scenarios corresponding to different wrappers are pre-configured in the deep learning inference engine.

[0102] Specifically, multiple types of wrappers are pre-configured, and the applicable scenarios and conditions are recorded during the initialization phase of the deep learning inference engine. At the start of the forward propagation, an auxiliary function (e.g., an aggregation selection function) is used to determine the target scenario corresponding to the input data at once. The target scenario includes at least one of the input data type, operating mode, and environment information / deployment scenario. Based on the target scenario, the matching wrapper type is determined and the corresponding wrapper instance is returned, mapping the target scenario to the corresponding wrapper instance. Subsequent multi-head attention calculations are all executed within the wrapper, skipping the cumbersome multiple if-else nestings, avoiding branch judgments in subsequent steps again, breaking the traditional scattered wrapper selection structure, constructing a pluggable and extensible wrapper management system, simplifying the wrapper selection logic, significantly simplifying the difficulty of later maintenance and transformation, and reducing the code complexity and maintenance cost.

[0103] In a specific embodiment, step S4 includes:

[0104] Check the type and dimension of the mixed attention calculation result;

[0105] In response to the type and dimension meeting the first preset condition, write the mixed attention calculation result into the output tensor for output;

[0106] In response to the type and / or dimension not meeting the first preset condition, reallocate the matching memory to form a new output tensor, and write the mixed attention calculation result into the new output tensor for output.

[0107] Specifically, after obtaining the mixed attention calculation result, quickly check the type and dimension of the mixed attention calculation result. If the type and dimension match the pre-allocated output tensor, reuse the original memory, write the mixed attention calculation result into the output tensor for output, without new memory allocation. Otherwise, reallocate the matching memory according to the dimension of the mixed attention calculation result to form a new output tensor, and then write the mixed attention calculation result into the new output tensor for output.

[0108] By pre-allocating the output tensor, memory management is advanced and unified, reducing data copying and memory allocation, reducing the impact of repeated memory applications in high-concurrency scenarios on system stability, and at the same time being able to reallocate memory according to the actual change in sequence length to meet the inference requirements of the deep learning inference engine.

[0109] In a specific embodiment, the running method further includes: in response to the end of the kernel function execution and the absence of subsequent call dependencies, cancel or postpone the display synchronous call to the execution result of the kernel function.

[0110] Specifically, after the kernel execution is completed, it is determined whether the output data has been completely written according to the dependency relationship. If the subsequent operators can automatically complete the dependency synchronization and there are no subsequent dependencies that must be waited for, there is no need for additional synchronization. The explicit synchronization calls such as torch.cuda.synchronize() (a function for explicit synchronization call) are cancelled or postponed through the debug switch. At the same time, synchronization can also be re-enabled in special scenarios (such as debugging, detecting exceptions, special dependencies, etc.). Through flexible dependency management, the previous default mode of "synchronization immediately after the operator ends" is changed, irrelevant or redundant synchronization is removed, and strict step-by-step synchronization is avoided, enabling pipelined processing on GPU parallel devices and improving the processing efficiency of GPU parallel computing.

[0111] In a specific embodiment, the running method further includes: checking the data input to the deep learning inference engine and the device attributes of the processor.

[0112] Specifically, a check module or common function is established at the entrance of the deep learning inference engine or in the common module to check information such as the data dimension, data type of the data input to the deep learning inference engine, and the device attributes of the processor. If it passes, the subsequent inference running steps are entered. Among them, the check of tensors such as K, V, and Q can also be completed uniformly in this link. By performing an overall check on the data input to the deep learning inference engine and the running hardware environment at the entrance stage of the deep learning inference engine, the assertion rules are centrally maintained, the maintainability is improved, and the repeated determination within each sub-module is reduced in multiple call levels, thereby optimizing the running overhead and execution path, keeping the inference process controllable and concise, and improving the overall performance of the system.

[0113] In a specific example, for the online dialogue scenario, the overhead of branch judgment is reduced through the centralized wrapper selection logic, and real-time response is achieved by combining the chunked attention calculation. For the batch text generation scenario, the memory fragmentation problem is significantly reduced in long sequence generation by using the pre-allocated output tensor and KV cache optimization. For the multi-scenario hybrid inference, different requirements of short text and long sequence scenarios can be adapted by dynamically adjusting the chunking and hybrid attention.

[0114] In this embodiment, by using a global or local tensor cache manager, tensors are reshaped and managed uniformly, storing to form an extensible data access layer, dynamically adapting to multi-step decoding or high-concurrency scenarios, reducing unnecessary data movement, reducing memory fragmentation and synchronization conflicts, making the deep learning inference engine more reliable and manageable, and through pre-allocation, memory management is advanced and unified, reducing the impact of repeated memory applications on the processor stability in high-concurrency scenarios. At the same time, by adaptively performing hybrid attention calculation in combination with the current sequence length and context requirements, it not only solves the redundancy problem of pure global attention calculation in long sequences, but also overcomes the bottleneck that pure local attention cannot capture long-range dependencies, and can adapt to diverse scenarios such as text, audio, or other sequence data, making the deep learning inference engine more general and extensible.

[0115] Embodiment 2

[0116] In a specific embodiment of the present disclosure, a running system of a deep learning inference engine is provided. The deep learning inference engine is installed in a processor, as Figure 3 shown, the running system includes an initialization module 100, a tensor reshaping module 200, an attention calculation module 300, and an output module 400;

[0117] The initialization module 100 is used to load the trained inference model, establish a tensor cache manager, and pre-allocate output tensors;

[0118] The tensor reshaping module 200 is used to receive input data and obtain the initial tensor corresponding to the input data, reshape the initial tensor to obtain a target tensor with a target dimension, and store the target tensor in the first cache variable of the tensor cache manager;

[0119] The attention calculation module 300 is used to call the target tensor in the first cache variable and perform hybrid attention calculation on the target tensor to obtain a hybrid attention calculation result;

[0120] The output module 400 is used to write the hybrid attention calculation result into the output tensor for output.

[0121] In a specific implementation manner, the hybrid attention includes global attention and local attention;

[0122] The attention calculation module 300 includes a calling unit, a calculation unit, and a merging unit;

[0123] The calling unit is used to call the target tensor in the first cache variable;

[0124] The computing unit is configured to, in response to the sequence length of the target tensor being greater than a preset threshold, perform global attention calculation on the elements in the target tensor that satisfy a first preset spacing to obtain a global attention calculation result, and perform local attention calculation on the elements in the target tensor that satisfy a second preset spacing to obtain a local attention calculation result;

[0125] wherein the first preset spacing is greater than the second preset spacing;

[0126] The merging unit is configured to perform a merging process on the global attention calculation result and the local attention calculation result to obtain a hybrid attention calculation result.

[0127] In a specific embodiment, the attention calculation module 300 further includes a splitting calculation unit and a splicing unit;

[0128] The splitting calculation unit is configured to split the target tensor into a plurality of blocks, and perform attention calculation on each block by using a scaled dot-product attention mechanism to obtain a plurality of block calculation results;

[0129] The splicing unit is configured to perform a splicing process on the plurality of block calculation results to obtain a hybrid attention calculation result.

[0130] In a specific embodiment, the splitting calculation unit is further configured to perform a linear transformation on each block according to the matrix dimension parameters of the query, key, and value in each head, and perform scaled dot-product attention and softmax operation calculations to obtain a feature matrix output by each head; in response to the end of the forward propagation, store the feature matrix in the second cache variable of the tensor cache manager; call the feature matrix in the second cache variable to perform a splicing process to obtain the block calculation result corresponding to each block.

[0131] In a specific embodiment, the running system further includes a wrapper selection module;

[0132] The wrapper selection module is configured to, in response to the start of the forward propagation, determine a target scenario corresponding to the input data by using an aggregation selection function; obtain a corresponding target wrapper based on the target scenario, and map the target scenario into the target wrapper;

[0133] wherein different scenarios corresponding to different wrappers are pre-configured in the deep learning inference engine.

[0134] In a specific embodiment, the output module 400 is further configured to check the type and dimension of the hybrid attention calculation result; in response to the type and dimension satisfying a first preset condition, write the hybrid attention calculation result into an output tensor for output; in response to the type and / or dimension not satisfying the first preset condition, reallocate a matching memory to form a new output tensor, and write the hybrid attention calculation result into the new output tensor for output.

[0135] In a specific embodiment, the operating system further includes a display synchronization call module;

[0136] The display synchronization call module is used to cancel or postpone the display synchronization call for the execution result of the kernel function in response to the end of the execution of the kernel function and the absence of subsequent call dependencies.

[0137] In a specific embodiment, the operating system further includes an inspection module;

[0138] The inspection module is used to inspect the data input into the deep learning inference engine and the device attributes of the processor.

[0139] In this embodiment, by modularizing functions such as wrapper selection, tensor reshaping, and attention calculation, they can be used independently or flexibly scheduled in a unified orchestration script. Among them, the modules are associated through interface definitions, and their respective functions are clear, facilitating subsequent expansion or migration.

[0140] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution.

[0141] In this embodiment, by using a global or local tensor cache manager, tensors are uniformly reshaped and managed, storing to form an extensible data access layer, dynamically adapting to multi-step decoding or high-concurrency situations, reducing unnecessary data movement, reducing memory fragmentation and synchronization conflicts, making the deep learning inference engine have higher reliability and manageability, and through pre-allocation, memory management is preposed and uniformly managed, reducing the impact of repeated memory applications in high-concurrency scenarios on the stability of the processor. At the same time, by adaptively performing hybrid attention calculation in combination with the current sequence length and context requirements, it not only solves the redundancy problem of pure global attention calculation in long sequences, but also overcomes the bottleneck that pure local attention cannot capture long-range dependencies, and can adapt to diverse scenarios such as text, audio, or other sequence data, making the deep learning inference engine have higher generality and scalability.

[0142] Embodiment 3

[0143] In a specific embodiment of the present disclosure, a deep learning inference platform is provided, and the deep learning inference platform includes the operating system of the deep learning inference engine described in any of the above embodiments.

[0144] Specifically, a deep learning inference engine is installed on the deep learning inference platform to perform prediction and inference.

[0145] In this embodiment, by using a global or local tensor cache manager, tensors are reshaped and managed uniformly, and a scalable data access layer is stored and formed, dynamically adapting to multi-step decoding or high-concurrency scenarios, reducing unnecessary data movement, reducing memory fragmentation and synchronization conflicts, making the deep learning inference engine more reliable and manageable. And through pre-allocation, memory management is advanced and unified, reducing the impact of repeated memory applications on the processor stability in high-concurrency scenarios. At the same time, hybrid attention calculation is adaptively performed in combination with the current sequence length and context requirements, which not only solves the redundancy problem of pure global attention calculation in long sequences, but also overcomes the bottleneck that pure local attention cannot capture long-distance dependencies, and can adapt to diverse scenarios such as text, audio, or other sequence data, making the deep learning inference engine more general and scalable.

[0146] Embodiment 4

[0147] Figure 4 It is a schematic structural diagram of an electronic device shown in an exemplary embodiment of the present disclosure. The electronic device includes a memory, a processor, and a computer program stored on the memory and configured to run on the processor. When the processor executes the computer program, it implements the operating system of the deep learning inference engine described in any of the above embodiments. Figure 4 The shown electronic device 30 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0148] As Figure 4 shown, the electronic device 30 may be presented in the form of a general computing device, for example, it may be a server device. The components of the electronic device 30 may include, but are not limited to: at least one of the above-mentioned processors 31, at least one of the above-mentioned memories 32, and a bus 33 connecting different system components (including the memory 32 and the processor 31).

[0149] The bus 33 includes a data bus, an address bus, and a control bus.

[0150] The memory 32 may include volatile memory, such as random access memory (RAM) 321 and / or cache memory 322, and may further include read-only memory (ROM) 323.

[0151] The memory 32 may further include a program tool 325 (or utility) having a set (at least one) of program modules 324. Such program modules 324 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0152] The processor 31 executes various functional applications and data processing by running the computer program stored in the memory 32, such as the operating system of the deep learning inference engine provided in any of the above embodiments.

[0153] The electronic device 30 can also communicate with one or more external devices 34 (such as a keyboard, a pointing device, etc.). Such communication can be carried out through the input / output (I / O) interface 35. Moreover, the electronic device 30 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through the network adapter 36. As shown in the figure, the network adapter 36 communicates with other modules of the electronic device 30 through the bus 33. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 30, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (redundant array of independent disks) systems, magnetic tape drives, and data backup storage systems, etc.

[0154] It should be noted that, although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0155] Embodiment 5

[0156] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the operating system of the deep learning inference engine provided in any of the above embodiments.

[0157] Among them, the computer-readable storage medium can more specifically include but not limited to: a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0158] Embodiment 6

[0159] The embodiments of the present disclosure also provide a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the operating system of the deep learning inference engine described in any one of the above.

[0160] Among them, the program code for executing the computer program product of the present disclosure can be written in any combination of one or more programming languages, and the program code can be executed entirely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or entirely on a remote device.

[0161] Although the specific embodiments of the present disclosure have been described above, those skilled in the art should understand that this is only an example, and the protection scope of the present disclosure is defined by the appended claims. Without departing from the principles and essence of the present disclosure, those skilled in the art can make various changes or modifications to these embodiments, but these changes and modifications all fall within the protection scope of the present disclosure.

Claims

1. A method for running a deep learning inference engine, characterized in that The deep learning inference engine is carried in a processor, and the running method includes: Loading a trained inference model, establishing a tensor cache manager, and pre-allocating an output tensor; Receiving input data and obtaining an initial tensor corresponding to the input data, performing tensor reshaping on the initial tensor to obtain a target tensor with a target dimension, and storing the target tensor in a first cache variable of the tensor cache manager; Invoking the target tensor in the first cache variable and performing hybrid attention calculation on the target tensor to obtain a hybrid attention calculation result; Writing the hybrid attention calculation result into the output tensor for output.

2. The operating method according to claim 1, characterized in that, The hybrid attention includes global attention and local attention; The step of invoking the target tensor in the first cache variable and performing hybrid attention calculation on the target tensor to obtain a hybrid attention calculation result includes: Invoking the target tensor in the first cache variable; In response to the sequence length of the target tensor being greater than a preset threshold, performing the global attention calculation on the elements of the target tensor that satisfy a first preset spacing to obtain a global attention calculation result, and performing the local attention calculation on the elements of the target tensor that satisfy a second preset spacing to obtain a local attention calculation result; Wherein, the first preset spacing is greater than the second preset spacing; Performing a merging process on the global attention calculation result and the local attention calculation result to obtain the hybrid attention calculation result.

3. The operating method according to claim 2, characterized in that The step of invoking the target tensor in the first cache variable and performing hybrid attention calculation on the target tensor to obtain a hybrid attention calculation result further includes: Dividing the target tensor into several blocks, and performing attention calculation on each block by using a scaled dot-product attention mechanism to obtain several block calculation results; Performing a splicing process on the several block calculation results to obtain the hybrid attention calculation result.

4. The operating method according to claim 3, wherein The step of performing attention calculation on each block by using a scaled dot-product attention mechanism to obtain several block calculation results includes: Performing linear transformation on each block with the matrix dimension parameters of query, key, and value in each head, and performing scaled dot-product attention and softmax operation calculations to obtain a feature matrix output by each head; In response to the end of forward propagation, storing the feature matrix in a second cache variable of the tensor cache manager; Invoking the feature matrix in the second cache variable for splicing to obtain the block calculation result corresponding to each block.

5. The operation method according to any one of claims 1 to 4, characterized in that, The running method further includes: In response to the start of forward propagation, determining a target scenario corresponding to the input data by using an aggregation selection function; Obtaining a corresponding target wrapper based on the target scenario and mapping the target scenario into the target wrapper; Wherein, different scenarios corresponding to different wrappers are pre-configured in the deep learning inference engine; And / or The step of writing the hybrid attention calculation result into the output tensor for output includes: Checking the type and dimension of the hybrid attention calculation result; In response to the type and the dimension satisfying a first preset condition, write the mixed attention calculation result into the output tensor for output; In response to the type and / or the dimension not satisfying the first preset condition, reallocate matching memory to form a new output tensor, and write the mixed attention calculation result into the new output tensor for output; and / or, The running method further includes: in response to the end of the execution of the kernel function and the absence of subsequent call dependencies, cancel or postpone the display synchronization call for the execution result of the kernel function; and / or, The running method further includes: checking the data input to the deep learning inference engine and the device attributes of the processor.

6. An operating system for a deep learning inference engine, characterized in that, The deep learning inference engine is carried in a processor, and the running system includes an initialization module, a tensor reshaping module, an attention calculation module, and an output module; The initialization module is used to load a trained inference model, establish a tensor cache manager, and pre-allocate an output tensor; The tensor reshaping module is used to receive input data and obtain an initial tensor corresponding to the input data, perform tensor reshaping on the initial tensor to obtain a target tensor with a target dimension, and store the target tensor in a first cache variable of the tensor cache manager; The attention calculation module is used to call the target tensor in the first cache variable and perform mixed attention calculation on the target tensor to obtain a mixed attention calculation result; The output module is used to write the mixed attention calculation result into the output tensor for output.

7. A deep learning inference platform, characterized in that, The deep learning inference platform includes the running system of the deep learning inference engine according to claim 6.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and adapted to run on the processor, characterized in that, When the processor executes the computer program, it implements the running method of the deep learning inference engine according to any one of claims 1 to 5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the running method of the deep learning inference engine according to any one of claims 1 to 5.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the running method of the deep learning inference engine according to any one of claims 1 to 5.

Citation Information

Cited By

  • End side model reasoning method and device based on RWKV architecture, electronic equipment and storage medium

    CN120725163A

  • Data processing method, processor, chip, display card and electronic equipment

    CN120950263A

  • Data processing method, processor, chip, graphics card and electronic device

    CN120950263B

  • Dynamic reconstruction method of deep learning accelerator and deep learning accelerator system

    CN121480579A

  • Deep learning accelerator dynamic reconfiguration method and deep learning accelerator system

    CN121480579B