Large language model reasoning method and device, electronic equipment and storage medium

By searching and converting historical high-storage-efficiency tensors in large language models, storage resource utilization is optimized, the problem of redundant computing in multi-round dialogue scenarios is solved, and reasoning efficiency and system response speed are improved.

CN120633822APending Publication Date: 2025-09-12TSINGHUA UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510549064.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Large language models have redundant computing problems in multi-round dialogue scenarios, resulting in wasted computing resources and decreased response speed, making it difficult to effectively support high-concurrency user requests.

Method used

By searching and converting historical high-storage-efficiency tensors when receiving the current request, obtaining and temporarily storing the current high-storage-efficiency tensors, and using the mapping table to manage the storage space, the utilization of storage resources is optimized and redundant calculations are reduced.

Benefits of technology

It achieves efficient management and reuse of historical intermediate data, reduces computing resource consumption, improves reasoning efficiency and system response speed, and supports high-concurrency user requests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633822A_ABST
    Figure CN120633822A_ABST
Patent Text Reader

Abstract

The invention provides a large language model reasoning method and device, electronic equipment and a storage medium, and the method comprises the steps: searching a storage space for a historical high-storage-efficiency tensor corresponding to each stored request according to a mapping table under the condition that a current request is received, and converting the historical high-storage-efficiency tensor into a historical target tensor; executing a current round of reasoning through the large language model according to the current request, obtaining an intermediate tensor corresponding to each layer of the large language model in the current round of reasoning process, and obtaining a current target tensor required by attention layer calculation of the large language model in the intermediate tensor; searching a current high-storage-efficiency tensor as a precursor of the current target tensor, storing the current request and the corresponding current high-storage-efficiency tensor in a storage space as historical intermediate data, and updating the mapping table; and inputting the historical target tensor and the current target tensor into the attention layer for calculation to obtain the reasoning text corresponding to the current request, thereby effectively reducing redundant calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a large language model inference method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of artificial intelligence (AI), large language models have demonstrated tremendous potential for application in a wide range of fields. By learning language patterns and semantic information from massive amounts of text data, these models can generate natural and fluent text responses, providing powerful support for applications such as intelligent customer service, automatic translation, and text summarization. However, large language models typically have a massive number of parameters and complex computational structures, which poses numerous challenges in the model's inference process.

[0003] In multi-turn conversation scenarios, to generate coherent and contextually consistent responses, the model needs to process and memorize past conversations. Traditional methods recalculate all the previous conversations during each inference, resulting in significant computational redundancy. Experiments show that this redundant computation can account for 93.1% of the total computation in multi-turn conversations, significantly slowing down inference speed and limiting the system's throughput and real-time performance.

[0004] To avoid this redundant computation, some research has attempted to temporarily store history tensors in GPU memory. However, due to the limited capacity of GPU memory, in addition to storing model weights, it can only store history tensors for a small number of user conversations, making this approach unsuitable for applications that require serving a large number of concurrent users. Furthermore, history tensors compete with ongoing inference tasks for GPU memory resources, further reducing the parallelism and overall efficiency of GPU-based inference.

[0005] Therefore, how to efficiently use the hierarchical storage system to temporarily store historical intermediate data and avoid redundant computing problems during large language model inference has become an urgent problem to be solved in the industry. Summary of the Invention

[0006] The present invention provides a large language model inference method, device, electronic device and storage medium to address the defects of the existing technology in large language model inference in multi-round dialogue scenarios, such as a large amount of redundant calculations, low storage resource utilization efficiency and difficulty in effectively supporting high-concurrency user requests. By temporarily storing and reusing historical high-storage-efficiency tensors, computational redundancy is effectively reduced and inference efficiency is improved.

[0007] The present invention provides a large language model inference method for an inference system, wherein the inference system is loaded with a large language model, and the method comprises: When a current request is received, searching the storage space for historical high-storage-efficiency tensors corresponding to each stored request according to a mapping table, and converting the historical high-storage-efficiency tensors into historical target tensors; the mapping table is used to record the correspondence between each request and the high-storage-efficiency tensors stored in the storage space; Performing a current round of reasoning using the large language model according to the current request, obtaining intermediate tensors corresponding to each layer of the large language model during the current round of reasoning, and obtaining a current target tensor required for attention layer calculation of the large language model from the intermediate tensors; Searching for a current high-storage-efficiency tensor that is a predecessor of the current target tensor, saving the current request and the corresponding current high-storage-efficiency tensor as historical intermediate data to a storage space, and updating the mapping table; The historical target tensor and the current target tensor are input into the attention layer for calculation to obtain the inference text corresponding to the current request.

[0008] According to the large language model inference method provided by the present invention, the historical high storage efficiency tensor is converted into a historical target tensor, specifically including: reading the historical high storage efficiency tensor from the storage space, and allocating temporary storage space in the hardware accelerator for temporary storage; allocating a low-priority computing flow in the hardware accelerator to distinguish it from the computing flow that performs inference; using the low-priority computing flow to convert the historical high storage efficiency tensor into the historical target tensor; and copying the historical target tensor to the cache area of ​​the corresponding user according to the order of historical characters.

[0009] According to the large language model inference method provided by the present invention, the storage space includes main memory and persistent storage; saving the current request and the corresponding current high-storage-efficiency tensor as historical intermediate data to the storage space, and updating the mapping table, specifically includes: Checking whether the remaining storage space in the main memory is less than a threshold; if the storage space in the main memory is greater than or equal to the threshold, saving the current high storage efficiency tensor as historical intermediate data in the main memory, and updating the mapping table; If the remaining storage space in the main memory is less than the threshold, the reuse distance probability distribution of each user is calculated. According to the reuse distance probability distribution of each user and the length of the inference text generated by the previous round of inference, the hit probability of each user accessing the historical high storage efficiency tensor next time is calculated. The historical high storage efficiency tensor with the lowest hit probability of each user is evicted to the persistent storage, the current high storage efficiency tensor is saved as historical intermediate data in the main memory, and the mapping table is updated.

[0010] According to the large language model inference method provided by the present invention, the reuse distance probability distribution of each user is calculated, specifically including: recording the total spatial scale of the historical high-storage-efficiency tensors stored in the storage space accessed by other users between the current request and the last request of each user as the reuse distance; recording the reuse distance corresponding to each request of each user, and obtaining the reuse distance probability distribution of the user's access to the historical high-storage-efficiency tensor.

[0011] According to the large language model inference method provided by the present invention, the hit probability of each user accessing the historical high storage efficiency tensor next time is calculated based on the reuse distance probability distribution of each user and the character length of the inference text generated by the previous round of inference, specifically including: For each user, predicting the reuse distance probability distribution of the user's next visit based on the reuse distance probability distribution of the user's access to the historical high storage efficiency tensor; Obtaining a correlation curve between the character length of the inference text output by the large language model in the current inference process and the predicted lower bound of the reuse distance of the user's next visit; According to the association curve and the character length of the inference text generated in the previous round of inference, the lower bound of the reuse distance of the user's next visit is modified; Online statistics and analysis of users' historical access requests are performed to obtain the main memory hit probability corresponding to different reuse distance ranges under the ideal eviction policy conditions; The hit probability of the user accessing the historical high storage efficiency tensor in the main memory next time is calculated by combining the main memory hit probabilities corresponding to different reuse distance ranges and the revised reuse distance probability distribution of the user's next access.

[0012] According to the large language model inference method provided by the present invention, the current high storage efficiency tensor is saved as historical intermediate data in the main memory, specifically including: adaptively allocating space in the main memory based on the input and output character length of the current request, and writing the current high storage efficiency tensor as historical intermediate data into the allocated space.

[0013] According to the large language model inference method provided by the present invention, the historical target tensor and the current target tensor are input into the attention layer for calculation to obtain the inference text corresponding to the current request, specifically including: when the input is calculated through the attention layer, the temporarily stored historical high storage efficiency tensor is pulled from the temporary space as the input of the attention layer; when the output is calculated through the attention layer, the temporarily stored historical high storage efficiency tensor and the current target tensor obtained by the current round of calculation are pulled from the temporary space as the input of the attention layer to obtain the inference text corresponding to the current request.

[0014] According to the large language model inference method provided by the present invention, the main memory is further provided with a waiting queue and a preparation queue; when multiple current requests are received, the method further includes: Upon receiving a current request, determining the storage location of the historical intermediate data corresponding to the current request, and if the historical intermediate data is located in the main memory, adding the current request to a standby queue; if the historical intermediate data is located in the persistent storage, adding the current request to a waiting queue; When the hardware accelerator has the ability to process a new request, a request is selected from the preparation queue for processing; For a request in the waiting queue, after the historical intermediate data corresponding to the request is loaded from the persistent storage to the main memory, the request is moved to the standby queue.

[0015] The present invention also provides a large language model inference device for use in an inference system, wherein the inference system is loaded with a large language model, and the method includes: a historical target tensor conversion unit, configured to, upon receiving a current request, search the storage space for historical high-storage-efficiency tensors corresponding to each stored request according to a mapping table, and convert the historical high-storage-efficiency tensors into historical target tensors; the mapping table being configured to record the correspondence between each request and the high-storage-efficiency tensors stored in the storage space; a current target tensor acquisition unit, configured to perform a current round of inference using the large language model according to the current request, obtain intermediate tensors corresponding to each layer of the large language model during the current round of inference, and obtain a current target tensor required for calculation of the attention layer of the large language model from the intermediate tensors; a target tensor storage unit, configured to search for a current high-storage-efficiency tensor that is a predecessor of the current target tensor, save the current request and the corresponding current high-storage-efficiency tensor as historical intermediate data to a storage space, and update the mapping table; The inference unit is configured to input the historical target tensor and the current target tensor into the attention layer for calculation, and obtain the inference text corresponding to the current request, including: The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the large language model inference method as described above is implemented.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the large language model inference methods described above.

[0017] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the large language model inference methods described above.

[0018] The large language model inference method and device provided by the present invention achieves efficient management and reuse of historical intermediate data by searching and converting historical high-storage-efficiency tensors according to a mapping table when receiving the current request, and obtaining and temporarily storing the current high-storage-efficiency tensors during the current round of inference: First, the historical high-storage-efficiency tensors in the storage space are quickly located according to the mapping table and converted into historical target tensors, avoiding repeated calculations of historical conversation content; second, when executing the current round of inference, the current target tensor in the intermediate tensors of each layer is obtained, and the current high-storage-efficiency tensors of its predecessor are searched for and temporarily stored to provide data support for subsequent inference; third, the historical target tensor and the current target tensor are input into the attention layer for calculation to obtain the inference text corresponding to the current request. By temporarily storing and reusing historical high-storage-efficiency tensors, redundant calculations are effectively reduced, the consumption of storage resources is reduced, and the inference efficiency is improved.

[0019] In addition, by rationally managing historical intermediate data in the storage space, fast access and efficient use of data are ensured, thereby significantly improving the performance and response speed of the large language model inference system while ensuring the accuracy of reasoning. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 It is a flowchart of the large language model inference method provided by the present invention.

[0022] Figure 2 It is a schematic diagram of the structure of the large language model provided by the present invention.

[0023] Figure 3 This is a schematic diagram of the distribution of historical intermediate data in the storage space provided by the present invention.

[0024] Figure 4 It is a schematic diagram of converting a historical high storage efficiency tensor into a historical target tensor according to the present invention.

[0025] Figure 5 Schematic diagram of the eviction strategy of the historical high storage efficiency tensor in the main memory provided by the present invention.

[0026] Figure 6 This is a schematic diagram of hierarchical storage of historical data provided by the present invention.

[0027] Figure 7 It is a schematic diagram of the request scheduling based on dual queues provided by the present invention.

[0028] Figure 8 It is a schematic diagram of the architecture of the model answer generation service platform provided by the present invention.

[0029] Figure 9 It is a structural diagram of the large language model inference device provided by the present invention.

[0030] Figure 10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0031] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0032] First, the terms involved in the embodiments of the present invention are explained accordingly.

[0033] High-storage-efficiency tensors: These are highly efficient among the tensors generated during large language model inference. They can be temporarily stored in a small amount of storage space while ensuring information availability. They are typically the precursors of the target tensor and can be converted to the target tensor through subsequent computations.

[0034] Intermediate tensors: Temporary tensors generated by each layer of a large language model during inference. These tensors are intermediate results of the model calculation process and are used in subsequent calculation steps.

[0035] Target tensor: The input tensor required for the attention layer calculation of the large language model. During inference, the target tensor is formed from specific data in the intermediate tensor and is used for the model's attention mechanism calculation.

[0036] Mapping table: A table that records the correspondence between each request and the high-storage-efficiency tensors stored in the storage space. This mapping table allows you to quickly locate the location of the historical high-storage-efficiency tensors corresponding to a specific request in the storage space.

[0037] Attention layer: This is the layer where the attention mechanism in large language models resides. The attention layer calculates the relationship between input tensors and assigns different weights to different input elements, allowing the model to focus on important parts when processing information.

[0038] Main memory: The main storage in a computer system, typically high-speed memory such as DRAM. Main memory has fast read and write speeds and is used to temporarily store data that is currently in use or frequently accessed.

[0039] Persistent storage: Secondary memory in a computer system, typically a large-capacity memory device such as a hard drive or solid-state drive. Persistent storage has a large storage capacity and retains data even after a power outage. It is used to store infrequently used data for a long period of time.

[0040] Hardware accelerator: A hardware device used to accelerate specific computing tasks, such as a GPU or TPU. Hardware accelerators can process large amounts of data in parallel, improving computing efficiency.

[0041] Reuse distance: This refers to the total size of the historical high-storage-efficiency tensors stored in the storage space accessed by other users between the current user's request and the previous request. This distance reflects the interval at which data is reused.

[0042] Reuse distance probability distribution: Describes the probability distribution of the reuse distance of a user accessing a historically high-storage-efficiency tensor. The reuse distance probability distribution reflects the likelihood of a user accessing data within different reuse distance ranges.

[0043] In today's digital age, artificial intelligence (AI) technology, particularly large language models, plays a vital role in numerous fields. Take intelligent customer service systems, for example. These systems respond to complex user queries in real time, gain a deep understanding of user needs through multiple rounds of conversation, and provide precise solutions. In practice, these systems must efficiently handle a large number of concurrent requests and quickly generate coherent and accurate responses. For example, on e-commerce platforms, users may inquire about product details, logistics progress, after-sales service, and other issues. These systems must maintain contextual consistency across multiple rounds of conversation, integrating historical conversation information to generate relevant and accurate responses. However, traditional reasoning methods waste computing resources and slow response times due to repeated computation of historical conversation content. This not only impacts user experience but also limits the system's scalability.

[0044] To address these challenges, the inference method proposed in this example specifically addresses the efficient inference of large language models in multi-turn conversation scenarios. This method aims to reduce redundant computation, improve inference efficiency, and ensure rapid system response even under high-concurrency requests. Its core focus is on optimizing the management and utilization of historical intermediate data, achieving efficient use of storage resources through intelligent data scheduling between main memory and persistent storage.

[0045] The hardware environment used in the reasoning method proposed in this embodiment includes multiple key components. First, the system is equipped with a high-performance processor for executing the main process of the reasoning task. Second, the system contains a large-capacity main memory for temporarily storing currently active intermediate calculation results and model weights to ensure fast data access. In addition, the system is also equipped with persistent storage devices, such as solid-state drives, for long-term storage of historical intermediate data so that it can be loaded into the main memory when needed. Hardware accelerators, such as GPUs or TPUs, are an important part of the system and can process large amounts of data in parallel, significantly improving computing efficiency. Finally, the system connects various components through a high-speed communication bus to ensure efficient and low-latency data transmission. In a multi-node deployment scenario, each node is interconnected through a high-speed network to support distributed reasoning tasks, further improving the scalability and throughput of the system.

[0046] The following combination Figures 1-8 The large language model inference method of an embodiment of the present invention is described to analyze how to combine mapping tables, hierarchical storage, and intelligent scheduling mechanisms to achieve efficient inference that is different from existing technologies.

[0047] The large language model inference method of the embodiment of the present invention is used in an inference system. In a typical inference system structure, its hardware includes: Processor (CPU): The core of the system, responsible for executing the main process and control logic of the inference task. The processor handles operations such as request scheduling and mapping table management.

[0048] Main memory (DRAM): High-speed memory used to temporarily store currently active intermediate computation results and model weights. Main memory supports fast data access, ensuring low latency during inference.

[0049] Persistent storage (SSD / HDD): Large-capacity storage devices used for long-term storage of historical intermediate data. They provide additional storage space when main memory is insufficient.

[0050] Hardware accelerators (GPU / TPU): Parallel computing units used to accelerate matrix and tensor operations in large language models. Hardware accelerators can significantly improve computing efficiency, especially when processing large amounts of data.

[0051] High-speed communication bus: Connects various hardware components, ensuring efficient and low-latency data transmission. The high-speed communication bus supports the rapid flow of data between main memory, persistent storage, and hardware accelerators.

[0052] Distributed nodes: In high-concurrency scenarios, the system can be expanded to multiple nodes, interconnected by a high-speed network. Distributed nodes support distributed reasoning tasks, further improving system throughput and scalability.

[0053] For software components, the reasoning system includes: Inference service interface: Receives user requests and performs preliminary processing. The inference service interface verifies the request format and parameters and prepares the input data required for inference.

[0054] Scheduling module: Responsible for the allocation and management of requests. The scheduling module assigns requests to the standby queue or waiting queue based on the location of the historical intermediate data corresponding to the request.

[0055] Mapping Table Management Module: Maintains the correspondence between requests and highly efficient tensors in storage. This module can quickly locate historical intermediate data and support efficient read and write operations.

[0056] Inference Computation Module: This module performs inference computations on large language models. It captures intermediate tensors at each layer and obtains the current target tensor required by the attention layer.

[0057] Storage Management Module: Manages the use of main memory and persistent storage. The storage management module determines the storage location and eviction policy of data based on the main memory space status and hit probability.

[0058] Data conversion module: Converts historical high-storage-efficiency tensors into historical target tensors. The data conversion module uses a low-priority computation flow to avoid interfering with regular inference computations.

[0059] In this embodiment of the present invention, the workflow of the reasoning system mainly includes: (1) Receiving requests: The inference service interface receives user requests and performs preliminary processing and verification.

[0060] (2) Request scheduling: The scheduling module assigns requests to the preparation queue or waiting queue based on the location of the historical intermediate data corresponding to the request.

[0061] (3) Data reading and conversion: According to the mapping table, the historical high storage efficiency tensor is searched in the main memory or persistent storage. The data conversion module converts the historical high storage efficiency tensor into the historical target tensor.

[0062] (4) Inference calculation: The inference calculation module uses the large language model to perform the inference calculation of the current request. It captures the intermediate tensors of each layer and obtains the current target tensor.

[0063] (5) Data temporary storage: Find the predecessor tensor of the current target tensor (the current high-storage-efficiency tensor). Save the current high-storage-efficiency tensor to main memory or persistent storage, and update the mapping table.

[0064] (6) Attention calculation: The historical target tensor and the current target tensor are input into the attention layer for calculation, and the inference text corresponding to the current request is generated and returned to the user as the output of the model.

[0065] Through the collaborative work of these hardware and software components, this inference system achieves efficient inference in multi-round conversation scenarios, reduces redundant computation, optimizes storage resource utilization, and supports high-concurrency user requests. The following details the specific steps of the inference method in this embodiment and demonstrates how it achieves these optimization effects in practice.

[0066] like Figure 1 As shown, the large language model inference method of the embodiment of the present invention includes the following: Step 101: When a current request is received, the historical high storage efficiency tensors corresponding to the stored requests are searched in the storage space according to the mapping table, and the historical high storage efficiency tensors are converted into historical target tensors.

[0067] In step 101, upon receiving a request, the system first receives and preliminarily processes the user request through the inference service interface. The request carries user input information, such as a text query. The system needs to generate an accurate and coherent response text for the request.

[0068] The system searches the storage space for the historical high-storage-efficiency tensor corresponding to the current request based on the mapping table. The mapping table records the correspondence between requests and high-storage-efficiency tensors, including information such as the request identifier and storage location. The system uses the request identifier to quickly locate the corresponding entry in the mapping table and obtain the storage location.

[0069] The storage location may be in main memory or persistent storage. Main memory provides fast data access and is suitable for temporarily storing currently active intermediate computation results; persistent storage is used for long-term storage of historical intermediate data. The system determines the data loading source based on this.

[0070] like Figure 2 As shown, a large language model consists of multiple sequentially connected decoders, each of which includes an attention module and a feedforward network module. During inference, each decoder sequentially generates multiple intermediate tensors, including a target tensor (such as a key-value pair), which serves as the input for the attention module. Existing methods typically directly store the target tensor, but this approach ignores storage and transmission efficiency issues.

[0071] The embodiment of the present invention innovatively chooses to temporarily store high-storage-efficiency tensors instead of directly storing target tensors. High-storage-efficiency tensors are tensors that require the least space per unit of computation among all intermediate tensors. In this way, the method of the embodiment of the present invention significantly reduces the space overhead and transmission overhead of temporarily storing historical tensors. Specifically, temporarily storing high-storage-efficiency tensors reduces the space overhead by half while keeping the computation overhead within an acceptable range.

[0072] This strategy achieves two key optimization goals: first, it reduces the main memory space requirements, allowing the main memory to accommodate more users' historical tensors and improving the system's concurrent processing capabilities; second, it reduces the amount of transmission when reading data from persistent storage, thereby reducing the latency of data loading.

[0073] After finding a historical tensor with high storage efficiency, the system converts it into a historical target tensor. This conversion process is typically performed on a hardware accelerator (such as a GPU or TPU) to improve computational efficiency. This conversion process may include operations such as data format conversion and dimensionality adjustment to ensure that the historical target tensor is compatible with the input requirements of the attention layer. The converted historical target tensor serves as one of the inputs to the attention layer, providing historical context for subsequent inference computations. This process enables efficient reuse of historical intermediate data, reduces recalculation, and improves inference efficiency.

[0074] Step 102: Perform a current round of reasoning through the large language model according to the current request, obtain the intermediate tensors corresponding to each layer of the large language model during the current round of reasoning, and obtain the current target tensor required for the attention layer calculation of the large language model from the intermediate tensors.

[0075] In step 102, the system performs inference calculations on the large language model based on the current request. The main challenges faced in this process include ensuring the efficiency and accuracy of inference and the ability to process large-scale data.

[0076] The system needs to efficiently process requests and quickly generate accurate responses, especially in multi-turn conversation scenarios where conversational coherence must be maintained. The model parameters and computational structure are complex, requiring high computing resources. This requires algorithm optimization and the use of hardware accelerators to improve efficiency. Furthermore, effective management of intermediate tensors is required to avoid out-of-memory issues and ensure smooth inference.

[0077] To this end, the system addresses these challenges through the following methods: efficient inference computing, capturing intermediate tensors, locating target tensors, and extracting target tensors. Efficient inference computing utilizes hardware accelerators to process data in parallel, improving computational efficiency; capturing intermediate tensors obtains the computational results of each layer and extracts the data required for the attention layer; locating the target tensor identifies key tensors to generate accurate responses; and extracting the target tensor ensures that the data format and content meet the computational requirements of the attention layer.

[0078] Specifically, the system uses the inference computation module to perform the inference task for the current request, sequentially computing intermediate tensors at each layer and storing them in main memory for use in subsequent steps. Based on the requirements of the attention layer, the system identifies and extracts the current target tensor, ensuring it contains the information necessary to generate an accurate response. After extraction and conversion, the current target tensor is fed into the attention layer's computation to generate the inference text for the current request.

[0079] Through these measures, the system effectively addresses the challenges in step 102, ensures efficient and accurate reasoning calculations, properly manages intermediate tensors, and supports rapid responses in multi-round dialogue scenarios.

[0080] Step 103: Search for a current high-storage-efficiency tensor that is a predecessor of the current target tensor, save the current request and the corresponding current high-storage-efficiency tensor as historical intermediate data to a storage space, and update the mapping table.

[0081] In step 103, the system first searches for a currently high-storage-efficiency tensor that serves as a predecessor to the current target tensor. To achieve this, the system analyzes the intermediate tensors generated by each layer of the model, identifying those that can be temporarily stored with minimal space cost while effectively supporting subsequent computations. This process involves an in-depth analysis of the characteristics of these intermediate tensors, assessing their storage efficiency and computational value.

[0082] After confirming the current high-storage-efficiency tensor, the system combines the current request with this tensor and saves it in the main memory as historical intermediate data.

[0083] At the same time, the system updates the mapping table, accurately recording the correspondence between the current request and the newly saved, highly storage-efficient tensor. This step is crucial for rapidly locating and retrieving the required data during subsequent inference. Updating the mapping table involves more than simply recording data; it also maintains data consistency and integrity, ensuring that multiple requests can accurately find the corresponding intermediate data in highly concurrent scenarios.

[0084] This process not only optimizes storage resource utilization but also lays a solid data foundation for subsequent inference calculations. By temporarily storing the current high-storage-efficiency tensors, the system can quickly recover and utilize these critical intermediate calculation results when processing future requests, thereby improving overall inference efficiency and system responsiveness. Furthermore, this optimization strategy reduces data loading latency and increases system throughput, enabling the system to better handle the complex demands of multi-round conversation scenarios.

[0085] Step 104: Input the historical target tensor and the current target tensor into the attention layer for calculation to obtain the inference text corresponding to the current request.

[0086] In step 104, the system inputs the historical target tensor and the current target tensor into the attention layer for computation to generate the inference text corresponding to the current request. Specifically, the historical target tensor provides contextual information about the previous conversation, while the current target tensor contains relevant data for the current input. The attention layer uses these two tensors as input to generate a coherent and accurate response based on both the historical context and the user's current question.

[0087] Combine Figure 3 As shown in the figure, when processing a request, the system first reads historically efficient tensors from storage. These tensors may be stored in main memory or persistent storage, and the system uses a mapping table to quickly locate and read the required data. The read historically efficient tensors are then transferred to a hardware accelerator (such as a GPU or TPU). During this process, the system allocates dedicated temporary storage space in the hardware accelerator to store these tensors, ensuring fast access and processing of the data during the conversion process.

[0088] like Figure 3 As shown in the figure, traditional fixed-block space allocation methods result in significant space waste, limiting the parallelism of inference computations and severely impacting the efficiency of large language model inference systems. Small-block space allocation methods, on the other hand, require massive amounts of small-block data transfer from main memory to the hardware accelerator, incurring additional transmission overhead.

[0089] The present invention allocates storage space in the hardware accelerator based on a combination of large blocks and small blocks. Large blocks are allocated for fixed-size input character parts, and small blocks are allocated for variable-size output character parts. This reduces the waste of video memory space while improving the efficiency of transferring high-storage-efficiency tensors from main memory to the hardware accelerator.

[0090] After the system of the present invention reads the historical high storage efficiency tensor read from the storage space into the hardware accelerator, a low priority computing flow is allocated in the hardware accelerator to distinguish it from the computing flow that performs inference, and the low priority computing flow is used to convert the historical high storage efficiency tensor into a historical target tensor.

[0091] During this process, the system of the present invention needs to prune the inherent weights of the model, such as Figure 4 As shown, the calculation only generates the required target tensor, avoiding the additional generation of unnecessary Q tensors. It should be explained that in the attention mechanism, the Q tensor is used to represent the query vector of the current input, the K tensor is used to represent the key vector, and the V tensor is used to represent the value vector. In some cases, some parts of the Q tensor may contain redundant information, or the complete Q tensor is not required for calculation at a specific calculation stage. In this embodiment, when converting the historical high storage efficiency tensor to the target tensor, only the K and V tensors are required for subsequent attention calculations, and the complete Q tensor is not required. By pruning off the unnecessary part of the Q tensor, the computational overhead is reduced by 1 / 3, the storage requirements and the amount of computation are reduced, and the computational efficiency is improved.

[0092] The computational process of converting a highly storage-efficient tensor into a target tensor can be parallelized with the normal inference process of other requests. This is because a significant amount of redundant computation (approximately 5 / 6) has already been saved, a significant portion of the streaming multiprocessors in the hardware accelerator are typically idle, and in most cases, inference computation is limited by the hardware accelerator's memory space rather than its compute units. Therefore, additional low-priority compute streams are required to perform the target tensor conversion computation to better utilize idle compute units.

[0093] The system of the present invention copies the historical target tensor into the corresponding user's cache area according to the order of historical characters. This step prepares data for subsequent attention calculations and ensures that historical context information can be effectively integrated into the current reasoning process.

[0094] At the same time, the system of the present invention combines the current target tensor (i.e., the tensor obtained during the current round of inference and required for the attention layer's computation) with the historical target tensor, which serve as input to the attention layer. Within the attention layer, the system performs specific computations, such as weighted summation, to comprehensively consider both the historical context and the current input. This allows the large language model to focus on the historical information most relevant to the current request and, in combination with the current input, generate the most appropriate response.

[0095] Finally, the system uses activation functions such as Softmax to convert the output of the attention layer into a probability distribution of each word in the vocabulary. Based on this probability distribution, the system selects the most likely word sequence to form the final inference text. This process not only demonstrates the model's ability to comprehensively process historical conversations and current input, but also demonstrates its ability to generate high-quality responses, ensuring the coherence and accuracy of multiple rounds of dialogue.

[0096] The large language model inference method provided by an embodiment of the present invention achieves efficient management and reuse of historical intermediate data by searching and converting historical high-storage-efficiency tensors according to a mapping table when receiving the current request, and obtaining and temporarily storing the current high-storage-efficiency tensor during the current round of inference. First, the historical high-storage-efficiency tensor in the storage space is quickly located according to the mapping table and converted into a historical target tensor, avoiding repeated calculation of historical conversation content. Second, when executing the current round of inference, the current target tensor in the intermediate tensors of each layer is obtained, and the current high-storage-efficiency tensor of its predecessor is searched for and temporarily stored to provide data support for subsequent inference. Third, the historical target tensor and the current target tensor are input into the attention layer for calculation to obtain the inference text corresponding to the current request. This process effectively reduces redundant calculations, reduces the consumption of storage resources, and improves inference efficiency.

[0097] Furthermore, the storage space in this embodiment consists of main memory and persistent storage. Main memory provides fast data access and is suitable for temporarily storing currently active intermediate computation results, while persistent storage provides greater storage capacity. After the system completes processing the current request and generates the current high-storage-efficiency tensor, this data needs to be saved as historical intermediate data to the storage space and the mapping table needs to be updated.

[0098] When saving the current request and the corresponding current high-storage-efficiency tensor as historical intermediate data to storage, the system first checks whether the remaining free space in main memory reaches a set threshold. This threshold is pre-set based on system performance and storage requirements to ensure that there is always a portion of free main memory available for quickly writing the high-storage-efficiency tensor for the current inference request. If there is sufficient main memory space (i.e., the remaining free space is greater than or equal to the threshold), the system saves the current high-storage-efficiency tensor directly to main memory. Specifically, the system adaptively allocates space in main memory based on the input and output character lengths of the current request and writes the current high-storage-efficiency tensor as historical intermediate data to the allocated space. This ensures that subsequent requests can quickly access this data, improving system responsiveness. Simultaneously, the system updates a mapping table to record the main memory locations of the current request and the newly saved high-storage-efficiency tensor, facilitating rapid subsequent location and access.

[0099] If main memory space is insufficient (i.e., remaining space is less than a threshold), the system initiates a data management strategy to optimize storage resource usage. At this point, the system calculates the probability distribution of reuse distance for each user. Reuse distance is defined as the total size of the historically efficient tensors stored in the storage space accessed by other users between the current request and the previous request. By analyzing user access patterns, the system can derive a probability distribution of the reuse distances of users accessing historically efficient tensors. This distribution reflects the likelihood that different users will reuse their historically efficient tensors in future requests.

[0100] Run a thread in the background that periodically scans the amount of free memory, such as Figure 5 As shown in the figure, when the free space in main memory is found to be less than 5%, an eviction operation is initiated. First, an association curve is generated based on the character length of the inference text output by the large language model during the current inference process and the predicted lower bound of the reuse distance of the user's next access. Then, based on the association curve and the character length of the inference text generated by the previous inference round, the system further calculates the hit probability of each user's next access to their historical high-storage-efficiency tensors. The calculation of the hit probability comprehensively considers the user's access pattern and the frequency of use of historical data. By analyzing this data, the system can predict which historical high-storage-efficiency tensors are least likely to be accessed in future requests.

[0101] Specifically, the lower bound of the reuse distance for the user's next visit is adjusted based on the association curve and the character length of the inference text generated during the previous inference round. The user's historical access requests are then statistically analyzed online to determine the main memory hit probability for different reuse distance ranges under an ideal eviction policy. Combining the main memory hit probabilities for different reuse distance ranges with the revised reuse distance probability distribution for the user's next visit, the main memory hit probability for the user's next access history high-storage-efficiency tensor is calculated.

[0102] The historical high storage efficiency tensor with the lowest hit probability will be evicted to persistent storage. This decision is based on the following logic: if the hit probability of a historical high storage efficiency tensor is very low, then it is less likely to be accessed in future requests, and moving it to persistent storage can free up main memory space for data that is more likely to be accessed. The current high storage efficiency tensor is then saved to the main memory, and the mapping table is updated to reflect the new storage status. After the eviction is completed, it is determined whether the free main memory is still less than 5% at this time. If it is still less than 5%, the historical data with the lowest theoretical hit probability will continue to be selected for eviction, otherwise the eviction process is completed. The eviction strategy of the embodiment of the present invention ensures that the data that is most likely to be accessed is always retained in the main memory, improves the utilization efficiency of storage resources, and at the same time ensures the security and recoverability of data through persistent storage.

[0103] When performing data eviction operations, this invention employs an intelligent and adaptive strategy designed to optimize storage resource utilization and ensure unimpeded system performance. First, the system adaptively identifies the hit rate for different ranges of reuse distances. This is because the hit rate is not fixed but varies dynamically with the size of main memory and the current workload pressure.

[0104] To achieve more accurate predictions, the system maintains a shadow cache using an optimal eviction policy. Specifically, the system uses the load data from the past five minutes to maintain this shadow cache. By counting the hit rates of requests with different reuse distance ranges, the reuse distance is divided into multiple intervals: a small bucket covering all reuse distances smaller than the main memory size, m medium buckets covering medium reuse distances, and a large bucket covering all extremely large reuse distances. For each medium bucket, the system calculates its hit rate under the optimal policy.

[0105] At the same time, the system continuously maintains the reuse distance distribution of each user's historical data. This allows the system to predict the probability that the reuse distance of the user's next request will fall into each bucket, denoted as prob_small, prob_promising(i) (i ∈ [1, m]), and prob_extreme, respectively. In addition, the system further modifies these predicted probabilities based on the length of the text output by the previous round of model.

[0106] Based on this analysis, the system assigns the hit rate under the optimal strategy as hit potential to different reuse distance buckets and calculates the overall hit potential of each user's historical data. Ultimately, the system evicts the historical data with the lowest overall potential, moving it from main memory to persistent storage. This process not only ensures that the most likely data is always retained in main memory, but also maximizes storage resource utilization through intelligent prediction and dynamic adjustment, while ensuring high system performance.

[0107] Through the above steps, the system can efficiently store and manage historical intermediate data even with limited storage resources, ensuring the efficiency and continuity of the inference process. This not only optimizes the use of storage resources but also ensures the rapid processing of subsequent requests.

[0108] like Figure 6 As shown in the figure, assume that the system is processing requests from multiple users, including User 1, User 2, and User 3. Each user has corresponding historical intermediate data stored in the hierarchical storage system. After receiving User 3's request, the system first searches the hierarchical storage for User 3's corresponding historical data. Hierarchical storage can simultaneously store the historical data of multiple users, such as User 1 and User 2. This data may be located in main memory or persistent storage. If the historical data is in main memory, it can be accessed directly and quickly; if it is in persistent storage, it must be loaded before use.

[0109] The system generates the current high-storage-efficiency tensor based on the current request and needs to save it as historical intermediate data. At this point, the system checks whether the remaining space in main memory has reached the set threshold. If there is sufficient main memory space, the system saves the current high-storage-efficiency tensor to main memory and updates the mapping table to record its location for subsequent fast access.

[0110] If main memory space is insufficient, the system calculates the hit probability of each user's historical data based on the user's reuse distance probability distribution and the length of the previous inference text. Data with a low hit probability is less likely to be accessed in the future. The system evicts this low-hit-probability data from main memory to persistent storage to make room for the new current high-storage-efficiency tensor. This ensures that only the most likely data to be accessed is retained in main memory, improving storage efficiency. After saving the current high-storage-efficiency tensor to main memory, the system updates the mapping table to reflect the new storage state. This allows subsequent requests to quickly locate and read the required data, improving overall system performance.

[0111] Through the above process, the system achieves efficient management of storage resources, ensuring rapid response to each user's request in multi-user scenarios, while optimizing storage resource utilization and reducing data loading latency. This process fully demonstrates the advantages of a tiered storage system, which achieves efficient data management and fast inference response by intelligently scheduling data between main memory and persistent storage.

[0112] In step 104, the system inputs the historical target tensor and the current target tensor into the attention layer for calculation to generate the inference text corresponding to the current request. Specifically: When the attention layer processes input information, the system reads previously saved historical high-memory tensors from a temporary storage space. This temporary space can be a cache in the hardware accelerator or other designated temporary storage area. The historical target tensor, which is a tensor converted from the historical high-memory tensor, provides contextual information about the historical conversation. By using the historical target tensor as part of the input, the attention layer can understand the content of the previous conversation, providing a basis for generating coherent responses.

[0113] When the attention layer generates output, the computation combines historical context with the current input. The system not only reads historical, highly memory-efficient tensors from the temporary storage space, but also uses the current target tensor computed during the current round of inference. The current target tensor contains key information about the current input, allowing the attention layer to focus on the key aspects of the current request. The attention layer combines the information from these two tensors to calculate the most relevant output, generating accurate and coherent reasoning text.

[0114] This computational approach, combining historical and current information, enables the model to maintain contextual coherence across multiple rounds of conversation while efficiently responding to the user's current request. In this way, step 104 not only generates inference text but also ensures conversational coherence and efficient model responses.

[0115] In an embodiment, Figure 7 As shown in the figure, to improve the system's processing efficiency for multiple requests, a waiting queue and a standby queue are set up in main memory. When the system receives multiple requests, the initial queue for each request is determined based on the storage location of the corresponding historical intermediate data. Specifically, if the historical intermediate data is located in main memory, the request is added to the standby queue; if it is located in persistent storage, the request is added to the waiting queue.

[0116] Requests in the standby queue can be quickly processed by the hardware accelerator because the data is in the main memory. When the hardware accelerator is idle, requests from the standby queue will be selected for processing first. Requests in the waiting queue need to load data from persistent storage, so the system will move them to the standby queue after the data is loaded into the main memory. To ensure fairness and improve efficiency, the standby queue adopts a first-come-first-served policy based on the order in which requests arrive, while the waiting queue adopts a maximum response ratio limited policy based on the request-response ratio. Through the dual queue mechanism of the waiting queue and the standby queue of this embodiment, it can be ensured that only requests from the standby queue can be scheduled to the hardware accelerator for inference, and requests from the waiting queue can only be promoted to the standby queue after their corresponding historical data are loaded from persistent storage to the main memory. In this way, when a request needs to load its historical data from persistent storage, subsequent requests whose historical data resides in the main memory will no longer be blocked, thereby minimizing overall latency.

[0117] This queue management mechanism enables the system to efficiently handle multiple requests. It not only ensures efficient data loading but also optimizes overall performance through a rational scheduling strategy. This process fully demonstrates the system's optimized design for multi-request scenarios, effectively improving inference efficiency and resource utilization.

[0118] Taking a usage scenario as an example, the intelligent customer service system receives concurrent requests from multiple users during peak hours. Each request involves multiple rounds of conversations and requires access to historical intermediate data to generate coherent responses.

[0119] The system is equipped with 200GB of main memory and 2TB of persistent storage. The main memory is used to temporarily store currently active intermediate calculation results and model weights, while the persistent storage is used to store historical intermediate data over the long term. The system is also equipped with a high-performance GPU as a hardware accelerator to accelerate inference calculations.

[0120] Initial state: There is enough space in the main memory to store the historical intermediate data of the new request.

[0121] The implementation steps are as follows: 1) Request reception and queue allocation.

[0122] User 1 sends a request. After receiving the request, the system checks and finds that user 1's historical intermediate data is stored in the main memory, so it adds the request to the standby queue.

[0123] User 2 sends a request. The system checks and finds that user 2's historical intermediate data is stored in persistent storage, so it adds the request to the waiting queue.

[0124] User 3 sends a request. The system checks and finds that user 3's historical intermediate data is stored in the main memory, so it adds the request to the standby queue.

[0125] 2) The hardware accelerator processes the request.

[0126] When the hardware accelerator completes its current task, the system selects a request from the standby queue (e.g., User 1's request) for processing. The hardware accelerator reads User 1's historical intermediate data from main memory and begins performing inference calculations.

[0127] 3) Waiting for requests in the queue to be processed.

[0128] The system starts loading user 2's historical intermediate data from persistent storage into main memory. During the data loading process, user 2's request remains in the waiting queue.

[0129] 4) Data loading is completed and the queue is updated.

[0130] Once user 2's historical intermediate data is completely loaded into the main memory, the system moves user 2's request from the waiting queue to the standby queue.

[0131] 5) Subsequent request processing.

[0132] When the hardware accelerator is idle again, the system selects another request (for example, the request from user 3) from the standby queue for processing. The hardware accelerator reads the historical intermediate data of user 3 from the main memory and begins performing inference calculations.

[0133] 6) Handling when main memory space is insufficient.

[0134] As more requests arrive, main memory space gradually decreases. When main memory space falls below a set threshold (for example, 5%), the system initiates a data eviction policy. The system calculates the hit probability of each user's historical intermediate data and evicts the data with the lowest hit probability to persistent storage, freeing up space for new requests.

[0135] 7) Continuous request scheduling.

[0136] The system continuously receives new requests and assigns them to a standby queue or a waiting queue based on the storage location of historical intermediate data. The hardware accelerator processes requests based on queue priority, ensuring efficient use of computing resources.

[0137] The present invention can provide a unified platform and service interface. Figure 8 As shown, the present invention supports distributed deployment of programs. Each distributed node uses a scheduler to parallelize the intermediate tensor matrix calculation units. After completing the matrix calculation task for an intermediate tensor, the intermediate tensor matrix calculation unit returns the result to the scheduler. The scheduler aggregates these results and returns them to the upper layer, achieving efficient completion of the overall task.

[0138] After receiving requests from multiple users, the system distributes them to different queues based on historical data storage locations. Through a unified platform and service interface, the system supports distributed program deployment. Within this architecture, multiple distributed nodes work together, each equipped with a scheduler to parallelize intermediate tensor matrix computation units.

[0139] The system receives concurrent requests from multiple users. Each request may require temporarily stored historical intermediate data, which may be stored in the main memory of a distributed node or in persistent storage. The system's goal is to process these requests efficiently, avoid blocking, and fully utilize hardware resources.

[0140] When the hardware accelerator is idle, the system selects requests from the standby queue for processing. Since the historical data corresponding to these requests is already in main memory, it can be quickly loaded into the hardware accelerator for inference execution. For requests in the waiting queue, the system loads the corresponding historical data from persistent storage into main memory. Once loaded, the request is moved to the standby queue for processing. When main memory space is insufficient, the system evicts some historical data to persistent storage based on the hit probability, ensuring efficient main memory utilization.

[0141] The method of this embodiment of the present invention distributes multiple inference requests to waiting queues and standby queues based on the location of their historical intermediate data in the storage hierarchy, preventing subsequent requests from being blocked by requests that need to read historical data from persistent storage. This method can significantly improve the overall throughput and reduce latency of large language model inference systems, efficiently utilizing hierarchical storage to avoid massive redundant computations.

[0142] The following describes a large language model inference device provided by an embodiment of the present invention. The large language model inference device described below and the large language model inference method described above can be referenced to each other.

[0143] The large language model inference device provided by the embodiment of the present invention is Figure 9 ,include: The historical target tensor conversion unit 901 is configured to, upon receiving a current request, search the storage space for historical high-storage-efficiency tensors corresponding to each stored request according to a mapping table, and convert the historical high-storage-efficiency tensors into historical target tensors; the mapping table is configured to record the correspondence between each request and the high-storage-efficiency tensors stored in the storage space; A current target tensor acquisition unit 902 is configured to perform a current round of inference using the large language model according to the current request, obtain intermediate tensors corresponding to each layer of the large language model during the current round of inference, and obtain a current target tensor required for the attention layer calculation of the large language model from the intermediate tensors; a target tensor storage unit 903 configured to search for a current high-storage-efficiency tensor that is a predecessor of the current target tensor, save the current request and the corresponding current high-storage-efficiency tensor as historical intermediate data to a storage space, and update the mapping table; The reasoning unit 904 is used to input the historical target tensor and the current target tensor into the attention layer for calculation to obtain the reasoning text corresponding to the current request.

[0144] The large language model inference device provided by an embodiment of the present invention achieves efficient management and reuse of historical intermediate data by searching and converting historical high-storage-efficiency tensors according to a mapping table when receiving the current request, and obtaining and temporarily storing the current high-storage-efficiency tensor during the current round of inference. First, the historical high-storage-efficiency tensor in the storage space is quickly located according to the mapping table and converted into a historical target tensor, avoiding repeated calculation of historical conversation content. Second, when executing the current round of inference, the current target tensor is obtained from the intermediate tensors of each layer, and the current high-storage-efficiency tensor of its predecessor is searched for and temporarily stored to provide data support for subsequent inference. Third, the historical target tensor and the current target tensor are input into the attention layer for calculation to obtain the inference text corresponding to the current request. This process effectively reduces redundant calculations, reduces the consumption of storage resources, and improves inference efficiency.

[0145] Figure 10 An example of a physical structure diagram of an electronic device is shown below. Figure 10As shown, the electronic device may include: a processor (processor) 1010 , a communication interface (Communications Interface) 1020 , a memory (memory) 1030 and a communication bus 1040 , wherein the processor 1010 , the communication interface 1020 , and the memory 1030 communicate with each other via the communication bus 1040 . The processor 1010 can call the logic instructions in the memory 1030 to execute the large language model inference method, which includes: when receiving the current request, searching the storage space for historical high storage efficiency tensors corresponding to each stored request according to the mapping table, and converting the historical high storage efficiency tensors into historical target tensors; the mapping table is used to record the correspondence between each request and the high storage efficiency tensors stored in the storage space; performing the current round of inference through the large language model according to the current request, obtaining the intermediate tensors corresponding to each layer of the large language model in the current round of inference, and obtaining the current target tensor required for the attention layer calculation of the large language model from the intermediate tensors; searching for the current high storage efficiency tensor that is a predecessor of the current target tensor, saving the current request and the corresponding current high storage efficiency tensor as historical intermediate data to the storage space, and updating the mapping table; inputting the historical target tensor and the current target tensor into the attention layer for calculation to obtain the inference text corresponding to the current request.

[0146] Furthermore, the logic instructions in the aforementioned memory 1030 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0147] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the large language model inference method provided by the above methods, the method including: when receiving a current request, searching the storage space for historical high storage efficiency tensors corresponding to each stored request according to a mapping table, and converting the historical high storage efficiency tensors into historical target tensors; the mapping table is used to record the correspondence between each request and the high storage efficiency tensors stored in the storage space; performing the current round of inference through the large language model according to the current request, obtaining the intermediate tensors corresponding to each layer of the large language model in the current round of inference process, and obtaining the current target tensor required for the attention layer calculation of the large language model from the intermediate tensor; searching for the current high storage efficiency tensor that is a predecessor of the current target tensor, saving the current request and the corresponding current high storage efficiency tensor as historical intermediate data to the storage space, and updating the mapping table; inputting the historical target tensor and the current target tensor into the attention layer for calculation to obtain the inference text corresponding to the current request.

[0148] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the large language model inference method provided by the above-mentioned methods, the method comprising: upon receiving a current request, searching the storage space for historical high storage efficiency tensors corresponding to each stored request according to a mapping table, and converting the historical high storage efficiency tensors into historical target tensors; the mapping table is used to record the correspondence between each request and the high storage efficiency tensors stored in the storage space; performing the current round of inference through the large language model according to the current request, obtaining the intermediate tensors corresponding to each layer of the large language model during the current round of inference, and obtaining the current target tensor required for the attention layer calculation of the large language model from the intermediate tensors; searching for the current high storage efficiency tensor that is a predecessor of the current target tensor, saving the current request and the corresponding current high storage efficiency tensor as historical intermediate data to the storage space, and updating the mapping table; inputting the historical target tensor and the current target tensor into the attention layer for calculation to obtain the inference text corresponding to the current request.

[0149] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0150] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A large language model inference method, characterized in that: Used in an inference system, wherein a large language model is loaded in the inference system, the method includes: When a current request is received, searching the storage space for historical high-storage-efficiency tensors corresponding to each stored request according to a mapping table, and converting the historical high-storage-efficiency tensors into historical target tensors; the mapping table is used to record the correspondence between each request and the high-storage-efficiency tensors stored in the storage space; Performing a current round of reasoning using the large language model according to the current request, obtaining intermediate tensors corresponding to each layer of the large language model during the current round of reasoning, and obtaining a current target tensor required for attention layer calculation of the large language model from the intermediate tensors; Searching for a current high-storage-efficiency tensor that is a predecessor of the current target tensor, saving the current request and the corresponding current high-storage-efficiency tensor as historical intermediate data to a storage space, and updating the mapping table; The historical target tensor and the current target tensor are input into the attention layer for calculation to obtain the inference text corresponding to the current request.

2. The large language model inference method according to claim 1, characterized in that: Converting the historical high-storage-efficiency tensor to a historical target tensor specifically includes: Reading the historical high storage efficiency tensor from the storage space, and allocating temporary storage space in the hardware accelerator for temporary storage; Allocating low-priority computation streams in the hardware accelerator to distinguish them from computation streams performing inference; Converting the historical high-storage-efficiency tensor into the historical target tensor using the low-priority computation flow; The historical target tensor is copied to the cache area of ​​the corresponding user according to the order of the historical characters.

3. The large language model inference method according to claim 1, characterized in that: The storage space includes main memory and persistent storage; Saving the current request and the corresponding current high-storage-efficiency tensor as historical intermediate data to a storage space, and updating the mapping table, specifically includes: Checking whether the remaining storage space in the main memory is less than a threshold; if the storage space in the main memory is greater than or equal to the threshold, saving the current high storage efficiency tensor as historical intermediate data in the main memory, and updating the mapping table; If the remaining storage space in the main memory is less than the threshold, the reuse distance probability distribution of each user is calculated. According to the reuse distance probability distribution of each user and the length of the inference text generated by the previous round of inference, the hit probability of each user accessing the historical high storage efficiency tensor next time is calculated. The historical high storage efficiency tensor with the lowest hit probability of each user is evicted to the persistent storage, the current high storage efficiency tensor is saved as historical intermediate data in the main memory, and the mapping table is updated.

4. The large language model inference method according to claim 3, characterized in that: Calculate the reuse distance probability distribution of each user, including: Record the total space size of historical high-storage-efficiency tensors stored in the storage space accessed by other users between the current request and the last request of each user as the reuse distance; The reuse distance corresponding to each request of each user is recorded to obtain a probability distribution of the reuse distance of the user accessing the historical high storage efficiency tensor.

5. The large language model inference method according to claim 3, characterized in that: Based on the reuse distance probability distribution of each user and the character length of the inference text generated by the previous round of inference, the hit probability of each user accessing the historical high storage efficiency tensor next time is calculated, specifically including: For each user, predicting the reuse distance probability distribution of the user's next visit based on the reuse distance probability distribution of the user's access to the historical high storage efficiency tensor; Obtaining a correlation curve between the character length of the inference text output by the large language model in the current inference process and the predicted lower bound of the reuse distance of the user's next visit; According to the association curve and the character length of the inference text generated in the previous round of inference, the lower bound of the reuse distance of the user's next visit is modified; Online statistics and analysis of users' historical access requests are performed to obtain the main memory hit probability corresponding to different reuse distance ranges under the ideal eviction policy conditions; The hit probability of the user accessing the historical high storage efficiency tensor in the main memory next time is calculated by combining the main memory hit probabilities corresponding to different reuse distance ranges and the revised reuse distance probability distribution of the user's next access.

6. The large language model inference method according to claim 3, characterized in that: Saving the current high-storage-efficiency tensor as historical intermediate data to the main memory specifically includes: Adaptively allocate space in the main memory according to the input and output character length of the current request, and write the current high storage efficiency tensor as historical intermediate data into the allocated space.

7. The large language model inference method according to claim 2, characterized in that: The historical target tensor and the current target tensor are input into the attention layer for calculation to obtain the inference text corresponding to the current request, specifically including: When the attention layer calculates the input, the temporarily stored historical high storage efficiency tensor is pulled from the temporary storage space as the input of the attention layer; When the output is calculated through the attention layer, the temporarily stored historical high storage efficiency tensor and the current target tensor calculated in the current round are pulled from the temporary space as the input of the attention layer to obtain the inference text corresponding to the current request.

8. The large language model inference method according to claim 3, characterized in that: The main memory is also provided with a waiting queue and a preparatory queue; In the case where multiple current requests are received, the method further includes: Upon receiving a current request, determining the storage location of the historical intermediate data corresponding to the current request, and if the historical intermediate data is located in the main memory, adding the current request to a standby queue; if the historical intermediate data is located in the persistent storage, adding the current request to a waiting queue; When the hardware accelerator has the ability to process a new request, a request is selected from the preparation queue for processing; For a request in the waiting queue, after the historical intermediate data corresponding to the request is loaded from the persistent storage to the main memory, the request is moved to the standby queue.

9. A large language model inference device, characterized in that: Used in an inference system, wherein a large language model is loaded in the inference system, the device comprises: a historical target tensor conversion unit, configured to, upon receiving a current request, search the storage space for historical high-storage-efficiency tensors corresponding to each stored request according to a mapping table, and convert the historical high-storage-efficiency tensors into historical target tensors; the mapping table being configured to record the correspondence between each request and the high-storage-efficiency tensors stored in the storage space; a current target tensor acquisition unit, configured to perform a current round of inference using the large language model according to the current request, obtain intermediate tensors corresponding to each layer of the large language model during the current round of inference, and obtain a current target tensor required for calculation of the attention layer of the large language model from the intermediate tensors; a target tensor storage unit, configured to search for a current high-storage-efficiency tensor that is a predecessor of the current target tensor, save the current request and the corresponding current high-storage-efficiency tensor as historical intermediate data to a storage space, and update the mapping table; An inference unit is used to input the historical target tensor and the current target tensor into the attention layer for calculation to obtain the inference text corresponding to the current request.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the large language model inference method according to any one of claims 1 to 8 is implemented.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the large language model inference method according to any one of claims 1 to 8 is implemented.

12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the large language model inference method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Large language model reasoning optimization method and related equipment

    CN121706977A

  • Large model reasoning acceleration method and system

    CN121920550A

  • A large model inference acceleration method and system

    CN121920550B