Scheduling method and device for key value cache data and large model reasoning method and device

Optimizing the migration of key-value cached data through prediction models and scheduling strategies, the problem of large bandwidth resource consumption in large model inference is solved, the inference efficiency is improved and resource waste is reduced, and it is suitable for heterogeneous systems.

CN120276667AActive Publication Date: 2025-07-08ILUVATAR COREX INC SHANGHAI

Patent Information

Application Number
CN202510219153.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-08
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

During the big model inference process, the bandwidth resources of key-value cached data are consumed relatively large, resulting in high latency and high costs, especially in long-context scenarios, repeatedly reading and writing high-bandwidth memory causes waste of computing power.

Method used

Through the prediction model, the target key-value cache data required for subsequent tokens of the large model is predicted, and a scheduling strategy is generated based on the environment characteristics, and the target key-value cache data is migrated from the first storage space to the second storage space, reducing frequent access to the entire amount of data.

Benefits of technology

It reduces the bandwidth resource usage, improves the efficiency of large-scale model inference, reduces resource waste, and is suitable for memory and processors of heterogeneous systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276667A_ABST
    Figure CN120276667A_ABST
Patent Text Reader

Abstract

The invention provides a key value cache data scheduling method and device and a large model reasoning method and device, and relates to the technical field of artificial intelligence. The method comprises the following steps: predicting target key value cache data required by a large model for reasoning a subsequent token by utilizing a prediction model; the subsequent tokens refer to tokens which are not reasoned by the large model; judging whether the target key value cache data needs to be scheduled or not; if scheduling is needed, generating a scheduling strategy; obtaining target key value cache data from the first storage space according to a scheduling strategy, and storing the target key value cache data to a second storage space; wherein the target key value cache data is used for enabling the large model to infer subsequent tokens. According to the method and the device, frequent access to the first storage space is reduced, and only the required target key value cache data is transmitted every time instead of full-amount key value cache data, so that occupation of bandwidth resources caused by transmission of the key value cache data is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method for scheduling key-value cache data, a large model inference method, and an apparatus. Background Art

[0002] In recent years, with the widespread application of large language models (LLMs) in scenarios such as search engines, intelligent dialogue systems, and real-time translation, the problems of high latency and high cost in their inference services have become increasingly prominent. The large model inference process is usually divided into a prefill stage and a decoding stage:

[0003] The prefill stage is a compute-intensive task that requires encoding the input sequence into key-value cache data (KV cache) at once, occupying a large amount of GPU video memory (for example, a 32B model requires 12GB of KV cache at a 50K context length).

[0004] The decoding stage is a process of serially generating tokens, and the performance bottleneck is concentrated on the data transfer between the GPU high-bandwidth memory (HBM) and the computing unit. Since the generation of each token depends on the historical KV cache, repeatedly reading and writing the HBM in long context scenarios consumes a large amount of bandwidth resources, resulting in wasted computing power. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide a method, an apparatus, an electronic device, and a storage medium for scheduling key-value cache data, so as to reduce the consumption of bandwidth resources through the scheduling of key-value cache data.

[0006] In a first aspect, the embodiments of the present application provide a method for scheduling key-value cache data, including:

[0007] Using a prediction model to predict the target key-value cache data required for the large model to infer subsequent tokens; subsequent tokens refer to tokens that the large model has not yet inferred;

[0008] Determining whether scheduling of the target key-value cache data is required;

[0009] If scheduling is required, generating a scheduling policy;

[0010] Obtaining the target key-value cache data from the first storage space according to the scheduling policy, and storing the target key-value cache data in the second storage space; wherein, the target key-value cache data is used for the large model to infer subsequent tokens; the first storage space is used to store all key-value cache data; during the inference process, the large model reads the required key-value cache data from the second storage space.

[0011] In the embodiments of the present application, a prediction model is used to predict the target key-value cache data required for the large model to infer subsequent tokens, and operations such as scheduling judgment, scheduling policy generation, and data migration are performed on the target key-value cache data, reducing frequent access to the first storage space. Moreover, only the required target key-value cache data is transmitted each time, rather than all the key-value cache data, greatly reducing the occupancy of bandwidth resources caused by transmitting the key-value cache data.

[0012] In any embodiment, using a prediction model to predict the target key-value cache data required for the large model to infer subsequent tokens includes:

[0013] Inputting the inferred tokens, the current token, and the position information of the current token into the prediction model to obtain the target key-value cache data required for the inference of subsequent tokens output by the prediction model; wherein, the current token refers to the token that the large model is currently inferring.

[0014] In the embodiments of the present application, by using the prediction model to analyze the inferred tokens, the current token, and the position information of the current token, context information can be captured, thereby more accurately predicting the target key-value cache data required for subsequent tokens.

[0015] In any embodiment, determining whether scheduling of the target key-value cache data is required includes:

[0016] Determining whether the target key-value cache data is stored in the second storage space;

[0017] If the target key-value cache data is not stored in the second storage space, it is determined that scheduling is required; otherwise, scheduling is not required.

[0018] In the embodiments of the present application, by determining whether the target key-value cache data is stored in the second storage space to determine whether scheduling of the target key-value cache data is required, unnecessary scheduling operations are reduced, and the overhead and time cost of data migration are reduced.

[0019] In any embodiment, generating a scheduling policy includes:

[0020] Inputting the target key-value cache data, the key-value cache data already stored in the second storage space, and the environmental characteristics into the scheduling model to obtain the scheduling policy output by the scheduling model; wherein, the environmental characteristics include at least one of the high-bandwidth memory (HBM) capacity, HBM occupancy rate, dynamic random access memory (DRAM) capacity, DRAM occupancy rate, central processing unit (CPU) direct read / write bandwidth, graphics processing unit (GPU) direct read / write bandwidth, CPU occupancy rate, and GPU occupancy rate.

[0021] In the embodiments of the present application, the scheduling model can consider the current environmental characteristics, so that storage and computing resources can be allocated more effectively, reducing resource waste; and it can be applied to heterogeneous systems with different types of memories and processors to make full use of the overall performance of the system.

[0022] In any embodiment, the scheduling model is obtained by training based on a reinforcement learning algorithm.

[0023] In the embodiments of the present application, the reinforcement learning algorithm can enable the scheduling model to self-adjust and optimize according to the feedback of the environment. Through continuous learning, it can consider the occupancy situations and performance characteristics of various resources, so as to make a more reasonable resource allocation strategy.

[0024] In any embodiment, before storing the target key-value cache data into the second storage space, the method further includes:

[0025] If the second storage space is insufficient, some key-value cache data is deleted from the second storage space.

[0026] In the embodiments of the present application, since the second storage space is limited, when the second storage space is insufficient, in order to be able to store the target key-value cache data required for inferring subsequent tokens, some key-value cache data in the second storage space is deleted to meet the needs of subsequent inference.

[0027] In any embodiment, deleting some key-value cache data from the second storage space includes:

[0028] Excluding the key-value cache data corresponding to the first preset number of tokens in the inference token sequence from the second storage space to obtain the excluded key-value cache data;

[0029] Deleting the key-value cache data that entered the second storage space earliest from the excluded key-value cache data according to the size of the target key-value cache data.

[0030] In the embodiments of the present application, by excluding the key-value cache data corresponding to the first preset number of tokens in the inference token sequence, the most relevant key-value cache data in the inference process is retained, and moreover, deleting the key-value cache data that entered the second storage space earliest makes the second storage space be utilized effectively.

[0031] In any embodiment, the first storage space is the storage space in the CPU, and the second storage space is the storage space in the GPU.

[0032] In a second aspect, the embodiments of the present application provide a large model inference method, including:

[0033] Obtain the target key-value cache data of the token to be inferred through a large model; wherein, the target key-value cache data is obtained by using the scheduling method of the key-value cache data described in the first aspect.

[0034] Infer the token to be inferred through the large model based on the target key-value cache data in the second storage space.

[0035] In the embodiment of the present application, the prediction model is used to predict the target key-value cache data required for the large model to infer subsequent tokens, and operations such as scheduling judgment, scheduling strategy generation, and data migration are performed on the target key-value cache data, reducing frequent access to the first storage space and improving the inference efficiency.

[0036] In any embodiment, the inference process of the large model and the scheduling process of the target key-value cache data are executed asynchronously.

[0037] In the embodiment of the present application, the inference of the large model and the scheduling of the target key-value cache data are executed asynchronously. During the inference process of the large model, the target key-value cache data required for subsequent tokens is predicted and scheduled. When the large model infers subsequent tokens, the target key-value cache data can be directly used without temporarily scheduling the target key-value cache data, improving the inference efficiency of the large model.

[0038] In a third aspect, an embodiment of the present application provides a scheduling device for key-value cache data, including:

[0039] A prediction module, configured to use a prediction model to predict the target key-value cache data required for the large model to infer subsequent tokens; the subsequent tokens refer to the tokens that the large model has not yet inferred.

[0040] A judgment module, configured to judge whether it is necessary to schedule the target key-value cache data.

[0041] A scheduling strategy generation module, configured to generate a scheduling strategy if scheduling is required.

[0042] A scheduling module, configured to obtain the target key-value cache data from the first storage space according to the scheduling strategy and store the target key-value cache data in the second storage space; wherein, the target key-value cache data is used for the large model to infer subsequent tokens; the first storage space is used to store all key-value cache data; during the inference process, the large model reads the required key-value cache data from the second storage space.

[0043] In a fourth aspect, an embodiment of the present application provides a large model inference device, including:

[0044] An acquisition module, configured to obtain target key-value cache data of a token to be inferred through a large model; wherein, the target key-value cache data is obtained by using the scheduling method of the key-value cache data as described in the first aspect.

[0045] An inference module, configured to infer the token to be inferred through the large model based on the target key-value cache data in the second storage space.

[0046] In a fifth aspect, an embodiment of the present application provides an electronic device, including: a processor, a memory, and a bus, wherein:

[0047] The processor and the memory communicate with each other through the bus;

[0048] The memory stores program instructions executable by the processor, and the processor can execute the methods of the first aspect or the second aspect by invoking the program instructions.

[0049] In a sixth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium, including:

[0050] The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the methods of the first aspect or the second aspect.

[0051] In a seventh aspect, an embodiment of the present application provides a computer program product, including computer program instructions, and when the computer program instructions are read and run by a processor, the methods of the first aspect are executed.

[0052] Other features and advantages of the present application will be described in the subsequent specification, and part of them will become obvious from the specification, or can be understood by implementing the embodiments of the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 It is a schematic flowchart of a scheduling method for key-value cache data provided by an embodiment of the present application;

[0055] Figure 2 It is a schematic flowchart of a large model inference method provided by an embodiment of the present application;

[0056] Figure 3 It is a diagram of the large model inference architecture provided by the embodiments of the present application;

[0057] Figure 4 It is a schematic diagram of the prefill stage provided by the embodiments of the present application;

[0058] Figure 5 It is a schematic diagram of the decoding stage provided by the embodiments of the present application;

[0059] Figure 6 It is a schematic diagram of the structure of a scheduling device for key-value cache data provided by the embodiments of the present application;

[0060] Figure 7 It is a schematic diagram of the structure of a large model inference device provided by the embodiments of the present application;

[0061] Figure 8 It is a schematic diagram of the entity structure of an electronic device provided by the embodiments of the present application. Detailed implementation manners

[0062] Next, embodiments of the technical solutions of the present application will be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present application, and thus are only examples and cannot be used to limit the protection scope of the present application.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion.

[0064] In the description of the embodiments of this application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary-secondary relationship of the indicated technical features. In the description of the embodiments of this application, "a plurality" means more than two unless otherwise specifically defined.

[0065] Referring to "embodiments" herein means that specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0066] In the description of the embodiments of the present application, the term "and / or" is merely a description of the associated relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0067] In the description of the embodiments of the present application, the term "plurality" refers to two or more (including two). Similarly, "multiple groups" refers to two or more groups (including two groups), and "multiple pieces" refers to two or more pieces (including two pieces).

[0068] In the description of the embodiments of the present application, unless otherwise clearly specified and limited, technical terms such as "installation", "connection", "connection", and "fixation" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may also be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific situations.

[0069] With the wide application of large models in online services, such as search engines, chatbots, and virtual assistants, their high operating costs have become a significant obstacle. The large model inference service includes a prefill stage and a decoding stage. Among them, the prefill stage is a computing power-intensive task, and the decoding stage is used to serially generate tokens. The performance bottleneck lies in the bandwidth between the HBM and the GPU. The decoding stage affects the inference latency between generated tokens, which is related to the key indicator TBT (time between tokens) of inference performance. Especially in the case of long contexts, such as the single-user inference scenario of a 32B model, a context length of 50K requires 12GB of key-value cache data (KV cache). If multiple users need to use the large model for inference simultaneously, the required KV cache will increase significantly.

[0070] Since in the decoding stage of large model inference, for each inferred token, all KV caches need to be read from the CPU to the GPU, and the computing power and bandwidth of the hardware are relatively large, resulting in waste of computing power resources. In subsequent research, it was found that large models often do not require all token information during the decoding process of inference services. The latest work of Microsoft shows that there are certain patterns in the context token information, such as A-shape, block-shape, etc.

[0071] Among them, A-shape represents a specific sparse pattern in which the distribution of attention takes a shape similar to the letter "A". Specifically, this pattern may show a tendency in some heads that the attention is mainly concentrated in a central area and gradually weakens from the central area to both sides. This reflects that when processing long contexts, the model may pay more attention to certain key information or specific parts in the context.

[0072] Block-shape represents another sparse mode, in which attention is concentrated on several block-shaped areas. In this mode, the model may focus on certain specific blocks or paragraphs of text when processing text.

[0073] Based on this, the embodiments of the present application provide a scheduling method for key-value cache data, a large model reasoning method and a device. The target key-value cache data (hereinafter referred to as the target KV cache) required for the large model to reason about subsequent tokens is predicted through a prediction model. If the target KV cache is not in the second storage space, the target KV cache is scheduled from the first storage space storing the full amount of KV cache, and the target KV cache is stored in the second storage space, so that the large model can read the target KV cache from the second storage space when reasoning about subsequent tokens. During the scheduling process, only the target KV cache is scheduled, rather than the full amount of KV cache, which reduces the bandwidth resources occupied by data scheduling.

[0074] The following is an introduction to the specific implementation methods:

[0075] It should be noted that the scheduling method for key-value cache data provided in the embodiments of the present application can be applied to electronic devices, which include terminals and servers; wherein the terminals can be specifically smart phones, tablet computers, computers, personal digital assistants (PDAs), etc.; the servers can be specifically application servers or web servers. The large model, prediction model, and scheduling model mentioned in the following embodiments are run in the electronic device.

[0076] Figure 1 A flowchart of a method for scheduling key-value cache data provided by an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes:

[0077] Step 101: Use the prediction model to predict the target key value cache data required for the large model to infer subsequent tokens; subsequent tokens refer to tokens that have not yet been inferred by the large model.

[0078] Among them, the prediction model is pre-trained, and its main function is to predict the target KV cache required for the tokens that the large model is about to infer. The prediction model may include a pooling layer, a fully connected layer, etc. During the training process, the pre-training of the small model can be carried out using the data of the commonly used KV cache sparsification scheme; for example: generating the tokens sequence data required during the inference service process of large models such as Minference. It can be understood that the prediction model can be deployed on the CPU of the electronic device or on the GPU of the electronic device.

[0079] A large language model (LLM) refers to a deep learning model trained using a large amount of text data, enabling the model to generate natural language text or understand the meaning of language text. These models can provide in-depth knowledge and language production on various topics by being trained on a vast dataset. The core idea is to learn the patterns and structures of natural language through large-scale unsupervised training, simulating the human language cognition and generation process to a certain extent. Currently commonly used large models include: GPT series, BERT series, Tongyi large model, etc. The embodiments of this application do not limit the specific type of the large model.

[0080] The subsequent token refers to the token that the large model has not yet inferred at the current moment. For example: if the large model is currently inferring the current token, then the subsequent token can be the next token that the large model is about to infer, or the next two tokens after the current token, etc. For the convenience of description, the embodiments of this application use the subsequent token as the next token of the currently inferred token.

[0081] Step 102: Determine whether it is necessary to schedule the target key-value cache data.

[0082] After obtaining the target KV cache, it is determined whether scheduling of the target KV cache is required. The main determination method is as follows: Determine whether the second storage space contains the target KV cache. If it does, it is determined that scheduling is not required; otherwise, it is determined that scheduling is required. Among them, for the case where only some of the target KV caches are included in the second storage space, it is also considered that the target KV cache is not included and scheduling is required. However, in the subsequent scheduling process, only the part of the target KV cache that is not included in the second storage space is stored in the second storage space. For example: If the target KV cache includes the KV caches corresponding to the 1st - 10th tokens, the 30th - 50th tokens, and the 100th - 110th tokens. Currently, the second storage space stores the KV caches corresponding to the 1st - 10th tokens and the 30th - 50th tokens, and does not store the KV cache corresponding to the 100th - 110th tokens, then it is determined that scheduling is required. When scheduling, the KV cache corresponding to the 100th - 110th tokens is stored in the second storage space.

[0083] It can be understood that the second storage space is the storage space for the target KV cache that needs to be read during the inference process of the large model. Specifically, it can be the HBM in the GPU, or other storage spaces in the GPU. The embodiments of the present application do not make specific limitations in this regard. In addition, the prediction model can output the index value of the target KV cache, and use the index value to check whether there is a corresponding KV cache with the same index value in the second storage space, so as to determine whether scheduling is required. And when scheduling is required, the corresponding KV cache can also be obtained from the first storage space based on the index value, and then stored in the second storage space.

[0084] Step 103: If scheduling is required, generate a scheduling policy.

[0085] Among them, the scheduling policy is used to indicate the way for the electronic device to schedule the target KV cache from the first storage space to the second storage space. For example: How large data blocks are used for scheduling, whether it is necessary to delete the KV caches in the second storage space, and how many KV caches to delete, etc. When generating the scheduling policy, environmental information such as the current bandwidth on the electronic device can be considered.

[0086] Step 104: Obtain the target key - value cache data from the first storage space according to the scheduling policy, and store the target key - value cache data in the second storage space; where the target key - value cache data is used for the large model to infer subsequent tokens.

[0087] Among them, after determining the scheduling policy, the target KV cache is obtained from the first storage space according to the scheduling policy, and the target KV cache is stored in the second storage space. The first storage space can be the DRAM on the CPU or other memories, and the embodiments of the present application do not make specific limitations on this. In the prefill stage of the large model, the full amount of KV cache is stored in the first storage space. When scheduling the target KV cache, the target KV cache can be migrated from the first storage space to the second storage space. After migration, the first storage space will no longer contain the target KV cache; alternatively, the target KV cache can be obtained from the first storage space and then stored in the second storage space. At this time, the first storage space still contains the target KV cache.

[0088] The embodiments of the present application predict the target key-value cache data required for the large model to infer subsequent tokens through a prediction model, and perform operations such as scheduling judgment, scheduling policy generation, and data migration on the target KV cache, reducing frequent access to the first storage space. Moreover, each time only the required target key-value cache data is transmitted, rather than the full amount of key-value cache data, greatly reducing the occupancy of bandwidth resources caused by transmitting key-value cache data.

[0089] Based on the above embodiments, predicting the target key-value cache data required for the large model to infer subsequent tokens by using a prediction model includes:

[0090] Inputting the inferred tokens, the current token, and the position information of the current token into the prediction model to obtain the target key-value cache data required for the subsequent token inference output by the prediction model; where the current token refers to the token that the large model is currently inferring.

[0091] In the specific implementation process, during the large model inference process, in order to improve the inference efficiency, KVcache is usually used to store the calculated hidden states or intermediate results, thereby avoiding repeated calculations.

[0092] The inferred tokens refer to the tokens that the large model has inferred, and can be specifically represented in the form of a token sequence. The hidden states or intermediate results of these tokens have been stored in the KV cache.

[0093] The current token refers to the token that the large model is currently inferring, and the position information of the current token refers to the position of the token that the large model is currently inferring in the inference service sequence, which can be represented by an index value index and is used for the prediction model to understand the sequential relationship of the tokens.

[0094] Combine the inferred tokens, the current token, and the position information of the current token into a complete input sequence, and input it into the prediction model. The prediction model generates the target KV cache required for the large model to infer the next token based on the input content.

[0095] In the embodiments of the present application, by using the prediction model to analyze the inferred tokens, the current token, and the position information of the current token, the context information can be captured, so as to more accurately predict the target key-value cache data required for subsequent tokens.

[0096] On the basis of the above embodiments, a scheduling policy is generated, including:

[0097] Input the target key-value cache data, the key-value cache data already stored in the second storage space, and the environmental characteristics into the scheduling model to obtain the scheduling policy output by the scheduling model; wherein, the environmental characteristics include at least one of the high-bandwidth memory (HBM) capacity, HBM occupancy rate, dynamic random access memory (DRAM) capacity, DRAM occupancy rate, central processing unit (CPU) direct read-write bandwidth, graphics processing unit (GPU) direct read-write bandwidth, CPU occupancy rate, and GPU occupancy rate.

[0098] In the specific implementation process, in the embodiments of the present application, by introducing the scheduling model, the target KV cache and the environmental characteristics are used as inputs to output a better scheduling policy to guide operations such as reading, writing, migration, and elimination of the KV cache.

[0099] Among them, the target KV cache is predicted by the prediction model in the above embodiments. The environmental characteristics refer to the hardware and software environmental characteristics of the electronic device, including at least one of the following:

[0100] High-bandwidth memory (HBM) capacity: The total capacity of the HBM, which reflects how much data the system can store.

[0101] HBM occupancy rate: The current usage of the HBM, which is used to evaluate the remaining space of the HBM.

[0102] Dynamic random access memory (DRAM) capacity: The total capacity of the DRAM, which also reflects the storage capacity.

[0103] DRAM occupancy rate: The current usage of the DRAM, which is used to evaluate the remaining space of the DRAM.

[0104] CPU direct read-write bandwidth: The rate at which the CPU accesses the memory (including HBM and DRAM), which affects the data processing speed.

[0105] GPU Direct Read / Write Bandwidth: The rate at which the GPU accesses memory, which is particularly important for tasks such as graphics processing and deep learning.

[0106] CPU Occupancy Rate: The current workload of the CPU, which reflects the usage of its processing capabilities.

[0107] GPU Occupancy Rate: The current workload of the GPU, which also reflects the usage of its processing capabilities.

[0108] The target key-value cache data refers to the KV cache required for the large model to infer the next token. The role of the tokens already stored in the second storage space is to determine whether the target KV cache has been stored in the second storage space, so as to determine whether KV cache scheduling is required.

[0109] Input the prepared target KV cache, the KV cache already stored in the second storage space, and the environmental characteristics into the scheduling model. The scheduling model can be a model based on machine learning, deep learning, or reinforcement learning, and learns from the input data and outputs the optimal scheduling strategy. That is, first determine whether KV cache scheduling is required, and in the case of required scheduling, how to perform KV cache scheduling.

[0110] The scheduling model can internally contain multiple processing layers, such as an input layer, a hidden layer, and an output layer, and processes the input data and generates an output through a complex network structure and algorithm.

[0111] The scheduling strategy output by the scheduling model can include the read / write order of key-value cache data, migration strategy, eviction rules, etc. The scheduling strategy aims to optimize system performance, such as reducing access latency, increasing data throughput, or balancing resource usage. In the embodiments of the present application, the scheduling model can consider the current environmental characteristics, so as to more effectively allocate storage and computing resources, reduce resource waste; and can be applied to heterogeneous systems with different types of memories and processors to fully utilize the overall performance of the system.

[0112] Based on the above embodiments, an adaptive reinforcement learning framework can be adopted for the scheduling model. Among them, adaptive means that it can adapt to reinforcement learning schemes with different hardware resources and different large models. In the embodiments of the present application, after performing one reinforcement learning on the scheduling model, it can adapt to different hardware resources. When the scheduling model is deployed on different hardware, in order to optimize the model performance, the scheduling model can be fine-tuned.

[0113] The core idea of reinforcement learning is that the agent learns the optimal policy by interacting with the environment. The agent selects actions in a given state and observes the rewards or punishments given by the environment, so as to adjust its policy in order to obtain greater cumulative rewards in the future. This process can be summarized as follows: the agent selects actions according to the current state, the environment returns the next state and reward according to the actions, and the agent updates the policy according to the rewards.

[0114] The reinforcement learning in the embodiments of this application is not limited to classical value-based or policy-based reinforcement learning methods, such as: Deep Q-Network, abbreviated as DQN; Asynchronous Advantage Actor Critic, abbreviated as A3C; Proximal Policy Optimization (PP0), etc.

[0115] The focus of reinforcement learning lies in the representation of the state space, reward function, and actions. The state space includes: the category of the large model framework adopted, the inferred tokens, the current token, the position information and storage location of the current token (including the storage locations of the inferred tokens and the current token), HBM capacity, HBM occupancy rate, DRAM capacity, DRAM occupancy rate, direct read and write bandwidth and occupancy rate of CPU and GPU, and the output tokens_needs (i.e., the target KV cache) of the pre-trained prefetch small model, etc.; it is not limited to heterogeneous storage media such as HBM and DRAM, and is not limited to heterogeneous computing power accelerators such as CPU and GPU. The state composition features are described as follows:

[0116] S = [Token_position, Token, whole_pre_tokens, whole_DRAM, DRAM_Usage, whole_HBM, HBM_Usage, Read_bandwidth, Write_bandwidth, Read_Usage_bandwidth, Write_Usage_bandwidth; tokens_needs].

[0117] Among them, Token_position is the position information of the current token; Token is the current token; whole_pre_tokens are the tokens that have been inferred; whole_DRAM is the DRAM capacity; DRAM_Usage is the DRMA occupancy rate; whole_HBM is the HBM capacity; HBM_Usage is the HBM occupancy rate; Read_bandwidth is the read bandwidth; Write_bandwidth is the write bandwidth; Read_Usage_bandwidth is the read occupancy rate; Write_Usage_bandwidth is the write occupancy rate; tokens_needs is the target KV cache.

[0118] The action is a communication or retention action for the data required for inferring the next token. The reward function includes: the negative value of the loss of the current inference, the HBM occupancy rate, and the bandwidth occupancy rate; this reward function can be expressed as:

[0119] Reward(s,index) = -loss + HBM_Usag + Read_Usag_bandwidth + Write_Usag_bandwidth

[0120] Among them, Reward(s,index) is the reward value; loss is the loss value of the current inference; HBM_Usag is the HBM occupancy rate; Read_Usag_bandwidth is the read bandwidth occupancy rate; Write_Usag_bandwidth is the write bandwidth occupancy rate.

[0121] The action space of reinforcement learning defines the operations that can be executed, mainly including: migration operations and maintaining the original state actions; the migration operations include: migrating a certain data block from HBM to DRAM; migrating a certain data block from DRAM back to HBM.

[0122] In the embodiments of the present application, the reinforcement learning algorithm can enable the scheduling model to self-adjust and optimize according to the feedback of the environment. Through continuous learning, it can consider the occupancy situations and performance characteristics of various resources, so as to make a more reasonable resource allocation strategy.

[0123] Based on the above embodiments, before storing the target key-value cache data in the second storage space, the method further includes:

[0124] If the second storage space is insufficient, some key-value cache data is deleted from the second storage space.

[0125] In the specific implementation process, the capacity of the second storage space is limited and it is impossible to store data without limit. Therefore, when it is necessary to schedule the target key-value cache data to the second storage space, if it is determined that the second storage space is insufficient, deleting some KV caches from the second storage space is an effective solution. It should be noted that the judgment of whether the second storage space is insufficient can be preset with judgment criteria. For example, when the occupancy rate of the second storage space reaches 80%, it can be considered that the storage space is insufficient; or, the remaining capacity of the second storage space is less than the size of the target KV cache to be scheduled.

[0126] When deleting some KV caches, the KV caches with the lowest access frequency can be deleted. Specifically, this can be achieved by maintaining an access counter or using the LRU (Least Recently Used) cache eviction policy. Deletion can also be based on importance or priority. For example, each KV cache in the second storage space has its own priority, and when deleting, the KV caches with lower priority are deleted.

[0127] There are also various strategies for the number of KV caches to be deleted. For example, deletion can be performed according to a preset size. For example, 5K of KV caches are deleted each time. If after deletion, it still does not meet the requirement of storing the target KV cache in the second storage space, then another 5K of KV caches are deleted until it meets the requirement of storing the target KV cache in the second storage space. Deletion can also be based on the size of the target KV cache to be scheduled, and the KV caches of the corresponding size are deleted from the second storage space.

[0128] In the embodiments of the present application, due to the limited capacity of the second storage space, when the second storage space is insufficient, in order to be able to store the target key-value cache data required for inferring subsequent tokens, some key-value cache data in the second storage space are deleted to meet the needs of subsequent inference.

[0129] Based on the above embodiments, when deleting some key-value cache data from the second storage space, the key-value cache data corresponding to the first preset number of tokens in the inference token sequence can be excluded to obtain the excluded key-value cache data; then, from the excluded key-value cache data, the key-value cache data that entered the second storage space earliest is deleted according to the size of the target key-value cache data.

[0130] In a specific implementation process, the inference token sequence refers to the index information corresponding to the context tokens. For example, for the 2nd token and the 9th token, the inference tokens sequence is [2, 9]. For instance, when inputting a question text to a large model, the KV cache corresponding to the first preset number (e.g., 20) of tokens obtained after tokenization by the large model is the KV cache to be excluded in the embodiments of the present application. Since during the inference process of the large model, the KV cache corresponding to the first preset number of tokens in the inference token sequence is relatively important and has a greater impact on the inference of subsequent tokens, this part of the KV cache needs to be retained in the second storage space. The specific value of the preset number can be 5, 10, 20, 25, etc.; it can also be a percentage value. For example, take the first 1% of the tokens in the tokens sequence, etc., which can be determined according to experience specifically. When deleting part of the KV cache in the second storage space, the KV cache corresponding to the first preset number of tokens in the inference token sequence can be excluded to obtain the excluded KV cache. Then, from the excluded KV cache, according to the size of the target KV cache to be scheduled, the KV cache that entered the second storage space earliest is deleted. That is, the size of the deleted KV cache is greater than or equal to the target KV cache to be scheduled. It can be understood that the target KV cache to be scheduled refers to the KV cache that needs to be stored in the second storage space, which can be all the target KV cache or part of the target KV cache, depending on whether the second storage space contains part of the target KV cache. If it contains, then the part of the target KV cache that is not included in the second storage space is stored in the second storage space.

[0131] In the embodiments of the present application, by excluding the key-value cache data corresponding to the first preset number of tokens in the inference token sequence, the most relevant key-value cache data during the inference process is retained. Moreover, by deleting the key-value cache data that entered the second storage space earliest, the effective utilization of the second storage space is achieved.

[0132] Figure 2 It is a schematic flowchart of a large model inference method provided by the embodiments of the present application, as Figure 2 shown, the method includes:

[0133] Step 201: Obtain the target key-value cache data of the token to be inferred through the large model; wherein, the target key-value cache data is obtained by using the key-value cache data scheduling method provided in the above embodiments;

[0134] Step 202: Use the large model to perform inference on the token to be inferred based on the target key-value cached data in the second storage space.

[0135] In a specific implementation process, Figure 3 This is the architecture diagram of the large model inference provided by the embodiment of the present application, as Figure 3 shown. The large model is deployed on an electronic device, which includes a CPU and a GPU. The CPU includes a CPU computing core and a DRAM. Among them, the DRAM stores the full KV cache, and can also store the relevant parameters of the data prefetch module. It can be understood that the data prefetch module may include a prediction model and a scheduling model. The GPU includes a GPU computing core and an HBM, and the HBM stores the weights of the large model LLM and the hot KV cache. It can be understood that the hot KV cache refers to the KV cache frequently used during the inference process of the large model, and the KV cache required for inferring subsequent tokens.

[0136] The prediction model is used to predict the target KV cache required for the large model to infer subsequent tokens, and the scheduling model is used to schedule the target KV cache, that is, to schedule the target KV cache from the CPU to the GPU. It should be noted that the prediction method of the prediction model can be referred to the above embodiments, and will not be elaborated here. Similarly, the scheduling strategy generated by the scheduling model and the specific method of scheduling according to the scheduling strategy can also be referred to the above embodiments, and will not be elaborated here.

[0137] The embodiment of the present application predicts the target key-value cached data required for the large model to infer subsequent tokens through the prediction model, and performs operations such as scheduling judgment, scheduling strategy generation, and data migration on the target key-value cached data, reducing frequent access to the first storage space and improving the inference efficiency.

[0138] Based on the above embodiments, the large model includes a prefill stage and a decoding stage during the inference process. Figure 4 This is the schematic diagram of the prefill stage provided by the embodiment of the present application, Figure 5 This is the schematic diagram of the decoding stage provided by the embodiment of the present application, as Figure 4 and Figure 5 shown. In the prefill stage, the GPU computing core of the GPU calculates and generates the full KV cache based on the input of the large model, and stores the full KV cache in the DRAM of the CPU. In the decoding stage, the hot KV cache scheduled from the CPU is stored in the HBM of the GPU. It can be understood that the hot KV cache includes the target KV cache required for the large model to infer subsequent tokens.

[0139] In addition, the scheduling of the target KV cache by the large model inference, prediction model, and scheduling model can be executed asynchronously. For example, during the inference of the current token by the large model, the prediction model can predict the target KV cache required for the large model to infer the next token. If the target KV cache is already stored in the second storage space (e.g., the HBM of the GPU), there is no need for scheduling; if the target KV cache is not in the second storage space, the scheduling model will schedule the target KV cache from the first storage space to the second storage space. When the large model infers the next token, it can directly read the target KV cache from the second storage space. At this time, the prediction model can predict the target KV cache required for the large model to infer the token after the next token, and so on. By executing asynchronously, when the large model infers a certain token, it can directly read the target KV cache required for the inference of that token from the GPU, saving the time of reading the full KV cache from the CPU and improving the inference efficiency; and reducing the number of KV cache reads and the occupancy of the bandwidth.

[0140] Figure 6 FIG. is a schematic structural diagram of a scheduling device for key-value cache data provided by an embodiment of the present application. The device may be a module, program segment, or code on an electronic device. It should be understood that the device corresponds to the above Figure 1 method embodiment and can execute Figure 1 each step involved in the method embodiment. The specific functions of the device can be referred to the description above. To avoid repetition, the detailed description is appropriately omitted here. The device includes: a prediction module 601, a judgment module 602, a scheduling policy generation module 603, and a scheduling module 604, where:

[0141] The prediction module 601 is used to predict the target key-value cache data required for the large model to infer subsequent tokens using a prediction model; the subsequent tokens refer to the tokens that the large model has not yet inferred;

[0142] The judgment module 602 is used to judge whether scheduling of the target key-value cache data is required;

[0143] The scheduling policy generation module 603 is used to generate a scheduling policy if scheduling is required;

[0144] The scheduling module 604 is used to obtain the target key-value cache data from the first storage space according to the scheduling policy and store the target key-value cache data in the second storage space; where the target key-value cache data is used to enable the large model to infer subsequent tokens; the first storage space is used to store all the key-value cache data; during the inference process, the large model reads the required key-value cache data from the second storage space.

[0145] Based on the above embodiments, the prediction module 601 is specifically configured to:

[0146] Input the inferred tokens, the current token, and the position information of the current token into the prediction model to obtain the target key-value cache data required for inferring the subsequent token output by the prediction model; wherein, the current token refers to the token being inferred by the large model.

[0147] Based on the above embodiments, the judgment module 602 is specifically configured to:

[0148] Judge whether the target key-value cache data is stored in the second storage space;

[0149] If the target key-value cache data is not stored in the second storage space, it is determined that scheduling is required; otherwise, scheduling is not required.

[0150] Based on the above embodiments, the scheduling policy generation module 603 is specifically configured to:

[0151] Input the target key-value cache data, the key-value cache data already stored in the second storage space, and the environmental characteristics into the scheduling model to obtain the scheduling policy output by the scheduling model; wherein, the environmental characteristics include at least one of the high-bandwidth memory HBM capacity, HBM occupancy rate, dynamic random access memory DRAM capacity, DRAM occupancy rate, central processing unit CPU direct read / write bandwidth, graphics processing unit GPU direct read / write bandwidth, CPU occupancy rate, and GPU occupancy rate.

[0152] Based on the above embodiments, the scheduling model is trained based on a reinforcement learning algorithm.

[0153] Based on the above embodiments, the device further includes a deletion module, which is used for:

[0154] If the second storage space is insufficient, delete some key-value cache data from the second storage space.

[0155] Based on the above embodiments, the deletion module is specifically configured to:

[0156] Exclude the key-value cache data corresponding to the first preset number of tokens in the inference token sequence from the second storage space to obtain the excluded key-value cache data;

[0157] Delete the key-value cache data that entered the second storage space earliest from the excluded key-value cache data according to the size of the target key-value cache data.

[0158] Based on the above embodiments, the first storage space is the storage space in the CPU, and the second storage space is the storage space in the GPU.

[0159] Figure 7 This is a schematic structural diagram of a large model inference device provided by an embodiment of the present application. The device can be a module, a program segment, or code on an electronic device. It should be understood that the device corresponds to the above Figure 2 method embodiment and can execute Figure 2 each step involved in the method embodiment. The specific functions of the device can be referred to the description above. To avoid repetition, the detailed description is appropriately omitted here. The device includes: an acquisition module 701 and an inference module 702, where:

[0160] The acquisition module is configured to obtain the target key-value cache data of the token to be inferred through the large model; wherein, the target key-value cache data is obtained by using the scheduling method of the key-value cache data as described in the first aspect;

[0161] The inference module is configured to infer the token to be inferred based on the target key-value cache data in the second storage space through the large model.

[0162] Based on the above embodiments, the inference process of the large model and the scheduling process of the target key-value cache data are executed asynchronously.

[0163] Figure 8 This is a schematic structural diagram of an electronic device entity provided by an embodiment of the present application. As Figure 8 shown, the electronic device includes: a processor 801, a memory 802, and a bus 803; where:

[0164] The processor 801 and the memory 802 communicate with each other through the bus 803;

[0165] The processor 801 is configured to call program instructions in the memory 802 to execute the methods provided in the above method embodiments. For example, it includes: using a prediction model to predict the target key-value cache data required for the large model to infer subsequent tokens; the subsequent tokens refer to the tokens that the large model has not yet inferred; determining whether scheduling of the target key-value cache data is required; if scheduling is required, generating a scheduling policy; obtaining the target key-value cache data from the first storage space according to the scheduling policy, and storing the target key-value cache data in the second storage space; wherein, the target key-value cache data is used for the large model to infer the subsequent tokens; the first storage space is used to store all key-value cache data; during the inference process, the large model reads the required key-value cache data from the second storage space. Or, obtaining the target key-value cache data of the token to be inferred through the large model; wherein, the target key-value cache data is obtained by using the key-value cache data scheduling method provided in the above embodiments; inferring the token to be inferred through the large model based on the target key-value cache data in the second storage space.

[0166] The processor 801 can be an integrated circuit chip with signal processing capabilities. The above processor 801 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0167] The memory 802 can include, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0168] This embodiment discloses a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above method embodiments. For example, it includes: using a prediction model to predict the target key-value cache data required for the large model to infer subsequent tokens; the subsequent tokens refer to the tokens that the large model has not yet inferred; determining whether it is necessary to schedule the target key-value cache data; if scheduling is required, generating a scheduling strategy; obtaining the target key-value cache data from a first storage space according to the scheduling strategy, and storing the target key-value cache data in a second storage space; wherein, the target key-value cache data is used to enable the large model to infer the subsequent tokens; the first storage space is used to store all key-value cache data; during the inference process, the large model reads the required key-value cache data from the second storage space. Or, obtaining the target key-value cache data of the token to be inferred through the large model; wherein, the target key-value cache data is obtained by using the scheduling method of the key-value cache data provided in the above embodiments; inferring the token to be inferred by the large model based on the target key-value cache data in the second storage space.

[0169] This embodiment provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions. The computer instructions cause the computer to execute the methods provided in the above method embodiments. For example, it includes: using a prediction model to predict the target key-value cache data required for the large model to infer subsequent tokens; the subsequent tokens refer to the tokens that the large model has not yet inferred; determining whether it is necessary to schedule the target key-value cache data; if scheduling is required, generating a scheduling strategy; obtaining the target key-value cache data from a first storage space according to the scheduling strategy, and storing the target key-value cache data in a second storage space; wherein, the target key-value cache data is used to enable the large model to infer the subsequent tokens; the first storage space is used to store all key-value cache data; during the inference process, the large model reads the required key-value cache data from the second storage space. Or, obtaining the target key-value cache data of the token to be inferred through the large model; wherein, the target key-value cache data is obtained by using the scheduling method of the key-value cache data provided in the above embodiments; inferring the token to be inferred by the large model based on the target key-value cache data in the second storage space.

[0170] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical or other form.

[0171] In addition, the units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0172] Furthermore, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0173] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0174] The above description is only for the embodiments of the present application and does not limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A scheduling method for key-value cached data, characterized in that Including: Using a prediction model to predict the target key-value cache data required for the large model to infer subsequent tokens; the subsequent tokens refer to the tokens that the large model has not yet inferred; Judging whether it is necessary to schedule the target key-value cache data; If scheduling is required, generating a scheduling policy; Obtaining the target key-value cache data from the first storage space according to the scheduling policy and storing the target key-value cache data in the second storage space; wherein, the target key-value cache data is used for the large model to infer the subsequent tokens; the first storage space is used to store all key-value cache data; During the inference process, the large model reads the required key-value cache data from the second storage space.

2. The method according to claim 1, wherein The using a prediction model to predict the target key-value cache data required for the large model to infer subsequent tokens includes: Inputting the inferred tokens, the current token, and the position information of the current token into the prediction model to obtain the target key-value cache data required for the subsequent token inference output by the prediction model; wherein, the current token refers to the token that the large model is currently inferring.

3. The method according to claim 1, wherein The judging whether it is necessary to schedule the target key-value cache data includes: Judging whether the second storage space stores the target key-value cache data; If the second storage space does not store the target key-value cache data, it is determined that scheduling is required; otherwise, scheduling is not required.

4. The method according to claim 1, wherein The generating a scheduling policy includes: Inputting the target key-value cache data, the key-value cache data already stored in the second storage space, and environmental characteristics into a scheduling model to obtain a scheduling policy output by the scheduling model; wherein, the environmental characteristics include at least one of the high-bandwidth memory (HBM) capacity, HBM occupancy rate, dynamic random access memory (DRAM) capacity, DRAM occupancy rate, central processing unit (CPU) direct read / write bandwidth, graphics processing unit (GPU) direct read / write bandwidth, CPU occupancy rate, and GPU occupancy rate.

5. The method according to claim 4, wherein The scheduling model is trained based on a reinforcement learning algorithm.

6. The method according to claim 1, wherein Before storing the target key-value cache data in the second storage space, the method further includes: If the second storage space is insufficient, deleting some key-value cache data from the second storage space.

7. The method according to claim 6, characterized in that, The deleting some key-value cache data from the second storage space includes: Excluding the key-value cache data corresponding to the first preset number of tokens in the inference token sequence in the second storage space to obtain the excluded key-value cache data; Deleting the key-value cache data that entered the second storage space earliest from the excluded key-value cache data according to the size of the target key-value cache data.

8. The method according to any one of claims 1 to 7, characterized in that, The first storage space is the storage space in the CPU, and the second storage space is the storage space in the GPU.

9. A large model inference method, characterized in that, Including: Obtaining the target key-value cache data of the token to be inferred through a large model; wherein, the target key-value cache data is obtained by using the key-value cache data scheduling method according to any one of claims 1-8. The large model performs inference on the token to be inferred based on the target key-value cached data in the second storage space.

10. The method according to claim 9, characterized in that, The inference process of the large model and the scheduling process of the target key-value cached data are executed asynchronously.

11. A scheduling device for key-value cached data, characterized in that It includes: A prediction module for predicting the target key-value cached data required for the large model to infer subsequent tokens using a prediction model; The subsequent tokens refer to the tokens that the large model has not yet inferred; A judgment module for judging whether it is necessary to schedule the target key-value cached data; A scheduling policy generation module for generating a scheduling policy if scheduling is required; A scheduling module for obtaining the target key-value cached data from the first storage space according to the scheduling policy and storing the target key-value cached data in the second storage space; wherein, the target key-value cached data is used for the large model to infer the subsequent tokens; the first storage space is used to store all key-value cached data; During the inference process, the large model reads the required key-value cached data from the second storage space.

12. An inference device for large models, characterized in that, It includes: An acquisition module for obtaining the target key-value cached data of the token to be inferred through the large model; wherein, the target key-value cached data is obtained by using the key-value cached data scheduling method according to any one of claims 1-8; An inference module for performing inference on the token to be inferred by the large model based on the target key-value cached data in the second storage space.

13. An electronic device, characterized in that, It includes: A processor, a memory, and a bus, wherein: The processor and the memory communicate with each other through the bus; The memory stores program instructions that can be executed by the processor, and the processor can execute the method according to any one of claims 1-10 by calling the program instructions.

14. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and when the computer instructions are run by the computer, the computer executes the method according to any one of claims 1-10.

15. A computer program product, characterized in that, It includes computer program instructions, and when the computer program instructions are read and run by the processor, the method according to any one of claims 1-10 is executed.

Citation Information

Patent Citations

  • Optimization method and system for dynamic reasoning memory allocation based on predictor

    CN118227336A

  • Multi-acceleration card multi-task scheduling method based on distributed parallel large model and medium

    CN118626233A

  • Token caching method and electronic equipment

    CN118747120A

  • Task scheduling method of large language model, data processing system and electronic equipment

    CN119311384A

  • Data pre-fetch for large language model (LLM) processing

    US20240370699A1

Cited By

  • Model reasoning method, electronic equipment and storage medium

    CN120450057A

  • Model reasoning method, electronic device and storage medium

    CN120450057B

  • GPU and HBM memory fusion management method and system

    CN120762935A

  • Data processing method, product, electronic equipment and computer readable storage medium

    CN120950009A

  • Access control method and device for heterogeneous storage system, equipment, medium and product

    CN122219856A