Data partitioning method and device, large model reasoning method and device, equipment and medium

By using a data partitioning method based on access frequency, time interval, and probability of future access, hot data is stored in the fastest video memory device, solving the problem of slow inference speed for large models and achieving more efficient data storage and inference.

CN120909791APending Publication Date: 2025-11-07INSPUR (SHANDONG) COMPUTER TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511071642.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

The slow inference speed of large models is mainly due to the limited memory resources of AI accelerators and the rapidly increasing storage requirements of KV Cache, which leads to increased computational latency.

Method used

By determining the access frequency parameters, time intervals, and future access probability of the key-value cache, a predictive model is used to partition the data, and the hot data set is stored in the fastest video memory device, thereby improving data storage efficiency.

Benefits of technology

It improves the inference speed of large models, reduces redundant calculations, and significantly reduces latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909791A_ABST
    Figure CN120909791A_ABST
Patent Text Reader

Abstract

The invention discloses a data division method, a large model reasoning method, devices, equipment and a medium, which are applied to the technical field of data processing, and the data division method comprises the following steps: based on historical access information corresponding to each key value cache, utilizing a prediction model to carry out prediction to obtain the probability that the key value cache associated with each key value cache is accessed in the future; dividing the key value caches based on the access frequency parameter, the time interval and the future access probability of each key value cache, and determining a hot data set and a cold data set, so as to store the key value caches in the hot data set to the video memory equipment with the fastest speed. According to the method, the access frequency is considered from the global aspect, the time interval is considered from the time local aspect, the future access probability is considered from the space local aspect, and the data is divided from different aspects, so that the key value can be obtained from the accurate hot data set in time during reasoning, and the accuracy of the data is improved. Therefore, the reasoning speed of the large model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a data partitioning method and a large model inference method, device, equipment and medium. BACKGROUND

[0002] In the LLM inference process, the KV Cache (Key-Value Cache) caches the Key (key) and Value (value) matrix of the historical token (token) to avoid repeated calculation, thereby significantly reducing the inference delay. As the Context Length (context length) increases, the storage requirement of the KV Cache increases sharply (may occupy tens of GB of memory), and the video memory resource of the AI accelerator is limited, so the large model inference speed is low.

[0003] Therefore, how to improve the large model inference speed is a technical problem that needs to be solved by those skilled in the art. SUMMARY

[0004] Therefore, the purpose of the present application is to provide a data partitioning method and a large model inference method, device, equipment and medium, which solve the technical problem of low large model inference speed in the prior art.

[0005] To solve the above technical problems, the present application provides a data partitioning method, comprising:

[0006] determining the access frequency parameter based on the access frequency parameter of each key-value cache, the time interval and the probability of future access of each key-value cache.

[0007] determining the access time corresponding to each key-value cache, and determining the time interval based on the access time;

[0008] based on the historical access information corresponding to each key-value cache, using a prediction model to predict the probability of future access of the key-value cache associated with each key-value cache; wherein the prediction model is a model for predicting based on the access association relationship between key-value caches;

[0009] based on the access frequency parameter, the time interval and the probability of future access of each key-value cache, partitioning the key-value cache to determine a hot data set and a cold data set, and storing the key-value cache in the hot data set to the fastest video memory device during inference. The fastest video memory device is the video memory device during inference.

[0010] On the one hand, determining the access frequency parameter based on the access frequency parameter of each key-value cache in the sliding time window, comprising:

[0011] determining an access type corresponding to each key-value cache, the access type including write and read;

[0012] determining an actual weight corresponding to an access frequency based on the access type, determining the access frequency parameter based on the actual weight and the access frequency; wherein the weight of write access is greater than the weight of read access.

[0013] In one aspect, the access time corresponding to each key-value cache is determined, and the time interval is determined based on the access time, including:

[0014] obtaining a time queue storing the access time of each key-value cache, and determining a preset number of time intervals based on the time queue;

[0015] determining a weight corresponding to each time interval, and determining the time interval based on each time interval and its corresponding weight.

[0016] In one aspect, before predicting based on the historical access information corresponding to each key-value cache using a prediction model to obtain the probability of future access of the key-value cache associated with each key-value cache, it further includes:

[0017] obtaining historical key-value caches;

[0018] determining an attention head corresponding to each of the historical key-value caches;

[0019] grouping and storing the historical key-value caches according to the attention head to obtain target grouped historical key-value caches;

[0020] determining a neighboring target grouped historical key-value cache corresponding to each of the target grouped historical key-value caches to obtain a neighboring relationship between the attention heads;

[0021] determining an association relationship between each historical key-value cache in two adjacent target grouped historical key-value caches based on the neighboring relationship between the attention heads; wherein the association relationship is the influence of the access of the current historical key-value cache on the access of the associated historical key-value cache.

[0022] training an initial prediction model based on the association relationship between each historical key-value cache to obtain the prediction model.

[0023] In one aspect, based on the access frequency parameter, the time interval and the probability of future access of each key-value cache, the key-value cache is divided to determine a hot data set and a cold data set, and the key-value cache in the hot data set is stored in the fastest visible memory device, and the fastest visible memory device is the visible memory device during inference, including:

[0024] determining a first hot data set based on the access frequency parameter;

[0025] determine a second hot data set based on the time interval;

[0026] determine a third hot data set based on the probability of future access;

[0027] determine the intersection of the key-value cache in the first hot data set, the second hot data set and the third hot data set, and construct the hot data set based on the cached key-value in the intersection;

[0028] data other than the hot data set as the cold data set.

[0029] In one aspect, after dividing the key-value cache based on the access frequency parameter, the time interval and the probability of future access of each key-value cache, determining the hot data set and the cold data set, further comprising:

[0030] determine the remaining storage space of the fastest GPU device, and pre-fetch the key-value cache from the hot data set based on the probability of future access.

[0031] The embodiment of the application further provides a large model inference method, comprising:

[0032] determine the hot data set in the fastest GPU device; wherein the hot data set is a set determined based on the above data division method;

[0033] obtain the target key-value required for current inference from the hot data set, and perform inference based on the target key-value using a large language model.

[0034] The embodiment of the application further provides a data division device, comprising:

[0035] An access frequency determination module is configured to determine the number of accesses of each key-value cache in a sliding time window, and determine an access frequency parameter based on the number of accesses;

[0036] A time interval determination module is configured to determine the access time corresponding to each key-value cache, and determine a time interval based on the access time;

[0037] A probability of future access determination module is configured to use a prediction model to predict based on the historical access information corresponding to each key-value cache, to obtain the probability of future access of the key-value cache associated with each key-value cache; wherein the prediction model is a model for predicting based on the access association relationship between key-value caches;

[0038] A hot data cache module is configured to divide the key-value caches based on the access frequency parameter, the time interval and the probability of future access of each key-value cache, determine a hot data set and a cold data set, and store the key-value caches in the hot data set to a fastest memory device, which is a memory device during inference.

[0039] The embodiment of the present application also provides a large model inference device, which comprises:

[0040] A hot data set determination module is configured to determine a hot data set in the fastest memory device, wherein the hot data set is a set determined based on the data division method.

[0041] An inference module is configured to obtain a target key-value required for current inference from the hot data set, and perform inference based on the target key-value using a large language model.

[0042] The embodiment of the present application also provides an electronic device, which comprises:

[0043] A memory is configured to store a computer program.

[0044] A processor is configured to execute the computer program to implement the steps of the above method.

[0045] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the above method.

[0046] The present application also provides a computer program product, which comprises a computer program / instruction, and the computer program / instruction is executed by a processor to implement the steps of the above method.

[0047] To solve the above technical problem, the embodiment of the present application provides a data division method, which comprises: determining the access frequency parameter of each key-value cache based on the access times of the key-value cache in a sliding time window; determining the time interval of each key-value cache based on the access time corresponding to the key-value cache; obtaining the probability of future access of the key-value cache associated with each key-value cache by using a prediction model based on the historical access information corresponding to each key-value cache; wherein the prediction model is a model for predicting based on the access association relationship between key-value caches; dividing the key-value caches based on the access frequency parameter, the time interval and the probability of future access of each key-value cache, determining a hot data set and a cold data set, and storing the key-value caches in the hot data set to a fastest memory device, which is a memory device during inference.

[0048] From the above technical solution can be seen, the beneficial effects of the present application lie in: the present application considers the access frequency parameter, time interval and the probability of future access three parameters, wherein the access frequency is considered from the global, the time interval is considered from the time local angle, and the probability of future access is considered from the space local, since the present application considers the cold and hot data from different angles, the accuracy of dividing the cold and hot data can be improved, so that the large model can obtain the key value from the accurate hot data set in time during reasoning, and the reasoning speed of the large model is improved. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0050] Figure 1 A flow chart of a data division method provided for the embodiments of the present application is shown in the figure.

[0051] Figure 2 A flow chart of a data division method provided for the embodiments of the present application is shown in the figure.

[0052] Figure 3 A schematic diagram for determining hot data based on access frequency provided for the embodiments of the present application is shown in the figure.

[0053] Figure 4 A schematic diagram for determining hot data based on time interval provided for the embodiments of the present application is shown in the figure.

[0054] Figure 5 A schematic diagram for determining hot data based on the probability of future access provided for the embodiments of the present application is shown in the figure.

[0055] Figure 6 A flow chart of a large model reasoning method provided for the embodiments of the present application is shown in the figure.

[0056] Figure 7 A structural schematic diagram of a data division device provided for the embodiments of the present application is shown in the figure.

[0057] Figure 8 A structural schematic diagram of a large model reasoning device provided for the embodiments of the present application is shown in the figure.

[0058] Figure 9 A structural schematic diagram of an electronic device provided for the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION

[0059] With reference to the drawings and specific embodiments described below, the technical solutions in the embodiments of the present application will be better understood.

[0060] The terms "comprise", "comprising", "include", "including", "have" and "having" in the specification and the above drawings, and any variations thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, system, product or device that includes a list of steps or units is not limited to the listed steps or units, but can include steps or units not listed.

[0061] In order for those skilled in the art to better understand the technical solutions of the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0062] Next, a data partitioning method provided by the embodiments of the present application will be described in detail. Figure 1 A flowchart of a data partitioning method provided by the embodiments of the present application, the method can include:

[0063] S101, determine the number of accesses of each key value cache in the sliding time window, and determine the access frequency parameter based on the number of accesses.

[0064] The steps in the embodiments can be executed by a designated electronic device, which can be a server, a portable terminal or other forms, and the electronic device has a memory bar, and the specific number is not limited. The KV Cache (Key-Value Cache) in the embodiments is a key technology for optimizing the self-attention mechanism. The embodiments do not limit the specific sliding time window, for example, the sliding time window in the embodiments can be 10 seconds, or the sliding time window can also be 5 seconds. The access frequency in the embodiments refers to the number of accesses of the client in a sliding time window. The embodiments do not limit the specific method of determining the access frequency parameter based on the number of accesses. For example, the embodiments can directly use the number of accesses as the access frequency, or the embodiments can also determine the weight corresponding to the number of accesses based on the access type corresponding to each access.

[0065] It needs to be further explained that in order to improve the accuracy of the determination of the access frequency parameter, the above determination of the number of accesses of each key value cache in the sliding time window, and the determination of the access frequency parameter based on the number of accesses, can include:

[0066] S1011, determine an access type corresponding to each key-value cache, the access type including write and read;

[0067] S1012, determine an actual weight corresponding to an access frequency based on the access type, determine the access frequency parameter based on the actual weight and the access frequency; wherein the write access weight is greater than the read access weight.

[0068] In this embodiment, a more accurate access frequency parameter can be determined according to the access type, for example, the write operation weight is set to 1.5 (the write weight is greater than the read weight) x read operation, for example, the access frequency of 1 read operation can be recorded as 1, and the access frequency of 1 write operation can be recorded as 1.5. This embodiment takes into account that in the Transformer inference process, the write operation corresponds to the generation of the KV (key-value) vector of the new token, which represents the update of the knowledge base. The read operation is more to query the existing knowledge. The writing of new knowledge often indicates that there will be intensive access in the future, just like the popular books newly added to the library. Secondly, the hardware efficiency is considered. The write operation needs to occupy the video memory bandwidth to execute the storage instruction, which is 1.5-2 times higher than the read operation. Giving a higher weight means recognizing its resource consumption cost, which will give priority to retaining this kind of data in the prefetch decision. Thirdly, the principle of time locality. Observing the actual inference mode will find that the newly written KV vector is accessed with a very high probability in the next few token generations. For example, after generating the key entity noun, it is likely to be referenced multiple times. This strong association makes the write operation the best indicator of predicting future hotspots. This access frequency statistical mechanism can improve the accuracy of determining the access frequency parameter by differentiating read and write weights. It should be noted that when determining hot data based on access frequency, the top 20% high-frequency access blocks can be initially marked as hot data, and self-adaptive adjustment can be performed, and according to the GPU memory pressure (such as HBM occupancy rate > 80%), it is automatically tightened to top 15% or relaxed to top 25%.

[0069] S102, determine an access time corresponding to each key-value cache, and determine a time interval based on the access time.

[0070] This embodiment does not limit the specific time interval. For example, the time interval in this embodiment can be the time interval between the last two accesses. Or the time interval in this embodiment can be determined based on the time interval between all access times within a set time. The time interval lower than the set time interval can be identified as hot data.

[0071] It should be further pointed out that, in order to improve the accuracy of determining the time interval, the above determining the access time corresponding to each key-value cache and determining the time interval based on the access time can include:

[0072] S1021, acquire a time queue storing access times of each key-value cache, determine a preset number of time intervals based on the time queue;

[0073] S1022, determine a weight corresponding to each time interval, and determine the time interval based on each time interval and the weight corresponding thereto.

[0074] The time queue of this embodiment stores the access times of each key-value cache, and it can be clearly known whether the access times consider the current time. The embodiment does not limit the specific preset number. For example, the preset number in this embodiment can be 10; or the preset number in this embodiment can be 15. The embodiment can set a higher weight for a time interval corresponding to a time closer to the current time, and a lower weight for a time farther from the current time. Compared with determining the time interval based on only the last two access times, the embodiment considers that using a time interval only once is not representative, so a weight is set for each time interval, and a more accurate time interval can be determined from a global perspective.

[0075] S103, predicting based on the historical access information corresponding to each key-value cache using a prediction model to obtain the probability of future access of the key-value cache associated with each key-value cache; wherein the prediction model is a model for predicting based on the access association relationship between key-value caches.

[0076] The prediction model in this embodiment can be a prediction model for determining the probability of future access of the associated key-value cache based on the access information of the current key-value cache. The historical access information in this embodiment can be a historical access frequency; or the historical access information can also be a historical access number. The access association relationship between key-value caches in this embodiment means that if A key-value cache is accessed, B key-value cache will also be accessed, at which time it can be determined that there is an association relationship between A key-value cache and B key-value cache.

[0077] It should be further explained that based on any of the above embodiments, before predicting based on the historical access information corresponding to each key-value cache using a prediction model to obtain the probability of future access of the key-value cache associated with each key-value cache, it can further include:

[0078] Step 1: acquire historical key-value caches;

[0079] Step 2: determine an attention head corresponding to each historical key-value cache;

[0080] Step 3: group and store the historical key-value caches according to the attention heads to obtain target grouped historical key-value caches;

[0081] Step 4: determine adjacent target grouped historical key-value caches corresponding to each target grouped historical key-value cache to obtain adjacent relationships between the attention heads;

[0082] Step 5: determining the association relationship between each historical key-value cache based on the adjacent relationship between the attention heads in the two adjacent target group history key-value caches; wherein the association relationship is the influence of the access of the current historical key-value cache on the access of the associated historical key-value cache;

[0083] Step 6: training the initial prediction model based on the association relationship between each historical key-value cache to obtain a prediction model.

[0084] This embodiment considers that the KV Cache of a large language model is usually stored in groups according to Attention Heads, if the Key / Value matrix of a certain Attention Head is accessed intensively, the KV Cache of the adjacent Head can also be preloaded, a lightweight model is trained based on historical data to predict the range of KV Block that can be accessed in the future. This embodiment gives a specific method for constructing a prediction model based on the specific association relationship between KV Caches, and improves the accuracy of the prediction model training.

[0085] S104, dividing the key-value caches based on the access frequency parameters, time intervals and future access probabilities of each key-value cache, determining a hot data set and a cold data set, and storing the key-value caches in the hot data set to the fastest memory device, which is the memory device during inference.

[0086] The embodiment does not limit the specific method of dividing the key-value cache based on the access frequency parameter, time interval and probability of future access of each key-value cache, determining the hot data set and the cold data set. For example, the embodiment can determine a hot data set based on the access frequency parameter, determine a hot data set based on the time interval, determine a hot data set based on the probability of future access, and the intersection of the three as the hot data set, and the other data as the cold data. The embodiment can set the key-value cache greater than the preset minimum access frequency threshold as hot data, set the key-value cache lower than the set time interval as hot data, and the key-value cache with a probability of future access greater than the set probability threshold as hot data. Or the embodiment can also set a weight for each key-value cache based on the access frequency parameter, time interval and probability of future access, and finally select the key-value cache with a weight greater than the set weight value as the hot data set, wherein the greater the weight, the greater the probability of data access. The embodiment does not limit the specific fastest visible memory device, for example, the fastest visible memory device in the embodiment can be a GPU (graphics processing unit). The embodiment does not limit the specific storage method of hot data and cold data, and the storage method of hot data has a higher visible memory speed than that of cold data. For example, in the embodiment, the storage method of hot data can be GPU, and the storage method of cold data can be SDRAM (synchronous dynamic random access memory) or SSD (solid state disk) or CPU (central processing unit) or (Compute Express Link, high-speed interconnection) extended memory. According to the existing definition, hot data refers to data with high access frequency, data higher than the set access frequency, and hot data refers to data lower than the set access frequency.

[0087] It should be further explained that based on any of the above embodiments, in order to improve the accuracy of determining the hot data set, the above dividing the key-value cache based on the access frequency parameter, time interval and probability of future access of each key-value cache, determining the hot data set and the cold data set, and storing the key-value cache in the hot data set to the fastest visible memory device, the fastest visible memory device is the visible memory device during inference, can include:

[0088] S1041, determining a first hot data set based on the access frequency parameter;

[0089] S1042, determining a second hot data set based on the time interval;

[0090] S1043, determining a third hot data set based on the probability of future access;

[0091] S1044, determining the intersection of the key-value caches in the first hot data set, the second hot data set and the third hot data set, and constructing a hot data set based on the cache key values in the intersection;

[0092] S1045, data other than the heat removal data set is taken as the cold data set.

[0093] The process of determining the first hot data set based on the access frequency parameter in this embodiment can be: the number of accesses within a sliding time window (such as 10 ms), and the read / write weight is distinguished, and the write operation weight is set to 1.5 (the write weight is greater than the read weight) x read operation. Initially, the top 20% high-frequency access block is marked as hot data, and adaptive adjustment is performed. According to the memory pressure of the memory device (such as HBM occupancy rate > 80%), the hot data proportion is automatically adjusted, for example, tightened to Top 15% or relaxed to Top 25%. The process of determining the second hot data set based on the time interval in this embodiment can be: maintaining an LRU (Least Recently Used) queue for each KV Cache, recording the last two access timestamps and calculating the time interval. The time interval lower than the set time interval belongs to hot data and can be saved in the memory device. The long interval belongs to cold data and can be migrated to CXL memory storage, maximizing the release of memory space. The process of determining the third hot data set based on the probability of future access in this embodiment can include: sorting the probability of future access of each key-value cache from large to small, and taking the first preset number of key-value caches as the third hot data set. The intersection of the first hot data set, the second hot data set and the third hot data set in this embodiment is taken as the hot data set, which improves the accuracy of determining the hot data set.

[0094] It needs to be further explained that based on any of the above embodiments, the key-value caches are divided based on the access frequency parameter, the time interval and the probability of future access of each key-value cache to determine the hot data set and the cold data set, which can include: based on the access frequency parameter and its corresponding weight of each key-value cache, the time interval and its corresponding weight of each key-value cache, and the probability of future access and its corresponding weight of each key-value cache, determining the hot data score corresponding to each key-value cache, sorting the hot data score from large to small, and determining the first preset number of key-value caches as the hot data set. Among them, the greater the access frequency parameter, the greater the access frequency score, the shorter the time interval, the greater the time locality score, and the greater the probability of future access, the greater the space locality score. Finally, based on the access frequency score and its weight, the time locality score and its weight, and the space locality score and its weight, the actual score corresponding to each key-value cache can be obtained. This embodiment can determine the weight of the three parameters based on the accuracy of dividing the hot data, so that the score of each key-value cache can be accurately determined.

[0095] It should be further pointed out that, in order to improve the efficiency of the cache, after dividing the key-value cache based on the access frequency parameter, time interval and probability of future access of each key-value cache, determining the hot data set and the cold data set, the method can further include: determining the remaining storage space of the current fastest video memory device, and pre-fetching the key-value cache from the hot data set based on the probability of future access. This embodiment can generate a pre-fetch candidate list (top-k high probability key-value cache), initiate a CXL pre-fetch request through a background thread, manage the pre-fetch task through a priority queue, and dynamically adjust the pre-fetch depth based on the feedback PID controller. This embodiment will pre-fetch the key-value cache from the hot data set, so that the key-value cache with a larger probability of future access can be extracted from the cache, preventing the defect that the key-value cache does not exist when the large model inference is required.

[0096] The data division method provided by the embodiment of the application can include: S101, determining the access frequency parameter of each key-value cache in a sliding time window based on the access frequency parameter; S102, determining the access time corresponding to each key-value cache, and determining the time interval based on the access time; S103, predicting the probability of future access of the key-value cache associated with each key-value cache based on the historical access information corresponding to each key-value cache using a prediction model; wherein the prediction model is a model for predicting based on the access association relationship between key-value caches; S104, dividing the key-value cache based on the access frequency parameter, time interval and probability of future access of each key-value cache, determining the hot data set and the cold data set, and storing the key-value cache in the hot data set to the fastest video memory device during inference. The fastest video memory device is the video memory device during inference. The application considers three parameters, access frequency, time interval and probability of future access. The access frequency is considered from a global perspective, the time interval is considered from a local time perspective, and the probability of future access is considered from a local spatial perspective. Since the application considers cold and hot data from different angles, the accuracy of dividing cold and hot data is improved, so that the large model can obtain key values from the accurate hot data set in time during inference, and the inference speed of the large model is improved.

[0097] In order to make the application more convenient to understand, please refer to Figure 2 , Figure 2 The flowchart of the data division method provided by the embodiment of the application can include:

[0098] S201, record the access frequency of each KV Cache in a sliding time window, determine the access type, and determine the actual access frequency based on the access type and the access frequency; wherein the access type includes read and write, and the write weight = 1.5x read weight.

[0099] S202, determine a first hot data set based on an actual access frequency.

[0100] For ease of understanding, please refer to Figure 3 , Figure 3 A schematic diagram for determining hot data based on access frequency provided by an embodiment of the application can include: monitoring the access of each KV Cache, counting the read and write times in each sliding window, determining the hot data proportion based on the GPU memory pressure, determining the KV Cache based on the access frequency parameter of the read and write type, sorting the KV Cache based on the access frequency parameter, taking the front hot data and the KV Cache on the side as the first hot data set.

[0101] S203, determine the last two access times of each KV Cache, and calculate the time difference between the last two access times, determine a second hot data set based on the time difference.

[0102] For ease of understanding, please refer to Figure 4 , Figure 4 A schematic diagram for determining hot data based on time interval provided by an embodiment of the application can include: recording the access timestamp of each KV Cache in the LRU queue, determining the access interval (time interval) based on the timestamp, taking the time interval as the decay factor, determining the KV Cache with a time interval not greater than a set threshold as hot data, and the KV Cache with a time interval greater than the set threshold as cold data. \n1 is the first step, \n2 is the second step.

[0103] S204, predict the probability of each KV Cache being accessed in the future, determine a third hot data set based on the probability of being accessed in the future.

[0104] For ease of understanding, please refer to Figure 5 , Figure 5A schematic diagram for determining hot data based on a probability of future access provided by an embodiment of the present application can include: monitoring each KV Cache Attention Head, grouping KV Caches according to Attention Heads, recording the access sequence of KV Caches in each Attention Head group, determining the association relationship of KV Caches based on the association relationship between Attention Heads, constructing a spatial relationship graph, training a prediction model based on the historical access sequence in the Attention Head, and obtaining the prediction model. Based on the KV Cache of the current Attention Head, the future access probability of the corresponding associated KV Cache is obtained, and the KV Cache with a future access probability greater than a set value is loaded asynchronously. In this embodiment, the prediction model is constructed by feature extraction, collection of historical attention patterns (head dimension distribution), recording of token (token) position access sequence, and statistics of each layer KV Cache hit rate. The input is the current attention score matrix, and the output is the KV Cache access probability of the future k tokens. The model size is limited within 1MB, and a lightweight prediction model is trained in this way.

[0105] S205, the intersection of the first hot data set, the second hot data set and the third hot data set is taken as a target hot data set.

[0106] S206, the KV Cache except the target hot data set is taken as a target cold data set.

[0107] S207, the target hot data set is stored to the GPU, and the target cold data set is stored to the memory.

[0108] Next, a large model inference method provided by an embodiment of the present application is introduced, Figure 6 A flowchart of a large model inference method provided by an embodiment of the present application can include:

[0109] S301, determining a hot data set in the fastest video memory device; wherein the hot data set is a set determined based on the above data division method.

[0110] The execution subject of this embodiment is an electronic device. In this embodiment, the fastest video memory device is the device that stores the hot data set.

[0111] S302, obtaining a target key value required for current inference from the hot data set, and inferring based on the target key value using a large language model.

[0112] The embodiment is not limited to a specific large language model, as long as the model needs to be based on the Key and Value of the historical token in the KV Cache for reasoning.

[0113] In the embodiment of the application, the hottest data stored in the fastest memory device is more accurate because the set is determined based on the above-mentioned data partitioning method, so that the Key and Value matrix of the historical token in the KV Cache is more accurate, thereby avoiding repeated calculation to the greatest extent and significantly reducing inference delay.

[0114] The data partitioning device provided by the embodiment of the application will be described below. The data partitioning device described below can be correspondingly referred to the data partitioning method described above.

[0115] Figure 7 The structure diagram of the data partitioning device provided by the embodiment of the application can include:

[0116] The access frequency determination module 100 is configured to determine the access frequency parameter based on the access frequency of each key-value cache in the sliding time window.

[0117] The time interval determination module 200 is configured to determine the time interval based on the access time corresponding to each key-value cache.

[0118] The future access probability determination module 300 is configured to obtain the future access probability of the key-value cache associated with each key-value cache by using a prediction model based on the historical access information corresponding to each key-value cache. The prediction model is a model for predicting based on the access association relationship between key-value caches.

[0119] The hot data cache module 400 is configured to partition the key-value cache based on the access frequency parameter, the time interval and the future access probability of each key-value cache, determine the hot data set and the cold data set, and store the key-value cache in the hot data set to the fastest memory device. The fastest memory device is the memory device during inference.

[0120] Further, based on the above-mentioned embodiment, the access frequency determination module 100 can include:

[0121] The access type determination unit is configured to determine the access type corresponding to each key-value cache, and the access type includes write and read.

[0122] The access frequency parameter determination unit is configured to determine an actual weight corresponding to the access frequency based on the access type, and determine the access frequency parameter based on the actual weight and the access frequency; wherein the weight of write access is greater than the weight of read access.

[0123] Further, based on any of the above embodiments, the time interval determination module 200 can include:

[0124] The preset number of time interval determination unit is configured to obtain a time queue storing access time of each key-value cache, and determine a preset number of time intervals based on the time queue.

[0125] The time interval determination unit is configured to determine a weight corresponding to each time interval, and determine the time interval based on each time interval and the weight corresponding thereto.

[0126] Further, based on any of the above embodiments, the data division device can further include:

[0127] The historical key-value cache obtaining module is configured to obtain historical key-value caches.

[0128] The attention head determination module is configured to determine an attention head corresponding to each of the historical key-value caches.

[0129] The target grouping historical key-value cache determination module is configured to group and store the historical key-value caches according to the attention heads, to obtain target grouping historical key-value caches.

[0130] The adjacent relationship between attention heads determination module is configured to determine adjacent target grouping historical key-value caches corresponding to each of the target grouping historical key-value caches, to obtain an adjacent relationship between attention heads.

[0131] The association relationship between historical key-value caches determination module is configured to determine an association relationship between historical key-value caches in two adjacent target grouping historical key-value caches based on the adjacent relationship between attention heads; wherein the association relationship is an influence of access of a current historical key-value cache on access of an associated historical key-value cache.

[0132] The prediction model training module is configured to train an initial prediction model based on the association relationship between historical key-value caches, to obtain the prediction model.

[0133] Further, based on any of the above embodiments, the hot data cache module 400 can include:

[0134] The first hot data set determination unit is configured to determine a first hot data set based on the access frequency parameter.

[0135] a second hot data set determination unit configured to determine a second hot data set based on the time interval;

[0136] a third hot data set determination unit configured to determine a third hot data set based on the probability of future access;

[0137] a hot data set determination unit configured to determine an intersection of the key-value cache in the first hot data set, the second hot data set and the third hot data set, and construct the hot data set based on the cached key-value in the intersection;

[0138] a cold data set determination unit configured to determine data other than the hot data set as the cold data set.

[0139] Further, based on any of the above embodiments, the data partitioning apparatus can further include:

[0140] a key-value cache prefetching module configured to determine a remaining storage space of the fastest GPU device, and prefetch key-value cache from the hot data set based on the probability of future access.

[0141] It should be noted that the order of the modules and units in the above data partitioning apparatus can be changed without affecting the logic.

[0142] Figure 7 The description of the features in the corresponding embodiments can be referred to Figure 7 The related description of the corresponding embodiments will not be repeated here.

[0143] The data division device provided by the embodiment of the application can include: an access frequency determination module 100, configured to determine the number of accesses of each key-value cache in a sliding time window, and determine an access frequency parameter based on the number of accesses; a time interval determination module 200, configured to determine the access time corresponding to each key-value cache, and determine a time interval based on the access time; a future access probability determination module 300, configured to use a prediction model to predict based on the historical access information corresponding to each key-value cache, to obtain the future access probability of the key-value cache associated with each key-value cache; wherein the prediction model is a model for predicting based on the access association relationship between key-value caches; a hot data cache module 400, configured to divide the key-value caches based on the access frequency parameter, the time interval and the future access probability of each key-value cache, to determine a hot data set and a cold data set, and to store the key-value caches in the hot data set to a memory device with the fastest speed, which is the memory device during inference. The application considers three parameters, namely the access frequency parameter, the time interval and the future access probability, wherein the access frequency is considered from a global perspective, the time interval is considered from a local time perspective, and the future access probability is considered from a local space perspective. Since the application divides cold and hot data from different perspectives, the accuracy of dividing cold and hot data is improved, so that the large model can obtain keys and values from the accurate hot data set in time during inference, and the inference speed of the large model is improved.

[0144] The large model inference device provided by the embodiment of the application is described below. The large model inference device described below can be correspondingly referred to the large model inference method described above.

[0145] Figure 8 The structure diagram of the large model inference device provided by the embodiment of the application can include:

[0146] The hot data set determination module 500 is configured to determine a hot data set in the memory device with the fastest speed; wherein the hot data set is a set determined based on the above data division method;

[0147] The inference module 600 is configured to obtain a target key-value required for current inference from the hot data set, and perform inference based on the target key-value using a large language model.

[0148] It should be noted that the order of the modules and units in the above large model inference device can be changed without affecting the logic.

[0149] Figure 8 The description of the features in the corresponding embodiment can be referred to Figure 8 the related description of the corresponding embodiment, which will not be repeated here.

[0150] The large model reasoning device provided by the embodiment of the present application can comprise: a hot data set determination module 500 configured to determine a hot data set in a memory device with the fastest speed; wherein the hot data set is a set determined based on the data division method described above; and a reasoning module 600 configured to obtain a target key value required for current reasoning from the hot data set and perform reasoning based on the target key value using a large language model. In the embodiment of the present application, since the set is determined based on the data division method described above, the hot data stored in the memory device with the fastest speed is more accurate, so the Key and Value matrix of the historical token cached by the KV Cache is more accurate, thereby avoiding repeated calculation to the greatest extent and significantly reducing reasoning delay.

[0151] Next, an electronic device provided by the embodiment of the present application is introduced. The electronic device described below can be correspondingly referred to the data division method and the large model reasoning method described above.

[0152] Figure 9 As shown in FIG. 1, the electronic device provided by the embodiment of the present application can comprise: a memory 60 configured to store a computer program; Figure 9

[0153] A processor 61 is configured to implement the steps of the data division method and / or the large model reasoning method of the above-described embodiments when executing the computer program.

[0154] The electronic device provided by the embodiment of the present application can include but is not limited to a smart phone, a tablet computer, a notebook computer or a desktop computer, etc.

[0155] ​The processor 61 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 61 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA), a central processing unit (CPU), and the like. The processor 61 can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also referred to as a central processing unit (CPU). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 61 can be integrated with a graphics processing unit (GPU) that is responsible for rendering and drawing content required to be displayed on the display screen. In some embodiments, the processor 61 can further include an artificial intelligence (AI) processor for processing computing operations related to machine learning.

[0156] The memory 60 can include one or more computer-readable storage media that can be non-transitory. The memory 60 can further include a high-speed random access memory, and a nonvolatile memory such as one or more disk storage devices, flash storage devices. In the present embodiment, the memory 60 is at least used to store the following computer program 601, wherein the computer program is loaded and executed by the processor 61, and can implement the related steps of the data partitioning method and / or the large model inference method disclosed in any of the preceding embodiments. In addition, the resources stored in the memory 60 can further include an operating system 602, data 603, and the like, and the storage manner can be temporary storage or permanent storage. The operating system 602 can include Windows, Unix, Linux, and the like.

[0157] In some embodiments, the electronic device can further include a display screen 62, an input / output interface 63, a communication interface 64, a power supply 65, and a communication bus 66.

[0158] Those skilled in the art can understand that the structure shown in the above embodiments does not constitute a limitation on the electronic device, and can include more or fewer components than those shown. Figure 9 The structure shown in the above embodiments does not constitute a limitation on the electronic device, and can include more or fewer components than those shown.

[0159] It can be understood that if the data partition method and / or large model inference method in the above embodiments are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and performs all or part of the steps of the methods of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), an electrically erasable programmable ROM, a register, a hard disk, a removable magnetic disk, a CD-ROM, a magnetic disk or an optical disk, and various media that can store program codes.

[0160] Based on this, the embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the data partition method and / or large model inference method are implemented.

[0161] The above describes a data partition method and / or large model inference method provided by the embodiment of the present application in detail. The embodiments in the specification are described in a progressive manner, and each embodiment mainly describes the differences from other embodiments. The same or similar parts of each embodiment can be referred to. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the related parts can be referred to the method part.

[0162] The skilled person can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of the two. In order to clearly show the interchangeability of hardware and software, the components and steps of each example have been described in the above description. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solutions. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0163] The above describes in detail the data division method and the large model inference method, the device, the equipment and the medium provided by the application. The principle and the implementation mode of the application are described by applying specific examples. The above description of the examples is only used to help understand the method of the application and the core idea. It should be pointed out that, for ordinary skilled persons in the technical field, some improvements and modifications can be made to the application without departing from the principle of the application, and these improvements and modifications also fall within the protection scope of the application.

Claims

1. A data partitioning method, characterized by, The method comprises: determining the access frequency parameter of each key-value cache based on the access frequency of each key-value cache in a sliding time window; determining the time interval based on the access time corresponding to each key-value cache; predicting the probability of future access of the key-value cache associated with each key-value cache based on the historical access information corresponding to each key-value cache using a prediction model; wherein the prediction model is a model for predicting based on the access association relationship between key-value caches; dividing the key-value caches based on the access frequency parameter, the time interval and the probability of future access of each key-value cache to determine the hot data set and the cold data set, so as to store the key-value caches in the hot data set to the fastest video memory device during inference.

2. The data partitioning method of claim 1, wherein, The method comprises: determining the access frequency parameter of each key-value cache based on the access frequency of each key-value cache in a sliding time window; determining the access type corresponding to each key-value cache, the access type including write and read; 3. The data partitioning method of claim 1, wherein, determining the actual weight corresponding to the access frequency based on the access type, and determining the access frequency parameter based on the actual weight and the access frequency; wherein the write access weight is greater than the read access weight. The method comprises: obtaining a time queue storing the access time of each key-value cache, and determining a preset number of time intervals based on the time queue; 4. The data partitioning method of any one of claims 1 to 3, wherein, determining the weight corresponding to each time interval, and determining the time interval based on each time interval and its corresponding weight. Before predicting the probability of future access of the key-value cache associated with each key-value cache based on the historical access information corresponding to each key-value cache using a prediction model, the method further comprises: obtaining historical key-value caches; determining the attention head corresponding to each historical key-value cache; grouping and storing the historical key-value caches according to the attention head to obtain target grouped historical key-value caches; determining the adjacent target grouped historical key-value cache corresponding to each target grouped historical key-value cache to obtain the adjacent relationship between the attention heads; determining the association relationship between each historical key-value cache in two adjacent target grouped historical key-value caches based on the adjacent relationship between the attention heads; wherein the association relationship is the influence of the access of the current historical key-value cache on the access of the associated historical key-value cache; 5. The data partitioning method of claim 1, wherein, training the initial prediction model based on the association relationship between each historical key-value cache to obtain the prediction model. The method comprises: determining the first hot data set based on the access frequency parameter; determining the second hot data set based on the time interval; determine a third hot data set based on the probability of future access; determine the intersection of key-value caches in the first hot data set, the second hot data set and the third hot data set, and construct the hot data set based on cached key-values in the intersection; data other than the hot data set is the cold data set.

6. The data partitioning method of claim 1, wherein, After dividing the key-value caches based on the access frequency parameter, the time interval and the probability of future access of each key-value cache, determining the hot data set and the cold data set, further comprising: determine the remaining storage space of the fastest GPU device, and pre-fetch key-value caches from the hot data set based on the probability of future access.

7. A large model inference method, comprising: Comprise: determine the hot data set in the fastest GPU device; wherein the hot data set is determined based on the data division method of any one of claims 1 to 6; get the target key-value required for current inference from the hot data set, and use a large language model to perform inference based on the target key-value.

8. A data partitioning apparatus characterized by comprising: Comprise: an access frequency determination module for determining the number of accesses of each key-value cache in a sliding time window, and determining the access frequency parameter based on the number of accesses; a time interval determination module for determining the access time corresponding to each key-value cache, and determining the time interval based on the access time; a future access probability determination module for predicting the probability of future access of key-value caches associated with each key-value cache based on historical access information corresponding to each key-value cache using a prediction model; wherein the prediction model is a model for predicting based on the access association relationship between key-value caches; a hot data cache module for dividing the key-value caches based on the access frequency parameter, the time interval and the probability of future access of each key-value cache, determining the hot data set and the cold data set, and storing the key-value caches in the hot data set to the fastest GPU device, which is the GPU device during inference.

9. An electronic device, comprising: Comprise: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium, and the computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.

Citation Information

Cited By

  • Big data stream processing method and system based on machine learning

    CN121957647A