A hardware system, method, and application for securely storing and restoring context information during the inference process of large models
Through self-developed hardware systems and slice reasoning methods, the context information volatile and data security problems in the process of large-model inference are solved, secure storage and rapid recovery are achieved, and inference efficiency is improved.
Patent Information
- Application Number
- CN202411033076.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-07-30
AI Technical Summary
The volatile problem of context information in the inference process of large models, and how to ensure data security is not leaked is a technical challenge that needs to be solved urgently.
It adopts a self-developed hardware system, communicates with the storage system through a USB interface, uses components such as ARM Cortex-M7 chip, DMA/FIFO module, SDMMC interface, etc., and combines USB 3.0 and SDIO 3.0 UHS-2 protocols to achieve secure storage and rapid recovery of K and V matrix data, and uses slice inference to reduce data loading delay.
It realizes the secure preservation of large-model context information, ensures user privacy without leakage, and significantly reduces data loading delay during recovery, improving inference efficiency by 18% to 20%.
Smart Images

Figure BDA0004970409830000071 
Figure BDA0004970409830000081 
Figure HDA0004970409840000011
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large model data security preservation and restoration, and relates to a system, method and application for securely preserving and restoring context information during the inference process of large models. Background Art
[0002] Currently, the hardware for inferring large models includes GPUs, NPUs, TPUs, etc. They all have a common feature of small on-chip storage, with typical capacities ranging from a few MB to several hundred MB, plus large-capacity DDR (Double Data Rate Synchronous Dynamic Random Access Memory) or HBM (High Bandwidth Memory), with typical storage capacities of dozens of GB. The algorithm frameworks of large models are usually based on the Transformer structure, with typical parameter quantities ranging from several B (billion) to several hundred B, and with the iteration and development of the models, their model contexts have also increased very significantly, from the initial context of 256 or 512 tokens to the current context of 200K, 300K or 500K tokens, showing a 1000-fold increase in just one year. The benefit of the increase in context is that the model can remember more historical interaction information during the inference process.
[0003] Based on the above software and hardware information, two problems arise during the inference process of large models: 1. During the inference process of the model, its context information is stored in the hardware's DDR or HBM. Once this type of hardware loses power, the information disappears. 2. The more context information the model has, the greater the significance of preserving it for later continued use, and it is also very important to ensure the security of the data without leakage. Therefore, there is an urgent need for a reliable and applicable technical solution to securely preserve the volatile context during the inference process of large models, reduce the latency of data loading, and at the same time protect user privacy. Summary of the Invention
[0004] In order to solve the deficiencies of the existing technology, the purpose of the present invention is to provide a hardware system, method and application for securely preserving and restoring context information during the inference process of large models. The technical solution in the present invention can solve the problem of context volatility during the inference process of large models, securely preserve the volatile context, and protect user privacy from leakage; in addition, when restoring context information, the data loading latency is significantly reduced through the tilling block loading method.
[0005] The present invention proposes a hardware system for securely preserving and restoring context information during the inference process of large models, and the hardware system includes: an inference host and a storage system;
[0006] The inference host communicates with the storage system through a USB interface;
[0007] The inference host includes: memory, central processing unit, hard disk, GPU / NPU / TPU module, etc.;
[0008] Among them, the memory is used to store temporary data and running programs, providing high-speed access; the central processing unit is responsible for executing computing tasks and is the core processing unit in the system; the hard disk is used to store data for a long time, including the operating system, application programs, user data, etc.; the GPU / NPU / TPU module includes a graphics processing unit, a neural network processing unit, and a tensor processing unit, which are used to accelerate the inference process of large models;
[0009] The storage system includes: ARM Cortex-M7 chip, DMA / FIFO module, SDMMC interface, USART interface, GPIO interface, etc.;
[0010] Among them, the ARM Cortex-M7 chip is used to handle control tasks and data management in the storage system; the I-Cache and D-Cache in the ARM Cortex-M7 chip are instruction cache and data cache, which accelerate data access, and JTAG / SWD are used for debugging and programming the ARM Cortex-M7, and BOOT and RST are used for system startup control and reset respectively;
[0011] The ARM Cortex-M7 chip is connected to FLASH, AXI SRAM, and SRAM through a bus. Among them, the FLASH provides persistent data storage to ensure that data is not lost in the event of a power outage, and the AXI SRAM and SRAM provide high-speed data storage and access, supporting complex computing tasks;
[0012] The USB interface of the inference host is connected to the USB switch in the storage system, and the USB switch then transmits data to the direct memory access / first-in-first-out buffer (DMA / FIFO) module through the USB high-speed physical layer; the USB switch can select the data transmission path, and the USB high-speed physical layer is responsible for the physical transmission of USB signals; the direct memory access / first-in-first-out buffer is used for high-speed data transmission;
[0013] The storage system mounts a Flash storage chip through the SDMMC interface to save the K and V matrix data in the large model inference process; after the data is processed by the DMA / FIFO module, it is transmitted to the Flash storage chip (SD FLASH) through the SDMMC interface for storage;
[0014] Data transmission is carried out between the storage system and the Flash storage chip or between the inference host and the storage system using the USB 3.0 and / or SDIO 3.0 UHS-2 protocol.
[0015] The ARM Cortex-M7 can control each module of the storage system and perform serial communication through the (Universal Synchronous / Asynchronous Receiver / Transmitter) USART interface; the USART interface can also be connected to a Print Log interface to print system logs;
[0016] The General-Purpose Input / Output (GPIO) interface can be connected to control external devices, including a USB Switch Check&Control module for checking and controlling the USB switch, an LED light for indicating the system status, etc.;
[0017] In use, the USB interface of the inference host is connected to the USB Switch of the storage system. The USB Switch transmits data to the DMA / FIFO module through the USB HS PHY. After being processed by the DMA / FIFO module, the data is transmitted to the SD FLASH for storage through the SDMMC interface;
[0018] Data is transmitted from the inference host to the USB Switch of the storage system through the USB interface. The USB Switch selects an appropriate transmission path and transmits the data through the USB HS PHY. The data enters the DMA / FIFO module for buffering and then is transmitted to the SD FLASH for storage through the SDMMC interface;
[0019] The ARM Cortex-M7 controls each module of the storage system and performs serial communication through the USART interface; the GPIO interface is used to control the LED and other external devices and monitor the status of the USB switch through the USB Switch Check&Control module.
[0020] When the hardware system of the present invention is specifically used, a USB driver is set on the inference host. The USB driver is the only external access interface. All access controls are processed through the USB driver, and illegal access and / or unauthorized access requests are intercepted; the judgment of whether the access is legal and / or authorized is realized by performing interactive verification through the set authorization ID during access.
[0021] When the hardware system switches contexts, it reloads the data saved in the Flash storage chip into the memory of the GPU / NPU / TPU module through the USB interface to optimize the data loading rate.
[0022] The hardware system adopts a method of slicing and inferring the K and V caches, enabling parallel data loading and inference calculations to reduce the inference latency caused by context switching.
[0023] The present invention also provides a method for saving and restoring context information during the inference process of a large model. The protection and restoration method includes the following steps:
[0024] Step 1: Obtain the historical K and V matrix data of the large language model and save it in the KV cache;
[0025] Step 2: Transmit the KV cache data to the storage system through the USB interface and save it in the Flash storage chip;
[0026] Step 3: When context switching occurs, reload the KV cache data from the Flash storage chip into the memory of the GPU / NPU / TPU module through the USB interface.
[0027] In Step 1, the capacity size of the KV cache is calculated by the following formula:
[0028] KV cache capacity size = 4000(max seq length) * 4096(dim) * 32(layer) * 2(K,V) * 2(bytes),
[0029] where max seq length represents the maximum sequence length, that is, the maximum number of tokens that the model can process during the inference process; dim represents the vector dimension of each token; layer represents the number of layers of the large language model; 2(K,V) represents that each of the K and V matrices occupies one storage space, for a total of 2; bytes represents the number of bytes occupied by each data item.
[0030] In Step 2, when it is necessary to save the current context information, the inference host prepares the data in the KV cache and transmits it to the external storage system through the USB interface; on the inference host, the USB driver is responsible for processing the data transmission request and performing security verification through the authorization ID; the FIFO module in the storage system temporarily caches the received data and then transmits it to the Flash storage chip through the SDMMC interface for persistent storage;
[0031] During the data transmission process, it passes through modules including a USB switch (USB Switch) and a USB high-speed physical layer (USB HSPHY) to ensure the stability and high speed of data transmission; the Flash storage chip can ensure that the context information is still available after power-off or system restart.
[0032] In step 3, when context switching occurs, the inference host sends a data recovery request to the storage system through the USB interface; the MCU (i.e., the central processing unit) in the storage system controls the SDMMC interface to read the KV cache data from the Flash storage chip, caches it through the FIFO module, and then transmits it back to the inference host through the USB interface; after receiving the data, the inference host loads the data into the memory of the GPU / NPU / TPU module to complete the context recovery for continued inference calculation.
[0033] During the operation of the method of the present invention, the method of slice inference is adopted to make data loading and inference calculation proceed in parallel to reduce the inference delay caused by context switching;
[0034] Specifically, when restoring context information, the KV cache data is sliced, and each slice contains partial K and V matrix data; the sliced data proceeds in parallel during loading and inference calculation, that is, while part of the data is being loaded, another part of the data is already performing inference calculation.
[0035] The present invention also provides a hardware device for implementing the above context information saving and method. The hardware device includes: a memory and a processor; a computer program is stored on the memory, and when the computer program is executed by the processor, the above context information saving and recovery method is implemented.
[0036] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above context saving and recovery method is implemented.
[0037] The present invention also provides the above hardware system, the above context saving and recovery method, the above hardware device, or the application of the above computer-readable storage medium in context information storage and extraction, etc.
[0038] The beneficial effects of the present invention include:
[0039] In the present invention, a self-developed hardware system is used to save the context information of the large model to ensure the security and non-disclosure of user information and privacy. The storage system in the present invention communicates with the inference host through the USB interface, and the USB driver on the inference host is the only external access interface. Illegal access or unauthorized access is intercepted in the driver program to protect the data in the Flash storage safely and effectively.
[0040] When restoring the inference context information, the slice inference method is adopted for the K, V cache (K, V cache) to make inference and loading proceed in parallel, thereby reducing the delay of the inference result.
[0041] The reason why large language models have memory function is that during the inference process, the model caches the K and V matrices in the attention mechanism. All historical information interacting with the large model is stored here. Therefore, saving the K and V matrices can preserve the historical token information during the user's inference process to prevent the loss of context information during the use of the large model. Usually, the K and V caches are temporarily stored in the DDR or HBM of the acceleration card, and once the hardware is powered off, the information is lost. Therefore, the present invention proposes a complete set of software and hardware solutions to permanently save the K and V cache matrices to the self-developed Flash, which not only avoids the loss of information but also ensures the security and controllability of the data.
[0042] Moreover, by performing sliced inference and saving on the K and V caches, it is possible to reduce the latency of the inference results caused by caching when restoring context information. As can be seen from the above data, this method has indeed achieved an 18% - 20% improvement in inference efficiency.
[0043] In the present invention, the self-developed Flash can solve the problem of context volatility during the inference process of large models, securely save the volatile context, and ensure that user privacy is not leaked. In addition, when restoring context information, the data loading latency is significantly reduced through the tilling block loading method. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.
[0045] Figure 1 It shows the hardware combination diagram of the storage system and the inference host in the present invention.
[0046] Figure 2 It shows the saving process diagram of the K and V caches of a 7B large language model.
[0047] Figure 3 It shows the sliced restoration process diagram of the K and V caches of a 7B large language model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] Combined with the following specific embodiments and drawings, the present invention will be further described in detail. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and well-known common sense in the art, and the present invention has no special limitations.
[0049] As the application of large models based on Transformer becomes more and more popular, the interaction between users and large models is becoming increasingly frequent, and the personal privacy and data security of users during the interaction with large models are also attracting more and more attention. The present invention uses an independently developed hardware system to quickly save and restore the context information during the inference process of the large model, providing strong software and hardware protection measures for the preservation and restoration of personal privacy information.
[0050] The present invention provides a hardware system that supports the secure preservation and restoration of context information during the inference process of a large model. The hardware system includes: an inference host and a storage system;
[0051] The inference host communicates with the storage system through a USB interface;
[0052] The inference host includes: a memory, a central processing unit, a hard disk, and a GPU / NPU / TPU module;
[0053] The storage system includes: an ARM Cortex-M7 chip, a DMA / FIFO module, an SDMMC interface, a USART interface, and a GPIO interface;
[0054] The storage system mounts a Flash storage chip through the SDMMC interface to save the K and V matrix data during the inference process of the large model;
[0055] Data is transmitted between the storage system and the Flash storage chip or between the inference host and the storage system using the USB 3.0 and / or SDIO 3.0 UHS-2 protocol.
[0056] The present invention also provides a method for preserving and restoring context information during the inference process of a large model, including the following steps:
[0057] Step 1: Obtain the historical K and V matrix data of the large language model and save it in the KV cache;
[0058] Step 2: Transmit the KV cache data to the storage system through the USB interface and save it in the Flash storage chip;
[0059] Step 3: When context switching occurs, reload the KV cache data from the Flash storage chip into the memory of the GPU / NPU / TPU module through the USB interface.
[0060] In the present invention, the method for slice inference of the K and V caches is specifically as follows:
[0061] 1. Data slicing:
[0062] Slicing Logic: First, slice the K and V matrix data in the KV cache according to preset rules. Each slice contains a part of the K and V matrix data, and these data blocks are divided according to the preset size and order.
[0063] Slice Size: The size of the slice can be adjusted according to the specific model and hardware configuration to ensure that the data volume of each slice is appropriate for parallel processing.
[0064] The slicing rules, slice size, order, etc. can all be set according to actual needs to meet actual usage.
[0065] 2. Data Loading:
[0066] Parallel Loading: When context switching, reload the KV cache data from the Flash storage chip into the memory of the inference host through the USB interface. To improve efficiency, this step loads the sliced data piece by piece, that is, while loading one slice, it can start loading the next slice.
[0067] Pipeline Operation: Adopt pipeline operation to ensure the continuity of data during loading. For example, while loading the first slice, the data of the second slice can be prepared, and immediately start loading the second slice after the first slice is loaded.
[0068] 3. Parallel Computing:
[0069] Slice Inference: Once the data of a slice is loaded into the memory of the inference host, start the inference calculation immediately. The GPU / NPU / TPU module of the inference host can perform inference calculation on the loaded data while loading the next slice.
[0070] Parallel Processing: This parallel processing method enables data loading and inference calculation to be carried out simultaneously, making the most of hardware resources and reducing waiting time.
[0071] 4. Data Transmission and Caching:
[0072] High-Speed Transmission: During slice inference, data is transmitted through the USB 3.0 and / or SDIO 3.0 UHS-2 protocol to ensure the high speed and stability of data transmission.
[0073] FIFO Caching: During data transmission, the data passes through the FIFO (First In First Out) module for caching to ensure that there is no bottleneck in the data transmission process and maintain the continuity of the data stream.
[0074] 5. Inference Result Merging:
[0075] Result merging: After the slice inference is completed, the inference results of each slice are merged to form a complete inference output. This step involves concatenating the calculation results of multiple slices in sequence to ensure the integrity and accuracy of the inference results.
[0076] The complete specific technical solution of the present invention is introduced in two parts:
[0077] I. Regarding how to save the volatile context during the large model inference process and ensure data security and non-disclosure
[0078] 1. Data source: For a large language model based on Transformer, its context information is actually the historical K, V matrices, which are obtained by multiplying the input matrix with the two weight matrices Wk and Wv. All historical K, V matrix values are saved in the K, V cache. The capacity of the K, V cache determines the length of the context information that the model can save. Based on the model parameters specified in the following table, the capacity size of the K, V cache can be calculated:
[0079]
[0080] The capacity size of the K, V cache is calculated by the following formula:
[0081] ① 4000 (max seq length) * 4096 (dim) * 32 (layer) * 2 (K, V) * 2 (bytes) = 1.95GB
[0082] ② 4000 (max seq length) * 5120 (dim) * 40 (layer) * 2 (K, V) * 2 (bytes) = 3.05GB
[0083] 2. How to save securely and ensure data non-disclosure
[0084] If the data needs to be saved for a long time, the quickest way is the hard disk. The present invention also provides a hardware system that supports the secure storage of context information during the large model inference process. In the hardware system, the inference host and the storage system communicate through USB. The storage system mainly consists of an ARM Cortex-M7 chip, a DMA / FIFO module, an SDMMC interface, a USART interface, a GPIO interface, etc. The Flash storage chip is mounted on the system through the SDMMC interface. To maximize the data transfer bandwidth as much as possible, this system selects USB 3.0 and / or SDIO 3.0 UHS-2 as the communication protocols between the inference host and the storage system and between the storage system and the Flash chip, and their bandwidth speeds are 500MB / S and 312MB / S respectively.
[0085] In the above system, the data is finally stored in the Flash storage chip, and its only access method is through the USB interface of the inference host. All access controls can be processed through the USB driver, and illegal access requests can be intercepted.
[0086] II. Regarding how to quickly restore the inference context and significantly reduce the latency of data loading during the large model inference process
[0087] During the large model inference process, it is necessary to reload the corresponding K, V cache data to switch the context. In the method of the present invention, the secure data stored in the Flash will be reloaded into the memory of the GPU / NPU / TPU module. From the perspective of the entire loading process, the speed bottleneck is on the SDIO bus, so the overall data transfer rate is 312MB / S.
[0088] Considering the two major factors affecting the model inference latency, one is data loading and the other is data operation time. During the process of re-switching the context for a new round of dialogue inference, this is the prefill stage of a large model, which belongs to the calculation-limited type, that is, the computing power bottleneck stage, and its operation time is related to the computing power of the acceleration card.
[0089] In order to minimize the impact of inference latency caused by context switching, the present invention will adopt the method of slicing inference for the K, V cache, so that calculation and loading can be parallelized. In this way, the overall latency time of context switching can be obtained by the following formula ②. If the slicing inference method is not adopted, the K, V cache loading time and the model inference time are completely serial, and its overall latency time will be greater, which can be calculated by the following formula ①.
[0090] ① Context switching latency = k, v cache loading time + inference time
[0091] ② Context switching latency = max(k, v cache loading time, inference time)
[0092] The following table summarizes the latency comparison of the method of the present invention under different parameter models
[0093] Note: The following data tests are sourced from dual NVIDIA A100 acceleration cards
[0094]
[0095] As can be seen from the above table, in the method adopted in this invention, for the sliced inference method of K, V Cache, the context switching latency is improved by 18% in the 7B model solution and 20% in the 13B model solution.
[0096] 1. Connect to the self-developed hardware circuit board, such as Figure 1 The figure shown is the framework diagram of the self-developed circuit board. Currently, there is no ready-made circuit board that meets the requirements, so it must be self-developed. The self-developed hardware circuit board in this invention can store data securely and is self-controlled and reliable. Connect to the USB port of the inference host. The box for installing the circuit board has a 9-pin USB female port.
[0097] 2. Connect the UART port of the self-developed hardware circuit board to the host (optional).
[0098] 3. Download the firmware program through the debug port to drive the Arm CM7 core of the circuit board.
[0099] 4. Install the USB driver program on the inference host.
[0100] At this point, the software and hardware environment is configured. Then, an authorization ID needs to be obtained offline. This ID needs to be given as a parameter to the interface when calling the interface of the USB driver program for credit verification. When the user needs to access the storage system, the following two interfaces need to be called to read / write the storage system respectively.
[0101] int_32 write_data_via_usb(int_32 id, int_32 ctx_id, void* data, uint_32 len) / / This interface writes data to the storage system
[0102] Where id is the authorization ID, ctx_id is the ID value of this data, and the following parameters are the data pointer and data length. If the data operation is successful, it returns 0; otherwise, it returns 1.
[0103] int_32 read_data_via_usb(int_32 id, int_32 ctx_id, void(*call_back)(void** data)) / / This interface reads data from the storage system. It is an asynchronous interface. After reading part of the data, the call_back function will be called to process the data.
[0104] In this invention, it is through call_back to achieve partial loading of the sliced K, V cache data;
[0105] Figure 2 and Figure 3Describes respectively the storage paths of the K, V Cache and the data flow of slice recovery during the inference process of a model with a 7B - sized parameter.
[0106] Figure 2 It is a schematic diagram of the slice storage process of the K, V cache of a 7B large - language model. Specifically,
[0107] I. Data generation and storage (inference host part):
[0108] Step 1: In the inference host, the GPU / NPU / TPU module is responsible for the inference process of the large model. The generated historical K, V matrix data during the process will be saved in the memory of the inference host.
[0109] Inference host components:
[0110] Central processing unit: Responsible for executing inference tasks and coordinating data transmission;
[0111] Memory: Temporarily stores the generated K, V matrix data during the inference process;
[0112] Hard disk: Stores data for a long time, including the operating system, application programs, and user data;
[0113] GPU / NPU / TPU: Accelerates the inference process of the large model.
[0114] II. Data transmission:
[0115] Step 2: Through the USB interface, the K, V matrix data saved in the inference host memory is transmitted to the FIFO module in the storage system. This step ensures high - speed data transmission, with a transmission rate of 500MB / s (USB 3.0) or 312MB / s (SDIO 3.0 UHS - 2).
[0116] III. Data storage (storage system part):
[0117] Step 3: After the FIFO module receives the data, the data is transmitted through the SDMMC interface to the Flash storage chip for persistent storage. The storage system includes the following components:
[0118] FIFO: Temporarily caches the transmitted data to ensure data stability and continuity;
[0119] Flash: Persistently stores the K, V matrix data to ensure that the data will not be lost in case of power failure or system restart;
[0120] SDMMC: Responsible for data transmission between the FIFO and the Flash storage chip;
[0121] MCU: Controls the operations of the entire storage system, including data transfer and storage.
[0122] Figure 3 It is a schematic diagram of the slice recovery process of the 7B large language model K, V cache. Specifically,
[0123] I. Data Reading (Storage System Part):
[0124] Step 1: When context information needs to be restored, the inference host sends a data recovery request to the storage system through the USB interface. The MCU in the storage system controls the SDMMC interface to read the K, V cache data from the Flash storage chip. These data are divided into multiple small slices and are cached through the FIFO module and prepared for transmission.
[0125] Storage System Components:
[0126] FIFO: Temporarily caches the data read from the Flash storage chip to ensure the stability and continuity of the data;
[0127] Flash: Persistently stores the K, V matrix data to ensure that the data will not be lost after power-off or system restart;
[0128] SDMMC: Responsible for reading data from the Flash storage chip;
[0129] MCU: Controls the operations of the storage system, including data reading and transmission;
[0130] II. Data Transfer:
[0131] Step 2: Transfers the cached data slices back to the inference host through the USB interface. In a specific embodiment, the total data volume is 40 slices of 50MB each, and a total of 2GB of data needs to be transferred. This step ensures high-speed data transfer, and the transfer rate can reach 500MB / s (USB 3.0) or 312MB / s (SDIO 3.0 UHS-2).
[0132] III. Data Loading and Inference (Inference Host Part):
[0133] Step 3: After receiving the data slices, the inference host loads these data into the memory. The data in the memory is then transferred to the inference acceleration card memory of the GPU / NPU / TPU module for preparing inference calculations.
[0134] Inference Host Components:
[0135] Central Processing Unit: Responsible for coordinating the execution of data transfer and inference tasks;
[0136] Memory: Temporarily stores the transferred data slices;
[0137] Hard disk: Long-term data storage, including operating systems, application programs, and user data;
[0138] GPU / NPU / TPU: Accelerate the inference process of large models.
[0139] Aspects of the present invention are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0140] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which causes a computer, a programmable data processing apparatus, and / or other devices to operate in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0141] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.
[0143] The protection scope of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept, variations and advantages that can be conceived by those skilled in the art are included in the present invention, and the scope of protection is defined by the appended claims.
Claims
1. A hardware system that supports the secure preservation and restoration of context information during the inference process of large models, characterized in that The hardware system performs sliced inference on the K and V caches, enabling parallel data loading and inference calculations, and reducing the inference latency caused by context switching. The capacity of the KV cache is calculated using the following formula: KV cache capacity = 4000 (max seq length) * 4096 (dim) * 32 (layer) * 2 (K, V) * 2 (bytes), where max seq length represents the maximum sequence length, i.e., the maximum number of tokens that the model can process during inference; dim represents the vector dimension of each token; layer represents the number of layers of the large language model; 2 (K, V) indicates that the K and V matrices each occupy one storage space; and bytes represents the number of bytes occupied by each data item. The system includes: an inference host and a storage system; The inference host communicates with the storage system through a USB interface. A USB driver is set on the inference host, which is the only external access interface. All access controls are processed through the USB driver, and illegal access and / or unauthorized access requests are intercepted. The legality of access and / or whether authorization has been obtained is determined by performing interactive verification using the authorization ID during access; The inference host includes: memory, a central processing unit, a hard disk, and a GPU / NPU / TPU module; The storage system includes: an ARM Cortex-M7 chip, a DMA / FIFO module, an SDMMC interface, a USART interface, and a GPIO interface; The storage system mounts a Flash storage chip through the SDMMC interface to persistently store the K and V matrix data during the inference process of the large model, ensuring that the context information remains available after a power outage or system restart; The storage system and the Flash storage chip or the inference host and the storage system use the USB3.0 and / or SDIO 3.0 UHS-2 protocol for data transmission; The USB interface of the inference host is connected to the USB Switch of the storage system. The USB Switch transmits the data to the DMA / FIFO module through the USB HS PHY. After the data is processed by the DMA / FIFO module, it is transmitted to the SDFLASH for storage through the SDMMC interface.
2. The hardware system according to claim 1, characterized in that When the hardware system performs context switching, it reloads the data saved in the Flash storage chip into the memory of the GPU / NPU / TPU module through the USB interface to optimize the data loading rate.
3. A method for saving and restoring context information during the inference process of a large model based on the hardware system described in claim 1 or 2, characterized in that, The method uses the sliced inference method to enable parallel data loading and inference calculations, reducing the inference latency caused by context switching. It includes the following steps: Step 1: Obtain the historical K and V matrix data of the large language model and save it in the KV cache; The capacity of the KV cache is calculated using the following formula: KV cache capacity = 4000 (max seq length) * 4096 (dim) * 32 (layer) * 2 (K, V) * 2 (bytes), where max seq length represents the maximum sequence length, i.e., the maximum number of tokens that the model can process during inference; dim represents the vector dimension of each token; layer represents the number of layers of the large language model; 2 (K, V) means that the K and V matrices each occupy one storage space; bytes represents the number of bytes occupied by each data item; Step 2: Transmit the KV cache data to the storage system through the USB interface and save it in the Flash storage chip; the Flash storage chip can ensure that the context information is still available after power-off or system restart; Step 3: When switching contexts, reload the KV cache data from the Flash storage chip into the memory of the GPU / NPU / TPU module through the USB interface; among them, the MCU in the storage system controls the SDMMC interface to read the KV cache data from the Flash storage chip, and these data are divided into multiple small slices and transmitted after being cached by the FIFO module.
4. The method according to claim 3, characterized in that, In Step 2, when it is necessary to save the current context information, the inference host prepares the data in the KV cache and transmits it to the external storage system through the USB interface; on the inference host, the USB driver is responsible for processing the data transmission request and performing security verification through the authorization ID; the FIFO module in the storage system temporarily caches the received data and then transmits it to the Flash storage chip through the SDMMC interface for persistent storage; and / or, During the data transmission process, it passes through a USB switch and a USB high-speed physical layer module to ensure the stability and high speed of data transmission.
5. The method according to claim 3, wherein In Step 3, when switching contexts, the inference host sends a data recovery request to the storage system through the USB interface; the KV cache data is transmitted back to the inference host through the USB interface; after receiving the data, the inference host loads the data into the memory of the GPU / NPU / TPU module to complete the context recovery for continuing the inference calculation.
6. A hardware device for implementing the method according to any one of claims 3-5, characterized in that, The hardware device includes: a memory and a processor; a computer program is stored on the memory, and when the computer program is executed by the processor, the context information saving and restoring method according to any one of claims 3-5 is implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the context saving and restoring method according to any one of claims 3-5 is implemented.