A GPU inference fast migration method and system based on model segmentation
Patent Information
- Application Number
- CN202511488460.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-10-17
AI Technical Summary
[0003]现有的基于智能推理加速处理器的智能推理服务的迁移方案包括显存快照、API拦截-重放,以及服务级日志,但都难以满足快速恢复推理服务的需求:显存快照方案通过存储整块GPU内存及相关上下文,并将这些信息写入另一节点实现迁移,但需要传输的数据量达到GB级,因此跨节点迁移耗时较长;API拦截-重放方案若不额外进行显存状态同步则需从花费大量时间重新推理以恢复状态;类似的,服务级日志同样需要在迁移后耗费大量时间重新推理完整模型
[0028]The beneficial effects of this invention are as follows: This invention aims to solve the problem of long migration time for intelligent inference services, which makes it difficult to quickly restore services. Existing technologies are not suitable for intelligent inference service scenarios, or require a long time to synchronize or restore service status, resulting in difficulties in solving the following problems: Existing CPU process checkpointing/recovery technology cannot capture the state of intelligent inference accelerator processors and is not suitable for intelligent inference service scenarios; Existing GPU memory snapshot technology requires saving and restoring the state of checkpoint files with GB-level data volume, resulting in long migration and transmission times and making it difficult to quickly restore services; Existing API interception-replay technology and service-level logging technology require repeating the entire inference process after migration, resulting in excessive service recovery time.
Smart Images

Figure CN121683998B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a fast GPU inference migration method and system, and more particularly to a fast GPU inference migration method and system based on model segmentation, belonging to the field of GPU inference migration technology. Background Technology
[0002] To address the computational demands of complex models, intelligent inference accelerators, represented by GPUs (Graphics Processing Units), have gradually become the mainstream hardware platform for intelligent inference services. However, if an inference node becomes inoperable due to hardware failure or program anomalies, all ongoing inference tasks will be interrupted for an extended period, making it difficult to meet the high-availability inference scenarios requiring rapid recovery, such as those involving unmanned equipment. Therefore, to reduce the impact of inference node failures on business operations, it is necessary to shorten the time required for cross-node migration and state recovery of inference services.
[0003] Existing migration solutions for intelligent inference services based on intelligent inference accelerators include memory snapshots, API interception-replay, and service-level logs. However, none of these solutions meet the need for rapid recovery of inference services. Memory snapshot solutions migrate by storing the entire GPU memory block and related context, then writing this information to another node. However, the amount of data transferred reaches gigabytes, resulting in lengthy cross-node migration times. API interception-replay solutions, without additional memory state synchronization, require significant time to re-infer the state. Similarly, service-level logs also require substantial time to re-infer the complete model after migration. Therefore, existing intelligent inference service migration solutions are generally time-consuming, necessitating a rapid migration technology to shorten the cross-node migration and state recovery time for inference services.
[0004] Existing technology one, CPU process checkpointing / restore technology, can serialize the complete state of a running CPU process (including registers, memory mappings, open file descriptors, etc.) into a disk file without modifying the application source code, and rebuild the process in-place or on another host with the same kernel and architecture to achieve migration. A representative work in this field, CRIU (Checkpoint / Restore In Userspace), has been widely used in traditional server disaster recovery and load balancing scenarios. Based on CRIU, researchers have further proposed various optimization schemes, such as VAS-CRIU, which uses copy-on-write virtual address space mapping, and Catalyzer, which reduces I / O latency through a lazy / on-demand recovery strategy. However, existing technology one is designed for general-purpose CPU processes and cannot capture the context and computational state of intelligent inference accelerators (such as GPUs and NPUs), therefore it is not suitable for intelligent inference service environments. Existing technology two, GPU device memory state saving / restoration technology (such as cuda-checkpoint), stores the entire GPU memory and internal objects (context) at the driver layer. The first method involves dumping checkpoint data (such as streams and events) into a checkpoint file in one go, and then writing it back to GPU memory as is during recovery, thus significantly shortening fault recovery time. However, the checkpoint data file generated by the second method can be several gigabytes in size, requiring a long transmission time to synchronize the state across nodes. This causes significant transmission time during cross-node migration, making it difficult to use for rapid migration of intelligent services. In addition, the second method requires that the source and target nodes have completely identical GPU models, driver versions, and PCIe topologies, further limiting its applicability to NVIDIA GPUs and making it difficult to adapt to other types of intelligent accelerator processors. The third method, API interception-replay technology (such as Cricket, Singularity, etc.), inserts an interception layer between the CUDA / ROCm runtime and the driver, recording the GPU data line by line. API call logs are used and replayed during recovery to reconstruct the computation state. However, the interception layer must record every API call and replay all API calls after migration to re-execute the complete inference process. The execution state can only be restored by reloading and inferring the complete model after migration, which causes significant waste and slows down the state recovery time. In addition, API interception introduces continuous runtime overhead and requires recompiling the software stack to enable the interception layer, which is highly intrusive to the environment. Existing technology 3 also has problems such as only being compatible with mainstream foreign GPUs and related drivers and requiring customized software environments.Existing technology four, the service-level log recovery mechanism, records state information to a log file after each inference or mini-batch execution. When the main inference node fails, the management component can restore the service by reading the latest log entry on the new node and re-executing, without needing to concern itself with low-level details such as registers and memory mapping. Existing technology four is widely used in stateless container clusters represented by Kubernetes and inference frameworks such as TensorFlow Serving and PyTorchTorchServe. However, this coarse-grained reconstruction method requires reloading the complete model weights and re-executing inference after a failure, resulting in slow inference recovery.
[0005] In summary, a fast GPU inference transfer method and system based on model segmentation is needed. Summary of the Invention
[0006] A brief overview of the invention is given below to provide a basic understanding of certain aspects of it. It should be understood that this overview is not an exhaustive summary of the invention. It is not intended to identify key or essential parts of the invention, nor is it intended to limit the scope of the invention. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description that follows.
[0007] In view of this, in order to solve the problems of slow inference recovery speed and narrow applicability of traditional GPU inference fast migration methods and systems in the prior art, the present invention provides a GPU inference fast migration method and system based on model segmentation.
[0008] Technical solution one is as follows: A fast GPU inference transfer method based on model segmentation, comprising the following steps:
[0009] S1. Design software components for inference devices in a GPU inference fast migration system;
[0010] S2. Design corresponding model segmentation mechanisms for typical inference models;
[0011] S3. Employ software components and a model segmentation mechanism to design a service migration process, enabling rapid migration of GPU inference.
[0012] Furthermore, in S1, the software components of the inference device in the GPU inference fast migration system include a migration control component, an inference interface component, an inference run manager, and a state management component. That is, the main inference device is equipped with a main migration control component, a main inference interface component, a main inference run manager, and a main state management component, and the backup inference device is equipped with a backup migration control component, a backup inference interface component, a backup inference run manager, and a backup state management component.
[0013] The migration control component is used to continuously monitor the node's operating status and, when the node currently providing inference services fails, call other components to perform service and route switching.
[0014] The inference interface component is used to receive external inference requests and forward them to the inference run manager. After the inference request is processed, the inference result is returned.
[0015] The inference run manager is used to execute the model inference process, and it includes an inference job queue, a model and block repository, and pre- and post-processing modules.
[0016] The state management component is used to realize service state saving and inference state synchronization between the primary inference device and the backup inference device. The backup state management component includes a state recovery module and a state synchronization module, and the primary state management component includes a state saving module and a state synchronization module.
[0017] Furthermore, in S2, for the CNN model, while avoiding splitting the fully connected layers, the splitting point of the CNN model is delayed, so that the execution time of other blocks except the last block n is close to the preset threshold.
[0018] For the Transformer model, it is divided into two parts: encoders and decoders. If the execution time of these two parts is greater than a preset threshold, the model is split proportionally between the encoder layer of the encoders and the decoder layer of the decoders until the execution time of the block is lower than the preset threshold.
[0019] For the LSTM model, it is divided into equal parts until the execution time of each block is lower than a preset threshold.
[0020] After the model is split, the start time of execution of each model block is set as the checkpoint.
[0021] Furthermore, in S3, the service migration process includes two stages: normal operation and service migration. The specific steps are as follows:
[0022] S31. Before the normal operation phase, elect a master node and a backup node, and bind the cluster virtual IP to the master node;
[0023] S32. During normal operation, the main inference interface component forwards the inference requests sent by the user to the main inference run manager. It executes model chunks for long requests and complete models for short requests. It parses the tasks into jobs and puts them into the inference job queue. If a long request is executed, before the master node executes a chunked inference job, the main state management component packages the inference job queue and intermediate inference results into an inference log and synchronizes it with the checkpoint to the standby node. Then the main inference manager loads the required model or chunk and executes the inference. The standby inference run manager synchronously preloads the required model or chunk but does not execute it. The above strategy process is repeated until the model inference is completed. If a short request is executed, the main state management component only synchronizes the inference log before the inference begins. Finally, when the model inference for both long and short requests is completed, the main inference interface component synchronizes the inference job queue again and returns the inference result to the user.
[0024] S33. During the service migration phase, the standby migration control component notifies other components of the standby inference device to rebuild the inference state. Upon receiving the notification, the standby inference interface component binds the cluster virtual IP to the current node and temporarily stores subsequent requests in the local cache. The standby state management component restores the CPU state and inference job queue based on the most recently received checkpoint and re-infers from the intermediate results contained in the checkpoint to restore the GPU state. After the inference state reconstruction is completed, the standby inference interface component releases the cached inference requests in sequence. Then the standby node enters the normal operation phase and continues to provide services as the new master node.
[0025] Technical Solution 2 is as follows: A GPU inference fast migration system based on model segmentation, used to execute the GPU inference fast migration method based on model segmentation described in Technical Solution 1, including a main inference device, a backup inference device, and a network switch;
[0026] The main inference device is connected to the backup inference device, the main inference device is connected to the network switch, and the backup inference device is connected to the network switch.
[0027] Both the primary inference device and the backup inference device include an intelligent inference acceleration processor, memory, and CPU.
[0028] The beneficial effects of this invention are as follows: This invention aims to solve the problem of long migration time for intelligent inference services, which makes it difficult to quickly restore services. Existing technologies are not suitable for intelligent inference service scenarios, or require a long time to synchronize or restore service status, resulting in difficulties in solving the following problems: Existing CPU process checkpointing / recovery technology cannot capture the state of intelligent inference accelerator processors and is not suitable for intelligent inference service scenarios; Existing GPU memory snapshot technology requires saving and restoring the state of checkpoint files with GB-level data volume, resulting in long migration and transmission times and making it difficult to quickly restore services; Existing API interception-replay technology and service-level logging technology require repeating the entire inference process after migration, resulting in excessive service recovery time.
[0029] This invention addresses intelligent inference scenarios and proposes a fast migration method and system for GPU inference based on model segmentation. Compared to existing methods, this invention can shorten the time required for cross-node migration and state recovery of inference services. Specifically:
[0030] 1) The GPU inference fast migration system designed in this invention consists of two inference nodes containing GPUs and CPUs with the same architecture and instruction set. The software part includes a migration control component, an inference interface component, an inference run manager and a state management component. The system divides models with long execution times into several model blocks according to their structure, and synchronizes CPU checkpoints and inference logs after each model or model block is executed, thereby achieving fast recovery through fine-grained saving of inference state.
[0031] 2) This invention proposes a method for selecting segmentation strategies based on typical model types (CNN, Transformer, LSTM). By analyzing the computational characteristics and data transfer volume of model layers, the location of segmentation points and the number of blocks are determined, thereby reducing the amount of intermediate result data and improving state synchronization efficiency. During execution, the CNN model delays the segmentation point as much as possible while avoiding segmentation of fully connected layers; the Transformer model preferably segments the encoder and decoder parts, and if the length after segmentation still exceeds the threshold, it segments proportionally between the encoder / decoder layers; while the LSTM model segments proportionally between LSTM units.
[0032] 3) This invention designs a service migration process, which includes a normal operation phase and a service migration phase. During the normal operation phase, the master node continuously synchronizes inference logs and checkpoints to the backup node. When a failure occurs, the backup node only needs to restore the most recent checkpoint to continue inference, achieving seamless service switching for users. Furthermore, the inference run manager first distinguishes inference tasks into long requests and short requests during scheduling. That is, short requests are directly processed as complete models, while long requests are split into several model blocks according to the aforementioned segmentation strategy. The execution job of each block is added to the inference job queue and executed sequentially. Whenever a model block is completed, the state saving / recovery module immediately synchronizes the current CPU state and inference logs to the backup node, shortening the service recovery time. Attached Figure Description
[0033] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:
[0034] Figure 1 This is a flowchart illustrating a fast transfer method for GPU inference based on model segmentation.
[0035] Figure 2 This is a schematic diagram of the structure of a GPU inference fast transfer system based on model segmentation;
[0036] Figure 3 This is a schematic diagram of the structural connections of an embodiment of a GPU inference fast migration system based on model segmentation;
[0037] Figure 4 A schematic diagram of the software component structure of an inference device;
[0038] Figure 5 The diagrams are schematic diagrams of model segmentation mechanisms, where (a) is a schematic diagram of the model segmentation mechanism of the CNN model, (b) is a schematic diagram of the model segmentation mechanism of the Transformer model, and (c) is a schematic diagram of the model segmentation mechanism of the LSTM model.
[0039] Figure 6 A timeline diagram of the service migration process.
[0040] Reference numerals: 1. Main inference device; 2. Backup inference device; 3. Network switch. Detailed Implementation
[0041] To make the technical solutions and advantages of the embodiments of the present invention clearer, the exemplary embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0042] Example 1: Reference Figures 1-6 This embodiment describes a fast GPU inference transfer method based on model segmentation, which specifically includes the following steps:
[0043] S1. Design software components for inference devices in a GPU inference fast migration system;
[0044] S2. Design corresponding model segmentation mechanisms for typical inference models to reduce the storage of intermediate results and improve the efficiency of state synchronization;
[0045] S3. Employ software components and a model segmentation mechanism to design a service migration process, enabling rapid migration of GPU inference.
[0046] Specifically, this invention is executed by two inference devices, a primary and a backup: during normal operation, the primary node executes inference requests, synchronizes checkpoints after short requests are completed, executes long requests in blocks, and synchronizes intermediate results and checkpoints after each block is completed; when the primary node fails, the backup node restores the service status according to the most recent checkpoint and migrates the cluster service IP to the backup node.
[0047] To address the issue of excessively long recovery times in existing inference service migration methods on intelligent inference accelerators, this invention provides a GPU inference fast migration system and method based on model segmentation. The redundant intelligent inference device comprises hardware and software components. The hardware component includes two inference devices connected by an internal network; the software component includes a migration control component, an inference interface component, an inference runtime manager, and a state management component. The fast inference service migration method consists of a model segmentation point selection strategy and a service migration process. The model segmentation process divides the original model with excessively long execution times into several model blocks whose execution times do not exceed a specified threshold based on their structure. The service migration process includes two phases: a normal operation phase for model / model block inference and state synchronization, and a service migration phase for state recovery and service migration.
[0048] Furthermore, in S1, reference Figure 4The software components of the inference device in the GPU inference fast migration system include a migration control component, an inference interface component, an inference run manager, and a state management component. Specifically, the main inference device is equipped with a main migration control component, a main inference interface component, a main inference run manager, and a main state management component, while the backup inference device is equipped with a backup migration control component, a backup inference interface component, a backup inference run manager, and a backup state management component.
[0049] The migration control component is used to continuously monitor the node's operating status and, when the node currently providing inference services fails, call other components to perform service and route switching.
[0050] The inference interface component acts as a bridge between the inside and outside of the system, receiving external inference requests and forwarding them to the inference run manager. After the inference request is processed, the inference result is returned.
[0051] The inference run manager is used to execute the model inference process, and it includes an inference job queue, a model and block repository, and pre- and post-processing modules.
[0052] The state management component is used to realize service state saving and inference state synchronization between the primary inference device and the backup inference device. The backup state management component includes a state recovery module and a state synchronization module, and the primary state management component includes a state saving module and a state synchronization module.
[0053] Specifically, the migration control component periodically sends heartbeat signals to the standby and primary inference devices and periodically probes the heartbeat information of the primary and standby inference devices. If no heartbeat information is received within a timeout, other components are invoked to migrate the service and cluster virtual IP to the current node.
[0054] Preferably, the inference interface component works as an RPC server and provides a unified virtual IP through KeepAlived. In addition to forwarding requests, the inference interface component is also responsible for virtual IP switching and request storage. That is, after receiving the service switching signal from the migration control component, the inference interface component performs virtual IP migration and stores the newly arrived inference requests. After receiving the ready signals from other components, the inference interface component sends the stored requests to the inference run manager one by one.
[0055] During the model inference process executed by the Inference Run Manager, when the main inference node (i.e., the main inference device) is running, the Inference Run Manager performs model inference. The pre- and post-processing modules complete data preprocessing and result post-processing respectively before and after the Inference Run Manager calls the model. The Inference Run Manager divides tasks into long request tasks or short request tasks according to whether the model inference time exceeds a specified threshold. Short request tasks are directly added to the complete model inference task as a job in the inference job queue, while long request tasks are mapped to a series of sequentially executed model block inference subtasks. Each model block inference subtask is added to the inference job queue. After addition, the model / model block is loaded from the model and block repository in the queue order and executed. Finally, the inference result or cached subtask result is returned. When the backup node (i.e., the backup inference device) is running, the Inference Run Manager only synchronously preloads the same model or model block as the main inference node, but does not perform inference.
[0056] The state saving module and the state recovery module are not only responsible for saving the CPU state of the inference-related processes as checkpoints and fully restoring them when necessary, but also for copying and restoring the inference log of the inference run manager. The inference log includes the inference job queue, the initial data sent by the user, or the intermediate results generated during the inference process. The checkpoint log synchronization module is responsible for synchronizing the above states to another node when a new task arrives or when the model or block inference is completed. Preferably, the synchronization process adopts Remote Direct Memory Access (RDMA) technology.
[0057] Furthermore, in S2, reference Figure 5 For CNN models, without splitting the fully connected layers, the splitting point of the CNN model is delayed, so that the execution time of other blocks except the last block n is close to the preset threshold.
[0058] For the Transformer model, it is divided into two parts: encoders and decoders. If the execution time of these two parts is greater than a preset threshold, the model is split proportionally between the encoder layer of the encoders and the decoder layer of the decoders until the execution time of the block is lower than the preset threshold.
[0059] For the LSTM model, it is divided into equal parts until the execution time of each block is lower than a preset threshold.
[0060] After the model is segmented, the start time of execution of each segment of the model is set as the checkpoint. Thus, the segments 1-n of the CNN model correspond one-to-one with checkpoint 1-checkpoint n, the encoder segments 1-m of the Transformer model correspond one-to-one with checkpoint 1-checkpoint m, and the decoders segments 1-n of the Transformer model correspond one-to-one with checkpoint m+1-checkpoint m+n.
[0061] Specifically, in order to improve the efficiency of state synchronization, when segmenting the model operator graph, it is necessary to select the segmentation points with a smaller number of edges and less data transmission. The reason for choosing different model segmentation mechanisms is: (1) Since the CNN model contains multiple convolutional and pooling layers, and these layers usually downsample the data, the data transmission between CNN model layers has an overall decreasing trend. Figure 5 (a) Delaying the segmentation point of the CNN model as much as possible while avoiding splitting the fully connected layer, so that the execution time of other blocks except the last block is as close as possible to the preset threshold, thereby reducing the amount of data transmission; (2) The Transformer model adopts an encoder+decoder structure, and the encoder and decoder are composed of multiple encoder layers and decoder layers respectively. The multiple encoder layers and decoder layers contain multi-head attention and fully connected structures, and the amount of data transmission between layers is relatively large, while the amount of data transmission between layers is relatively low. Figure 5 (b) First, the transformer model is divided into encoders and decoders. If the execution time of these two parts is greater than the preset threshold, the model is divided proportionally between the encoder and decoder layers until the execution time of each block is lower than the preset threshold. (3) In LSTM models (such as RNN, LSTM, etc.), the processing structure within the LSTM unit is relatively complex and the amount of data transmitted is relatively large. Refer to Figure 5 (c) The model is divided into equal parts among the LSTM units until the execution time of each block is lower than the preset threshold. After the division is completed, the system sets the start time of each model block as the checkpoint.
[0062] Furthermore, in S3, reference Figure 1 The service migration process includes two phases: normal operation and service migration. The specific steps are as follows:
[0063] S31. Before the normal operation phase, elect a master node and a backup node, and bind the cluster virtual IP to the master node;
[0064] S32. During normal operation, the master node executes inference requests, synchronizes checkpoints after short requests are completed, executes long requests in blocks, and synchronizes intermediate results and checkpoints after the sub-blocks are completed.
[0065] In step S32, the main inference interface component forwards the inference request sent by the user to the main inference run manager. It executes model chunks for long requests (implemented through a model segmentation mechanism) and executes complete models for short requests. The task is parsed into a job and placed in the inference job queue. If a long request is executed, before the main node executes a chunked inference job, the main state management component packages the inference job queue and intermediate inference results (or initial data) into an inference log and synchronizes it to the backup node along with the CPU checkpoint. Then, the main inference manager loads the required model or chunks and executes the inference. Meanwhile, the backup inference run manager synchronously preloads the required model or chunks but does not execute them. This process repeats until the model inference is complete. If a short request is executed, the main state management component only synchronizes the inference log before inference begins. Finally, when the model inference for both long and short requests is completed, the main inference interface component synchronizes the inference job queue again and returns the inference result to the user.
[0066] S33. During the service migration phase, the standby node restores the service status according to the most recent checkpoint and migrates the cluster service IP to the standby node.
[0067] In step S33, the backup migration control component notifies other components of the backup inference device to rebuild the inference state. Upon receiving the notification, the backup inference interface component binds the cluster virtual IP to the current node and temporarily stores subsequent requests in the local cache to prevent loss. The backup state management component restores the CPU state and inference job queue based on the most recently received checkpoint and re-infers from the intermediate results contained in the checkpoint to restore the GPU state. After the inference state reconstruction is completed, the backup inference interface component releases the cached inference requests in sequence. Then, the backup node enters the normal operation phase and continues to provide services as the new master node.
[0068] For details, please refer to Figure 6 Before operation begins, a primary and backup node are elected, and the cluster virtual IP is bound to the primary node. During normal operation, the primary node executes user inference requests and synchronizes the inference status and checkpoints to the backup node after completing the inference job. If the backup node does not receive a heartbeat message within a timeout period, it enters the service migration phase. During this phase, the backup node performs virtual IP migration and service state switching, and caches incoming inference requests. After the switchover is complete, the backup node becomes the new primary node and re-enters the normal operation phase.
[0069] Example 2: Reference Figure 2 and Figure 3This embodiment describes a GPU inference fast migration system based on model segmentation, used to execute the GPU inference fast migration method based on model segmentation described in Embodiment 1, including a main inference device 1, a backup inference device 2, and a network switch 3.
[0070] The main inference device 1 is connected to the backup inference device 2, the main inference device 1 is connected to the network switch 3, and the backup inference device 2 is connected to the network switch 3.
[0071] Both the main inference device 1 and the backup inference device 2 include an intelligent inference acceleration processor (GPU), memory, and CPU.
[0072] For details, please refer to Figure 3 In this system, IP Addr A represents the IP address configuration information of the network interface named A, IP Addr B represents the IP address configuration information of the network interface named B, and IP Addr C represents the IP address configuration information of the network interface named C. The GPU inference fast migration system consists of two peer inference devices. Each inference device includes an intelligent inference accelerator processor (i.e., GPU / NPU, etc.), memory, and a CPU with the same architecture and instruction set, which is used to execute inference tasks. The primary inference device 1 and the backup inference device 2 are connected by an internal network cable, which can transmit key information such as liveness detection messages, checkpoint status, and inference logs (model block inference), thereby achieving state synchronization and work coordination between devices. During operation, the two inference devices dynamically determine which one will become the primary inference device 1, obtain the cluster virtual IP, and provide services to the outside world through node election, while the other one, as the backup inference device 2, obtains the cluster virtual IP and takes over the services when the primary inference device 1 fails. External users always access the inference service through the network switch 3 via the cluster virtual IP, making the migration process seamless for users.
[0073] Although the invention has been described with reference to a limited number of embodiments, those skilled in the art will understand from the foregoing description that other embodiments are conceivable within the scope of the invention described herein. Furthermore, it should be noted that the language used in this specification has been chosen primarily for readability and edibility purposes, and not for the purpose of interpreting or limiting the subject matter of the invention. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. The scope of the invention is as follows.
Claims
1. A GPU inference fast migration method based on model segmentation, characterized in that, Includes the following steps: S1. Design software components for inference devices in a GPU inference fast migration system; S2. Design corresponding model segmentation mechanisms for typical inference models; S3. Employ software components and a model segmentation mechanism to design a service migration process and complete rapid migration of GPU inference; In S2, for the CNN model, without splitting the fully connected layer, the splitting point of the CNN model is delayed, so that the execution time of the other blocks except the last block n is close to the preset threshold. For the Transformer model, it is divided into two parts: encoders and decoders. If the execution time of these two parts is greater than a preset threshold, the model is split proportionally between the encoder layer of the encoders and the decoder layer of the decoders until the execution time of the block is lower than the preset threshold. For the LSTM model, it is divided into equal parts until the execution time of each block is lower than a preset threshold. After the model is split, the start time of execution of each model block is set as the checkpoint; In S3, the service migration process includes two phases: normal operation and service migration. The specific steps are as follows: S31. Before the normal operation phase, elect a master node and a backup node, and bind the cluster virtual IP to the master node; S32. During normal operation, the main inference interface component forwards the inference requests sent by the user to the main inference run manager. It executes model chunks for long requests and complete models for short requests. It parses the tasks into jobs and puts them into the inference job queue. If a long request is executed, before the master node executes a chunked inference job, the main state management component packages the inference job queue and intermediate inference results into an inference log and synchronizes it with the checkpoint to the standby node. Then the main inference manager loads the required model or chunk and executes the inference. The standby inference run manager synchronously preloads the required model or chunk but does not execute it. The above strategy process is repeated until the model inference is completed. If a short request is executed, the main state management component only synchronizes the inference log before the inference begins. Finally, when the model inference for both long and short requests is completed, the main inference interface component synchronizes the inference job queue again and returns the inference result to the user. S33. During the service migration phase, the standby migration control component notifies other components of the standby inference device to rebuild the inference state. Upon receiving the notification, the standby inference interface component binds the cluster virtual IP to the current node and temporarily stores subsequent requests in the local cache. The standby state management component restores the CPU state and inference job queue based on the most recently received checkpoint and re-infers from the intermediate results contained in the checkpoint to restore the GPU state. After the inference state reconstruction is completed, the standby inference interface component releases the cached inference requests in sequence. Then the standby node enters the normal operation phase and continues to provide services as the new master node.
2. The GPU inference fast migration method based on model segmentation according to claim 1, characterized in that, In S1, the software components of the inference device in the GPU inference fast migration system include a migration control component, an inference interface component, an inference operation manager, and a state management component. That is, the main inference device is equipped with a main migration control component, a main inference interface component, a main inference operation manager, and a main state management component, and the backup inference device is equipped with a backup migration control component, a backup inference interface component, a backup inference operation manager, and a backup state management component. The migration control component is used to continuously monitor the node's operating status and, when the node currently providing inference services fails, call other components to perform service and route switching. The inference interface component is used to receive external inference requests and forward them to the inference run manager. After the inference request is processed, the inference result is returned. The inference run manager is used to execute the model inference process, and it includes an inference job queue, a model and block repository, and pre- and post-processing modules. The state management component is used to realize service state saving and inference state synchronization between the primary inference device and the backup inference device. The backup state management component includes a state recovery module and a state synchronization module, and the primary state management component includes a state saving module and a state synchronization module.
3. A model-based segmentation GPU inference fast migration system, characterized in that, The method for executing a fast migration method for GPU inference based on model segmentation as described in any one of claims 1-2 includes a main inference device (1), a backup inference device (2), and a network switch (3). The main inference device (1) is connected to the backup inference device (2), the main inference device (1) is connected to the network switch (3), and the backup inference device (2) is connected to the network switch (3); Both the main inference device (1) and the backup inference device (2) include an intelligent inference acceleration processor, memory, and CPU.
Citation Information
Patent Citations
Workload configuration optimization method and system under multi-instance GPU
CN119902885A
Application migration method and device, equipment, readable storage medium and program product
CN120104289A