Distributed checkpoint access method for large model
Through the distributed checkpoint shard storage and index call method, the problem of low checkpoint storage and loading efficiency in large model training is solved, and an efficient training process and improved computing resource utilization are achieved.
Patent Information
- Application Number
- CN202510279986.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
During the training of large-scale model, checkpoints have low storage and loading efficiency, resulting in increased training waiting time and reduced computing resource utilization.
The distributed checkpoint access method is adopted, and the checkpoint shard storage and index calls are used to avoid whole storage and retrieval. The log manager LogMgr and the history manager Log hisMgr are used to reduce the access pressure and ensure the continuity and efficiency of training.
It significantly improves the efficiency of large-scale model training, reduces storage pressure and training waiting time, and improves the utilization rate of computing resources and the overall performance and stability of model training.
Smart Images

Figure CN120216098A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and relates to a distributed checkpoint access method for large models, and more specifically to a method for distributed storage, extraction and disaster recovery of checkpoints based on the characteristics of large model sharding and scheduling distribution. Background Art
[0002] With the rapid development of generative artificial intelligence, large language models have become a new paradigm for data-driven applications and are widely used in fields such as text and image generation. However, large model training not only requires a large amount of GPU hardware resources, but also the training duration is normally about 3 to 6 months. In order to tolerate faults, it is necessary to store checkpoint information regularly. During the training of large models, a checkpoint refers to an intermediate snapshot point that periodically saves the current training state completely. This snapshot contains key data such as model parameters, optimizer state, and gradient information, which is used to quickly resume from this point after a training interruption or failure, thereby avoiding repeated calculations, shortening the training interruption time, and improving the utilization rate of computing resources and the overall efficiency of model training. Different from the training data mainly consisting of small files, the checkpoint of a single node of a large model can usually reach dozens or hundreds of GB. When multiple training nodes write simultaneously and need to read simultaneously during recovery, it poses a high throughput requirement for storage. At the same time, a key issue is that the entire training is interrupted during checkpointing. Therefore, by improving throughput and controlling the checkpoint time-consuming within a minimal proportion, it is also very important for reducing training waiting and improving the utilization rate of computing resources.
[0003] The size of a checkpoint is usually a multiple of the number of model parameters. For example, if the model parameters are stored in 32-bit floating-point numbers (FP32), then each parameter requires 4 bytes of storage space. In addition, information such as optimizer state and gradients also needs to be stored, which may increase the storage requirement to several times the size of the model parameters. How to optimize this process to improve training efficiency and reduce the storage pressure during large model learning is the current research focus. The current optimization methods mainly include the following several:
[0004] 1. Parameter quantization: By using a lower-bitwidth data type (such as FP16 or INT8), the storage requirement of model parameters can be significantly reduced, thereby reducing storage space and accelerating the loading speed. However, quantization may introduce precision loss and affect the performance of the model, especially in scenarios that require high-precision calculations.
[0005] 2. Sparsification: In many deep learning models, a large number of parameters are actually close to zero. By sparsifying, only non-zero parameters are stored and updated, which can reduce the use of storage and computing resources. However, sparsification may increase the computational complexity because special handling of zero and non-zero values is required, and the storage and access of sparse data may not be as efficient as that of dense data.
[0006] 3. Gradient Accumulation: By updating the model parameters only after multiple training steps are executed, the write operations required for each update can be reduced, thereby reducing I / O operations and improving training efficiency. However, gradient accumulation may affect the convergence speed and stability of the model and needs to be carefully adjusted to ensure the training effect.
[0007] 4. Delayed Update: Similar to gradient accumulation, delayed update can reduce the write frequency of checkpoints, reduce I / O operations, and improve training efficiency. However, delayed update may lead to inconsistencies in the model state, and it is necessary to ensure that the states can be correctly synchronized during recovery.
[0008] In summary, the existing model training acceleration methods integrate various technical means such as software, hardware, network, and parameter exchange. These methods are applicable to general deep learning models but do not consider optimizing the model learning efficiency by leveraging large-scale self-characteristics. Summary of the Invention
[0009] The purpose of the present invention is to solve the deficiencies in the prior art and propose a distributed checkpoint access method for large models, which can achieve efficient storage and loading of checkpoints during the training process of large-scale machine learning models. By means of sharded storage and indexed invocation of checkpoints, the whole storage and retrieval of checkpoints can be avoided, and the training efficiency of large models can be significantly improved.
[0010] To achieve the above object, the present invention is implemented through the following technical solutions: A distributed checkpoint access method for large models, the method comprising the following steps:
[0011] Step S1, constructing a distributed checkpoint access system, the system including a log manager LogMgr and a history record manager Log hisMgr. The log manager LogMgr records the information and status of each checkpoint in chronological order. The history record manager Log hisMgr is deployed on the GPU server and receives instructions from the scheduling dispatcher and the log manager LogMgr to complete the maintenance of checkpoints for long-duration training tasks; the log manager LogMgr includes a LogList management linked list.
[0012] LogList management linked list: Utilize multiple LogGraph graph structures to record and manage information of a series of checkpoints;
[0013] LogGraph graph structure: Used to describe the topological structure and attributes of the LogChunk shard data structure of a single checkpoint;
[0014] LogChunk shard data structure: Used to represent a specific checkpoint shard and its storage and association information;
[0015] Step S2, distributed checkpoint storage
[0016] Step S2.1, according to large model training, configure the large model path, large model parallel mode, training hyperparameters, and storage period for the checkpoint, and record the relevant configuration information as checkpoint_setting; After receiving checkpoint_setting, the historical record manager Log hisMgr on the GPU server parses it, confirms whether to perform the checkpoint storage task, and waits for the large model training to start after the configuration information is globally synchronized;
[0017] Step S2.2, when the large model training starts, the scheduling dispatcher shards the network of the large model according to the large model parallel mode to obtain multiple shards ModelChunk, and schedules the shards ModelChunk to different GPU servers for execution respectively. At the same time, the scheduling dispatcher passes the shards ModelChunk and scheduling information to the log manager LogMgr. The log manager LogMgr extracts the information and dependency relationships of the LogChunk shard data structure according to the shards ModelChunk and scheduling information according to the LogGraph graph structure. Then, the historical record managers Log hisMgr on each GPU server perform status synchronization, and the log manager LogMgr maintains the LogList management linked list;
[0018] Step S2.3: When the number of iterations of the large model meets the storage cycle configuration requirements, the historical record manager Log hisMg on each GPU server triggers the checkpoint sharding storage task. The historical record manager Log hisMgr calls the local parameter storage service of the GPU server to store the parameters and status corresponding to the sharded ModelChunk, and returns the information and status of the checkpoint shards. The historical record manager Log hisMgr passes them to the log manager LogMgr. The log manager LogMgr parses and stores the information and status of the passed checkpoint shards according to the LogGraph graph structure and the LogChunk sharding data structure at the corresponding storage location, and simultaneously completes the information synchronization of the LogList management linked list. The log manager LogMgr sends an acknowledgment completion message to the historical record manager Log hisMgr. After the checkpoint backup is also stored, the log manager LogMgr notifies the scheduling dispatcher to re-trigger Step S2.1 to start a new cycle of checkpoint storage;
[0019] Step S3: Distributed checkpoint loading
[0020] When the log manager LogMgr receives the checkpoint loading instruction, it retrieves the corresponding checkpoint in the LogList management linked list, parses the LogGraph graph structure, and obtains the physical storage information of all shards of the checkpoint. Then it sends a shard loading instruction to the historical record manager Log hisMgr where the physical storage information of all checkpoint shards is located, and obtains the data set of all shards and status of the checkpoint. The data set is compressed, transmitted, decompressed, and assembled to obtain the final result.
[0021] Preferably, the LogList management linked list is used to manage and record a list structure of a series of checkpoints. Each element represents the checkpoint information corresponding to a specific training iteration cycle and contains the detailed description structure LogGraph graph structure of the checkpoint;
[0022] Each group of elements in the LogGraph graph structure contains the LogChunk sharding data structure and its index and timestamp;
[0023] The LogChunk sharding data structure contains the training iteration cycle, shard naming, physical storage location, and link information between this shard and adjacent shards.
[0024] Preferably, the specific steps of Step S2.1 are as follows:
[0025] According to the large model training, configure the large model path, large model parallel mode, training hyperparameters, and storage period for the checkpoint, and record the relevant configuration information as checkpoint_setting; the scheduling dispatcher distributes checkpoint_setting to the GPU servers corresponding to the checkpoint storage in the previous round and this round. After receiving checkpoint_setting, the historical record manager Log hisMgr deployed on the GPU server parses it. If it no longer participates in checkpoint storage, it stops triggering the storage task; if it participates in the new round of checkpoint storage, it obtains the configuration information. The historical record manager Log hisMgr generates a checkpoint shard storage task according to the configuration information, executes the local LogChunk shard data structure storage operation through the task trigger, and waits for the large model training to start after configuring global synchronization.
[0026] Preferably, in step S2.2, the log manager LogMgr extracts the information and dependencies of the LogChunk shard data structure according to the sharded ModelChunk and scheduling information in the LogGraph graph structure. The specific extraction principle is: there is a one-to-one correspondence between the sharded ModelChunk and the LogChunk shard data structure, the dependency relationship of the sharded ModelChunk is the dependency relationship of the LogChunk shard data structure, and the physical location of the GPU server where the execution physical location of the sharded ModelChunk is located is the physical storage location of the LogChunk shard data structure.
[0027] Preferably, in step S2.3, the specific steps are as follows:
[0028] When the number of iterations of the large model meets the storage cycle configuration requirements, the historical record manager Log hisMg on each GPU server triggers the checkpoint shard storage task. The historical record manager Log hisMgr calls the local parameter storage service of the GPU server to store the parameters and status corresponding to the shard ModelChunk responsible for calculation on this GPU server, and returns the specific information and status of the checkpoint shard. The historical record manager Log hisMgr that completes the storage task passes the specific information and status of the checkpoint shard to the log manager LogMgr. The log manager LogMgr extracts the specific information and status of the checkpoint shard according to the LogGraph graph structure and the LogChunk shard data structure for parsing and stores it in the corresponding storage location. At the same time, it completes the information synchronization of the LogList management linked list. After the synchronization is completed, the log manager LogMgr sends an acknowledgment completion message to the historical record manager Log hisMgr. When the checkpoint backup is also stored, the log manager LogMgr notifies the scheduling dispatcher to re-trigger step S2.1 to start a new cycle of checkpoint storage.
[0029] Preferably, during the parallel training of the large model, there are multiple large model shards in the GPU cluster, and each large model shard runs on different GPU servers. Considering system fault tolerance, when constructing a checkpoint distributed access system, a checkpoint dual-backup balancing strategy is adopted: select two different GPU servers to store the checkpoint shards respectively or according to the usage of storage resources, select a GPU server to store the checkpoint shard itself and its backup.
[0030] Preferably, the distributed checkpoint loading described in step S3 is specifically as follows:
[0031] Step S3.1, Checkpoint loading trigger and execution
[0032] After the LogMgr (Log Manager) receives the checkpoint loading instruction, it retrieves the corresponding checkpoint in the LogList management linked list, parses the LogGraph graph structure, obtains the physical storage information of all shards of the checkpoint, and then sends a shard loading instruction to the Log hisMgr (Historical Record Manager) where the physical storage information of each LogChunk shard data structure is located. After each Log hisMgr receives the shard loading instruction, it parses the training iteration cycle information, looks up the specific information and status of the corresponding checkpoint shard, and executes the shard loading instruction according to the records in the specific LogChunk shard data structure in this information to obtain the data set of the corresponding shard and status, denoted as ChunkSetitersPeroidSeq;
[0033] Step S3.2, checkpoint compression transmission and decompression
[0034] The Log hisMgr (Historical Record Manager) compresses the data set ChunkSetitersPeroidSeq, denoted as ZipChunkSetitersPeroidSeq. After completing the hybrid compression task, the Log hisMgr (Historical Record Manager) performs similarity detection on the compression parameters on different GPU servers, quickly filters out similar data blocks, and integrates the similar data blocks. After compression, the Log hisMgr (Historical Record Manager) triggers the parameter transmission mechanism and sequentially transmits the parameters to the LogMgr (Log Manager) according to the configuration in checkpoint_setting. After the LogMgr (Log Manager) receives the signal indicating the completion of transmission, it decompresses the ZipChunkSetitersPeroidSeq data set;
[0035] Step S3.3, checkpoint shard assembly and result return
[0036] After the decompression of ZipChunkSetitersPeroidSeq is completed, it is stored as ChunkSetiterPeroidSeqDecomposed, assembled and stored according to the LogGraph graph structure and the LogChunk shard data structure, and the result is returned.
[0037] Preferably, in order to avoid the situation where GPU overload or crash occurs during large-scale GPU cluster training, resulting in the inability to perform the corresponding computing tasks on a certain node normally, a fault tolerance and disaster tolerance mechanism is set as follows:
[0038] Step S4.1, triggering of checkpoint shard fault tolerance and disaster tolerance
[0039] According to the fault situation, the scheduling dispatcher re - adjusts the sharding and scheduling of the large model, updates and generates a new large - model sharding strategy, and waits to migrate the historical checkpoint of the faulty node to the new node;
[0040] Step S4.2, Disaster - tolerant migration of checkpoint shards
[0041] The historical record manager Log hisMgr on the GPU server where the faulty node is located calls the local parameter storage service of the GPU server to compress the stored large - model sharding parameters, and transfers the compressed model sharding parameters to the new GPU server according to the LogGraph graph structure in the model sharding and scheduling information of the scheduling dispatcher. The Log hisMgr in the new GPU server decompresses and updates the storage information of the local LogChunk sharding data structure;
[0042] Step S4.3, Checkpoint status synchronization and update
[0043] After the disaster - tolerant migration of the checkpoint shards is completed, the historical record manager Log hisMgr on each GPU server sends a successful disaster - tolerant migration message to the log manager LogMgr to synchronize the migration information. After receiving the migration information, LogMgr parses its content and updates the storage information of the migrated shards in the corresponding historical checkpoint in the LogList.
[0044] Preferably, in step S4.1, the faults are divided into GPU faults and GPU server faults; if it is a GPU fault, the historical checkpoint of the GPU server where the GPU fault occurs is migrated to the new GPU server; if it is a GPU server fault, the historical checkpoint of the backup server is migrated to the new GPU server.
[0045] The present invention has the following beneficial effects: By means of checkpoint sharding storage and index call, the present invention avoids the whole - storage and whole - retrieval of checkpoints, which can significantly improve the training efficiency of large models; by introducing modules such as LogMgr and Log hisMgr, the access pressure of checkpoints is reduced, ensuring the continuity and efficiency of distributed training. The distributed sharding storage mechanism of this method not only improves the storage efficiency, but also facilitates the rapid integration of checkpoint data as needed, thus realizing the flexibility and controllability of the training process in practical applications, and greatly improving the performance and stability of model training.
[0046] The present invention proposes a simple and feasible block storage strategy, which distributes checkpoints in a distributed manner in the GPU cluster. Each LogChunk is associated with its corresponding computing task and the shards of the large model. During the model training process, if it is necessary to restore or load a checkpoint, the specific block can be quickly located according to the position index in the LogGraph, thereby improving the loading efficiency and restoration speed.
[0047] The present invention creatively proposes a method for managing checkpoint blocks with dual indexing based on a graph structure. The LogGraph graph structure is used to record the LogChunk index and physical dependency relationships to record and track the distribution positions and statuses of each LogChunk, providing a structured way to store the index of the checkpoint. The dual-indexing method can not only meet the needs of fault tolerance and disaster tolerance but also minimize the storage requirements to the greatest extent.
[0048] The present invention constructs a dual-level fault tolerance and disaster tolerance mechanism for GPU failure levels and GPU server failure levels during large-scale learning processes. When a hardware failure occurs, dual backups of LogChunk can always be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic diagram of the sharding and scheduling process of the large model training model;
[0050] Figure 2 It is a schematic diagram of the distributed checkpoint system;
[0051] Figure 3 It is a schematic diagram of the distributed checkpoint storage process;
[0052] Figure 4 It is a schematic diagram of the distributed checkpoint loading process;
[0053] Figure 5 It is a schematic diagram of the distributed checkpoint fault tolerance and disaster tolerance mechanism process. DETAILED DESCRIPTION OF THE INVENTION
[0054] The following further describes the present invention in conjunction with embodiments, but it is not used as a basis for limiting the present invention.
[0055] A distributed checkpoint access method for a large model includes the following steps:
[0056] Step 1, construction of the distributed checkpoint system
[0057] The system includes a Log Manager (LogMgr) and a Log History Manager (Log hisMgr). LogMgr records the information and status of each checkpoint in chronological order through a management linked list LogList. LogList uses multiple LogGraph graph structures to record and manage the information of a series of checkpoints. LogGraph is used to describe the topological structure and attributes of each LogChunk shard data structure in a single checkpoint. Each checkpoint consists of multiple LogChunk shard data structures, and different LogChunk shard data structures are stored in different GPU servers respectively. At the same time, the LogChunk shard data structure corresponds to the ModelChunk shard of the large model, and there are also physical structure dependencies among the LogChunk shard data structures. Among them, the formal definitions of LogList, LogGraph, and LogChunk are as follows:
[0058] LogList = { <itersperoidseq> <loggraph> <...>],...,[ <itersperoidseq> <loggraph><...>]};
[0059] LogGraph = { <index> <timestamp> <logchunk> ],...,[ <index> <timestamp> <logchunk>};
[0060] LogChunk= <itersperoidseq><chunk_name><store_pos1><store_pos2> <uplink> <downlink> <leflink> <rightlink>;
[0061] Among them, LogList is a list structure used to manage and record a series of checkpoints. Each group of elements represents the information of the checkpoint corresponding to a specific training iteration period (itersPeroidSeq), and contains data such as the detailed description structure LogGraph of the checkpoint. Among the elements <itersperoidseq>The identification of the training iteration cycle corresponding to this checkpoint (such as the number of times of storage), <loggraph>Graph structure description for this checkpoint, used to represent which shards the checkpoint consists of and their relationships; <...> may contain other extended information;
[0062] LogGraph is used to describe the topological structure and its properties of each shard in a single checkpoint. It is a collection composed of multiple record items, and each record item contains the index, timestamp of the shard, and the corresponding shard data structure (LogChunk). <index>Used to uniquely identify the position of the sharded data structure in this checkpoint, <timestamp>Indicates the time information corresponding to the storage or update of this shard data structure (which can be used for version management or consistency checking);
[0063] LogChunk is used to represent a specific checkpoint shard and its storage and association information. Among them, itersPeroidSeq represents the training iteration period identifier corresponding to this shard, which is consistent with the period in LogList, facilitating quick retrieval and matching of which specific checkpoint version this shard belongs to. chunk_name is the name or identifier of this shard, used to distinguish the corresponding part of the shard in the model (such as corresponding to a certain network module or parameter block). <store_pos1> and <store_pos2> are the information of two different physical storage positions of the shard at the same position, used to implement double backup or redundant storage, thereby improving data fault tolerance and reliability. When a storage point fails, another backup can be used for recovery. <uplink> <downlink> <leftlink> <rightlink>Link information pointing to adjacent shards, indicating the relationship between this shard and its associated upper, lower, left, and right shards in the model topology. Through these links, the physical and logical dependencies of the checkpoint can be constructed, facilitating recovery and data consistency maintenance.
[0064] During the actual training process, LogMgr and the task scheduler TaskScheduler cooperate with each other to complete operations related to the access and storage of distributed checkpoints. According to the structure and computational load of the large model, it determines on which GPU servers in the GPU cluster each shard model should be executed, and at the same time cooperates with LogMgr, etc. to complete the access and management of checkpoints. The task scheduler ensures that when the training iteration reaches the set period, it triggers each GPU node to store and update the checkpoint. LoghisMgr is deployed on the GPU server, receives instructions from the Dispatcher (scheduling dispatcher) and LogMgr, is responsible for the management of checkpoint historical records and shard index management, and completes the maintenance of checkpoints for long-duration training tasks. The distributed checkpoint system avoids the whole-storage and whole-retrieval mode of traditional methods, but distributes the LogChunk shard data structure of the complete checkpoint among each GPU server, reducing the storage pressure on a single server and shortening the time-consuming for single checkpoint access, thereby improving the overall efficiency of model learning.
[0065] Step S2, Distributed checkpoint storage
[0066] Implement the distributed storage of each shard of the checkpoint in the GPU training cluster according to the LogGraph relationship.
[0067] Step S2.1, Checkpoint storage configuration, dynamic balancing, and global synchronization. The administrator configures configuration information such as the model path, parallel mode, training hyperparameters, and checkpoint storage period through the large model training cluster management interface. The checkpoint storage period is usually set to a fixed number of training iterations, such as 10,000 times. The checkpoint-related configuration information is denoted as checkpoint_setting, and the specific configuration can be obtained by LogMgr from the model training control center node. This is a common method and will not be elaborated here.
[0068] During the parallel training of large models, there are usually multiple model shards in the GPU cluster, and each model shard runs on different GPU servers. After the LogMgr obtains the checkpoint_setting, it triggers the dynamic balancing process of checkpoint storage for the large model learning scheduler Dispatcher. Dynamic balancing is a common method, and the difference lies in the balancing strategy. In the present invention, a checkpoint dual-backup balancing strategy is adopted: a. The two servers selected for one large model shard do not overlap; or b. Considering the usage of storage resources and making full use of the storage utilization rate of the cluster, select servers to store the model shard itself and its backup. After the large model learning scheduler Dispatcher selects the corresponding large model shard, it distributes the checkpoint_setting to the GPU servers corresponding to the checkpoint storage in the previous round and this round. The distribution process is a common message passing process.
[0069] After receiving the checkpoint_setting, the Log hisMgr deployed on the GPU server parses the configuration. If it no longer participates in checkpoint storage, it stops triggering the storage task. If it participates in the new round of checkpoint storage, it obtains configuration information such as the checkpoint storage period. The Log hisMgr generates a storage trigger rule according to the checkpoint configuration information and executes the local LogChunk shard data structure storage operation through the task trigger. After configuring global synchronization, it waits for the start of the large model training task execution.
[0070] Step S2.2, Start of large model training and synchronization of sharding scheduling information. The administrator starts the large model training through the large model training cluster management interface. The large model learning scheduling dispatcher Dispatcher shards the large model according to information such as the model parallelism method and the dual-backup balancing strategy, and schedules the task shards to different GPU servers for execution. The shards of the large model ModelChunk and the scheduling information (denoted as ModelParseInfo) include the specific network module information of different shards ModelChunk of the model, the physical dependency relationship between ModelChunks, and the GPU server and GPU card information for shard execution. Dispatcher passes ModelParseInfo to LogMgr. After receiving ModelParseInfo, LogMgr extracts the LogChunk shard information and the dependency relationship of LogChunk according to the LogGraph data structure respectively. The specific extraction process follows the following rules: 1. There is a one-to-one correspondence between ModelChunk and LogChunk. Similarly, the ModelChunk dependency relationship is also the LogChunk dependency relationship; 2. The GPU server where the physical location of ModelChunk execution is located is also the physical location where LogChunk is stored. The above message passing and information parsing are common methods in the industry and will not be elaborated. After this step is completed, LogMgr waits for the status synchronization of Log hisMgr on each GPU server and maintains LogList.
[0071] Step S2.3, Checkpoint sharding storage and status synchronization. When the number of model training iterations meets the requirements of the checkpoint storage period setting, the Log hisMgr of each GPU server triggers the checkpoint sharding storage task. Log hisMgr calls the local parameter storage service of the GPU server to store the parameters and status corresponding to the model shards responsible for calculation on this server. After the storage is completed, the local parameter service of the GPU server returns the specific information and status of the shard LogChunk, denoted as LocalCheckpointInfo. LocalCheckpointInfo is in the form of an array, specifically: LocalCheckpointInfo= <state> <loggraph>, where state is the status information of the shard, and the GPU servers at the other backup locations also use the LogGraph structure to manage the local shard LogChunk. The Log hisMgr that has completed the storage task passes the LocalCheckpointInfo to the LogMgr. After receiving the LocalCheckpointInfo, the LogMgr first extracts the itersPeroidSeq information. The LogMgr checks whether there is a corresponding management object in the LogList according to the itersPeroidSeq information. If there is no such information, a new management object is created, and the data structure is <itersperoidseq> <loggraph><...>, if the object already exists, there is no need to create a new one. After accurately matching the management object, the information in LocalCheckpointInfo is parsed according to the LogGraph and LogChunk structures and stored in the management object. The above operations complete the information synchronization work from the new LocalCheckpointInfo to the LogList. To improve the mechanism's fault tolerance, the CheckpointMeta file can be backed up to the persistent storage synchronously. After the synchronization is completed, LogMgr sends a confirmation completion message to Log hisMgr. When both checkpoints are synchronized, and <store_pos1> and <store_pos2> in LogChunk are updated, the storage of this checkpoint is completed. LogMgr notifies Dispatcher to re-trigger the 2.1 checkpoint dual-backup balancing strategy process and start a new cycle of checkpoint storage. The above transfer, comparison, and synchronization methods are common methods and will not be elaborated.
[0072] Step S3, distributed checkpoint loading. The administrator can extract historical checkpoints through the management interface to test the model performance at a specified historical time point during the training process, or continue training from a specified checkpoint to flexibly handle training exceptions, training interruptions, etc., as follows:
[0073] Step S3.1, checkpoint loading trigger and execution. The administrator loads the checkpoint corresponding to the specified itersPeroidSeq through the large model training cluster management interface. After receiving the checkpoint loading instruction, LogMgr retrieves the corresponding checkpoint in the LogList, parses the LogGraph, and obtains the physical storage information of all shards of this checkpoint. By default, <store_pos1> is used. If an exception occurs during the reading, it will automatically switch to <store_pos2>. Then, a shard loading instruction is sent to all Log hisMgrs where the physical storage information of LogChunk is located. After receiving the loading instruction, the LoghisMgrs on each GPU server parse the itersPeroidSeq information, find the corresponding LocalCheckpointInfo information, and perform the shard loading operation according to the specific LogChunk records in the LocalCheckpointInfo. The data set of the corresponding shard and its status obtained is denoted as ChunkSetitersPeroidSeq, and the structure is as follows: { <logchunk> <status><>iterPeroidSeq]}, the meaning of which is the same as that mentioned above.
[0074] Step S3.2, checkpoint compression transmission and decompression. Log hisMgr compresses the data set ChunkSetitersPeroidSeq through a method of hybrid compression of FP8 quantization and Huffman coding, denoted as ZipChunkSetitersPeroidSeq, and the data structure is: ZipChunkSetitersPeroidSeq = {[<compressed_Logchunk><compressed_status><compressed_iterPeroidSeq><scaling_factor> <huffmantree>....}, corresponding to a set of 8-bit floating-point numbers, such as [12.3, 0.1, -1.4,...]. After completing the hybrid compression task, Log hisMgr performs similarity detection on the compression parameters on different GPU servers, quickly filters out potentially similar data blocks, and integrates the similar data blocks to reduce the occupancy of the transmission bandwidth. The above hybrid compression, similarity detection, and data block integration are all commonly used methods in the industry and will not be elaborated here. After compression is completed, Log hisMgr triggers the parameter transmission mechanism and sequentially transmits the parameters to LogMgr according to the configuration in checkpoint_setting. After receiving the signal indicating the completion of transmission, LogMgr decompresses the ZipChunkSetitersPeroidSeq dataset;
[0075] Step S3.3, checkpoint shard assembly and result return. After decompression is completed, it is stored as ChunkSetiterPeroidSeqDecomposed. LogMgr extracts the iterPeroidSeq information and checks whether there is a corresponding management object in the LogList based on this information. If there is no such information, a new management object is created, and the data structure is <itersperoidseq> <loggraph><...>, where the meaning is the same as in Phase 1. If the object already exists, there is no need to create a new one. After accurately matching the management object, ChunkSetiterPeroidSeqDecomposed is assembled according to the LogGraph and LogChunk structures and stored in the management object, and compared with the checkpoint in LogList to avoid parameter defects, complete the synchronization process, return the result, and back up the CheckpointMeta file in the persistent storage for easy extraction next time.
[0076] The above process completes the checkpoint extraction and assembly work, and returns the final assembly result to the administrator operation interface, waiting for further operations by the administrator.
[0077] Step S4, distributed checkpoint fault tolerance and disaster recovery mechanism. During the large-scale GPU cluster training process, GPU overload or crash is likely to occur, resulting in the inability to perform the corresponding computing tasks on a certain node. For checkpoints, a fault tolerance mechanism is also required to handle abnormal situations. When a GPU fails, the transfer or reallocation of training parameters can be completed in a timely manner to ensure the normal operation of the checkpoint system. The specific steps are as follows:
[0078] Step S4.1, triggering of checkpoint shard fault tolerance and disaster recovery. When one or more GPUs in the GPU server fail, the training cluster management system can obtain the fault information. The large model learning scheduling dispatcher Dispatcher reschedules the training tasks according to the fault situation, and at the same time triggers the checkpoint shard fault tolerance and disaster recovery process. The specific steps are as follows: According to the specific fault situation, Dispatcher adjusts the model sharding and scheduling based on the available computing resources. Dispatcher updates and generates a new model sharding strategy. This step is the same as the processing in Step 2.2. In addition, in order to migrate the historical checkpoint of a faulty node to a new GPU server, the migration is carried out along the path of shard rescheduling. The format of a single migration message is { <logchunk> <srcpath> <destpath>, where ScrPath is the storage path on the source GPU server, and DestPath is the storage path on the new GPU server, denoted as ModelParseAdjusted. If a node in the GPU server fails, then SrcPath is the storage path on the failed GPU server; if the GPU server fails, at this time, the Log hisMgr of the failed server is unable to communicate, so SrcPath is the corresponding address of the non-failed GPU server in <store_pos1> and <store_pos2>, that is, in any case, the dual backup of the checkpoint can be guaranteed. Then the Dispatcher passes the ModelParseAdjusted information to the Log hisMgr on the source GPU server (if the GPU server fails, the source GPU server is the server where the backup is located) and the new GPU server. After receiving the ModelParseAdjusted information, the Log hisMgr on the new GPU server parses it and triggers the checkpoint shard disaster recovery migration task, and waits for synchronization after the migration is completed.
[0079] Step S4.2, checkpoint shard disaster recovery migration. The Log hisMgr on the source GPU server where the failed GPU is located calls the local parameter storage service of the GPU server, compresses the stored model shard parameters in the same way as step 3.2, and transfers the compressed model shards to the new GPU server through the network according to the data structure information of LogGraph in ModelParseInfo. The Log hisMgr in the new GPU server triggers decompression, performs decompression processing in the same way as step 3.2, and updates the local LogChunk storage information LocalCheckpointInfo.
[0080] Step S4.3, checkpoint status synchronization and update. After each GPU server included in ModelParseAdjusted completes the checkpoint shard migration, their respective Log hisMgrs send a disaster recovery migration success message to the LogMgr and synchronously transmit the ModelParseAdjusted information. After receiving ModelParseAdjusted, the LogMgr parses the content of ModelParseAdjusted, updates the storage information of the migrated shards in the corresponding historical checkpoint in the LogList, and backs up the CheckpointMeta file to the persistent storage.
[0081] The basic principles, main features and advantages of the present invention have been shown and described above. However, the above are only specific embodiments of the present invention, and the technical features of the present invention are not limited thereto. Any other embodiments obtained by those skilled in the art without departing from the technical solution of the present invention should be covered within the patent scope of the present invention.< / destpath> < / srcpath> < / logchunk> < / loggraph> < / itersperoidseq> < / huffmantree> < / status> < / logchunk> < / loggraph> < / itersperoidseq> < / loggraph> < / state> < / rightlink> < / leftlink> < / downlink> < / uplink> < / timestamp> < / index> < / loggraph> < / itersperoidseq> < / rightlink> < / leflink> < / downlink> < / uplink> < / itersperoidseq> < / logchunk> < / timestamp> < / index> < / logchunk> < / timestamp> < / index> < / loggraph> < / itersperoidseq> < / loggraph> < / itersperoidseq>
Claims
1. A distributed checkpoint access method for a large model, characterized in that: The method comprises the following steps: Step S1, constructing a distributed checkpoint access system, the system includes a log manager LogMgr and a history manager Log hisMgr, the log manager LogMgr records the information and status of each checkpoint in chronological order, the history manager Log hisMgr is deployed on the GPU server, receives instructions from the scheduler and the log manager LogMgr, and completes the maintenance of the checkpoint of the long-span training task; the log manager LogMgr includes a LogList management linked list; LogList management list: uses multiple LogGraph structures to record and manage a series of checkpoint information; LogGraph graph structure: used to describe the topological structure and properties of each LogChunk shard data structure of a single checkpoint; LogChunk shard data structure: used to represent a specific checkpoint shard and its storage and associated information; Step S2: Distributed checkpoint storage Step S2.1, according to the large model training, configure the large model path, large model parallel mode, training hyperparameters, and storage period for the checkpoint, and record the relevant configuration information as checkpoint_setting; after receiving the checkpoint_setting, the history record manager Log hisMgr on the GPU server parses it to confirm whether to perform the checkpoint storage task. After the configuration information is globally synchronized, wait for the large model training to start; Step S2.2, the large model training starts, the scheduling distributor slices the large model network according to the large model parallel mode, obtains multiple slices ModelChunk, and schedules the slices ModelChunk to different GPU servers for execution. At the same time, the scheduling distributor passes the slice ModelChunk and scheduling information to the log manager LogMgr. The log manager LogMgr extracts the information and dependencies of the LogChunk slice data structure according to the slice ModelChunk and scheduling information according to the LogGraph graph structure. Then, the history record manager Log hisMgr on each GPU server synchronizes the status, and the log manager LogMgr maintains the LogList management list. Step S2.3, when the number of iterations of the large model meets the storage cycle configuration requirements, the history record manager Log hisMg on each GPU server triggers the checkpoint shard storage task, and the history record manager Log hisMgr calls the GPU server local parameter storage service to store the corresponding parameters and status of the shard ModelChunk, and returns the information and status of the checkpoint shard, which is passed to the log manager LogMgr by the history record manager Log hisMgr. The log manager LogMgr parses the information and status of the passed checkpoint shard according to the LogGraph graph structure and the LogChunk shard data structure and stores them in the corresponding storage location, and completes the information synchronization of the LogList management linked list at the same time. The log manager LogMgr sends a confirmation completion message to the history record manager Log hisMgr. When the checkpoint backup is also stored, the log manager LogMgr notifies the scheduling distributor to re-trigger step S2.1 and start a new cycle of checkpoint storage; Step S3: Distributed checkpoint loading The log manager LogMgr receives the checkpoint loading instruction, retrieves the corresponding checkpoint in the LogList management list, and parses the LogGraph graph structure to obtain the physical storage information of all shards of the checkpoint. It then sends a shard loading instruction to the historical record manager Log hisMgr where the physical storage information of all checkpoint shards is located to obtain the data set of all shards and states of the checkpoint. The data set is compressed, transmitted, decompressed and assembled to obtain the final result.
2. The distributed checkpoint access method for large models according to claim 1, characterized in that: The LogList management list is used to manage and record a series of checkpoint list structures. Each element represents the checkpoint information corresponding to a specific training iteration cycle and contains a detailed description structure LogGraph of the checkpoint. Each group of elements in the LogGraph graph structure includes each LogChunk shard data structure and its index and timestamp; The LogChunk shard data structure includes the training iteration period, shard naming, physical storage location, and link information between the shard and adjacent shards.
3. The distributed checkpoint access method for large models according to claim 1, characterized in that: Step S2.1, the specific steps are as follows: According to the large model training, the checkpoint is configured with the large model path, large model parallel mode, training hyperparameters, and storage cycle, and the relevant configuration information is recorded as checkpoint_setting; the scheduler distributes checkpoint_setting to the corresponding GPU servers participating in the previous round and the current round of checkpoint storage. The history record manager Log hisMgr deployed on the GPU server receives the checkpoint_setting and parses it. If it no longer participates in the checkpoint storage, the storage task trigger is stopped; if it participates in the new round of checkpoint storage, the configuration information is obtained. The history record manager LoghisMgr generates a checkpoint shard storage task according to the configuration information, and executes the local LogChunk shard data structure storage operation through the task trigger. After configuring global synchronization, wait for the large model training to start.
4. The distributed checkpoint access method for large models according to claim 1, characterized in that: In step S2.2, the log manager LogMgr extracts the information and dependency of the LogChunk shard data structure according to the shard ModelChunk and the scheduling information according to the LogGraph graph structure. The specific extraction principle is: the shard ModelChunk and the LogChunk shard data structure are in a one-to-one correspondence, the dependency of the shard ModelChunk is the dependency of the LogChunk shard data structure, and the GPU server where the execution physical location of the shard ModelChunk is located is the physical location where the LogChunk shard data structure is stored.
5. The distributed checkpoint access method for large models according to claim 1, characterized in that: In step S2.3, the specific steps are as follows: When the number of iterations of the large model meets the storage cycle configuration requirements, the history record manager LoghisMg on each GPU server triggers the checkpoint shard storage task. The history record manager Log hisMgr calls the local parameter storage service of the GPU server to store the parameters and status corresponding to the shard ModelChunk calculated by this GPU server, and returns the specific information and status of the checkpoint shard. The history record manager Log hisMgr that completes the storage task passes the specific information and status of the checkpoint shard to the log manager LogMgr. The log manager LogMgr extracts the specific information and status of the checkpoint shard according to the LogGraph graph structure and the LogChunk shard data structure, parses them and stores them in the corresponding storage location, and completes the information synchronization of the LogList management linked list at the same time. After the synchronization is completed, the log manager LogMgr sends a confirmation completion message to the history record manager Log hisMgr. When the checkpoint backup is also stored, the log manager LogMgr notifies the scheduling distributor to re-trigger step S2.1 and start a new cycle of checkpoint storage.
6. The distributed checkpoint access method for large models according to claim 1, characterized in that: When training large models in parallel, there are multiple large model shards in the GPU cluster, and each large model shard is distributed and runs on different GPU servers. Considering the system fault tolerance, a checkpoint dual-backup balancing strategy is adopted when building a checkpoint distributed access system: select two different GPU servers to store the checkpoint shards respectively or select a GPU server to store the checkpoint shard itself and its backup based on the usage of storage resources.
7. The distributed checkpoint access method for large models according to claim 1, characterized in that: The distributed checkpoint loading in step S3 is specifically as follows: Step S3.1, checkpoint loading trigger and execution After receiving the checkpoint loading instruction, the log manager LogMgr retrieves the corresponding checkpoint in the LogList management list, parses the LogGraph structure, obtains the physical storage information of all shards of the checkpoint, and then sends the shard loading instruction to the history record manager Log hisMgr where the physical storage information of all LogChunk shard data structures is located. After receiving the shard loading instruction, each history record manager Log hisMgr parses the training iteration cycle information and searches for the specific information and status of the corresponding checkpoint shard. According to the specific LogChunk shard data structure record in the information, the shard loading instruction is executed to obtain the data set of the corresponding shard and status, which is recorded as ChunkSetitersPeroidSeq; Step S3.2: checkpoint compression, transmission and decompression The history record manager Log hisMgr compresses the data set ChunkSetitersPeroidSeq, recorded as ZipChunkSetitersPeroidSeq. After completing the mixed compression task, the history record manager Log hisMgr performs similarity detection on the compression parameters on different GPU servers, quickly screens similar data blocks, and integrates similar data blocks. After the compression is completed, the history record manager Log hisMgr triggers the parameter transmission mechanism and transmits the parameters to the log manager LogMgr in sequence according to the configuration in checkpoint_setting. After receiving the signal that the transmission is completed, the log manager LogMgr decompresses the ZipChunkSetitersPeroidSeq data set; Step S3.3, checkpoint shard assembly and result return After ZipChunkSetitersPeroidSeq is decompressed, it is stored as ChunkSetiterPeroidSeqDecomposed, assembled and stored according to the LogGraph graph structure and LogChunk shard data structure, and the result is returned.
8. The distributed checkpoint access method for large models according to claim 6, characterized in that: In order to avoid GPU overload or crash during large-scale GPU cluster training, which may cause the corresponding computing task of a certain node to fail to proceed normally, a fault-tolerance and disaster-tolerance mechanism is set up as follows: Step S4.1, checkpoint shard fault tolerance and disaster recovery trigger The scheduler and distributor readjusts the sharding and scheduling of the large model according to the fault situation, updates and generates a new large model sharding strategy, and waits for the historical checkpoints of the faulty node to be migrated to the new node; Step S4.2: Checkpoint shard disaster recovery migration The history record manager Log hisMgr on the GPU server where the failed node is located calls the local parameter storage service of the GPU server to compress the stored large model shard parameters, and passes the compressed model shard parameters to the new GPU server according to the model shard of the scheduler distributor and the LogGraph graph structure in the scheduling information. The LoghisMgr in the new GPU server decompresses and updates the local LogChunk shard data structure storage information. Step S4.3, checkpoint status synchronization and update After completing the checkpoint shard disaster recovery migration, the history record manager Log hisMgr on each GPU server sends a disaster recovery migration success message to the log manager LogMgr and synchronously transmits the migration information. After receiving the migration information, LogMgr parses its content and updates the migrated shard storage information in the corresponding historical checkpoint in LogList.
9. The distributed checkpoint access method for large models according to claim 8, characterized in that: In step S4.1, the failure is divided into GPU failure and GPU server failure. If it is a GPU failure, the historical checkpoint of the GPU server with GPU failure is migrated to the new GPU server. If it is a GPU server failure, the historical checkpoint of the backup server is migrated to the new GPU server.