Online switching method of parallel training strategy and deep learning model training system

CN122021969BActive Publication Date: 2026-09-25BEIJING WUWEN CORE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610096887.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-09-25
Estimated Expiration
2046-01-23

AI Technical Summary

Technical Problem

然而,随着训练规模的扩大,这种检查点方案的开销变得极高,可能需要数十分钟的时间来加载和恢复状态,极大影响了训练效率

Benefits of technology

[0016]由此,本公开通过摒弃传统离线切换方法中的磁盘读写,转而基于参数重切片和迁移及GPU间直接通信实现了高效的在线并行切换,并且支持复杂并行策略及扩缩容功能,对大模型训练框架具有很强的适配性。该技术使得大规模模型训练能够在毫秒级别内完成并行策略切换,相比于检查点方案所需的几十分钟,极大提升了训练效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122021969B_ABST
    Figure CN122021969B_ABST
Patent Text Reader

Abstract

The present disclosure provides an online switching method of parallel training strategies and a deep learning model training system. The method comprises: determining new communication group information and a model parameter slicing manner according to new parallel dimension division information and new worker node topology information, and further generating a mapping scheme for model parameter slicing migration according to old parallel dimension division information and old worker node topology information; and migrating the model parameters to new worker nodes based on the slicing manner and the mapping scheme. The new communication group information can be used to enable communication groups, and the enabled communication groups and the model parameters migrated to the new worker nodes are used for the new worker nodes to perform parallel model training under new parallel dimension division. Thus, the present disclosure realizes online fast switching of parallel training strategies by enabling corresponding communication groups at runtime, reconstructing parameter mapping, and performing minimum-cost reslicing and migration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence, and in particular to an online switching method for parallel training strategies and a deep learning model training system. Background Technology

[0002] In training large-scale distributed deep learning models, the training task typically relies on specific parallel strategies to allocate computing resources (e.g., GPUs) and achieves efficient model training through GPU clusters. Current mainstream training frameworks (e.g., Megatron) depend on static parallel strategies, meaning that once the training task starts, its parallel strategy and GPU resource configuration cannot be changed. If a failure occurs during training, or if it's necessary to adjust the parallel strategy, increase or decrease GPU resources, the relevant solutions usually rely on checkpointing mechanisms. This mechanism saves information such as the current model and state to storage when training is interrupted and loads the most recent checkpoint when the task restarts. However, as the training scale increases, the overhead of this checkpointing scheme becomes extremely high, potentially taking tens of minutes to load and restore the state, significantly impacting training efficiency.

[0003] Furthermore, in dynamic workloads such as reinforcement learning, the long-tail effect of inference components often leads to reduced GPU utilization. This inefficient utilization not only affects the training process but also restricts the full use of computing resources. Therefore, addressing the dynamic allocation needs of computing resources has become an important direction for improving training efficiency and resource utilization, especially in training platforms based on large-scale GPU clusters (such as cloud platforms). Summary of the Invention

[0004] To address this, this disclosure proposes a deep learning model training scheme, which can be implemented as an online switching method for parallel training strategies and a deep learning model training system. By enabling the required communication groups at runtime, reconstructing parameter mappings, and performing minimum-cost resharding and transfer operations, the scheme achieves rapid online switching of parallel training strategies. This scheme can dynamically adjust parallel strategies at extremely low cost, allowing for flexible selection of suitable parallel strategies at different stages of training. Furthermore, it reduces training task interruptions due to faults, thereby improving the continuity and efficiency of the training process.

[0005] According to a first aspect of this disclosure, an online switching method for parallel training strategies is proposed. The method is used for deep learning model training and includes: determining new communication group information based on new parallel dimension partitioning information and new worker node topology information; determining the sharding method for each model parameter of the deep learning model based on the new parallel dimension partitioning information and the new worker node topology information, and further generating a mapping scheme for model parameter sharding migration based on the old parallel dimension partitioning information and the old worker node topology information; and migrating the model parameters to the new worker node based on the sharding method and the mapping scheme; wherein the new communication group information is used to enable the communication group, and the enabled communication group and the model parameters migrated to the new worker node are used by the new worker node to perform parallel model training under the new parallel dimension partitioning.

[0006] Optionally, the new parallel dimension partitioning information includes at least one of the following: tensor parallel (TP) partitioning information, pipeline parallel (PP) partitioning information, and expert parallel (EP) allocation information; determining the new communication group information based on the new parallel dimension partitioning information and the new working node topology information includes: determining a new global communication group based on the new working node topology information, and determining at least one of the following based on the new parallel dimension partitioning information and the new working node topology information: TP sub-communication group, PP sub-communication group, and EP sub-communication group.

[0007] Optionally, determining the sharding method for each model parameter of the deep learning model based on the new parallel dimension partitioning information and the new worker node topology information, and further generating a mapping scheme for model parameter sharding migration based on the old parallel dimension partitioning information and the old worker node topology information, includes: extracting parallel metadata of each virtual parameter in the global virtual parameter space based on the new parallel dimension partitioning information, wherein the global virtual parameter space is constructed as a unified logical view of the model parameters of the deep learning model; binding the physical shards of the model parameters held by the old worker nodes performing parallel model training under the old parallel dimension partitioning with the corresponding virtual parameter entries in the global virtual parameter space to obtain binding information; and traversing the global virtual parameter space based on the parallel metadata and the binding information to determine the sharding method and mapping scheme.

[0008] Optionally, extracting parallel metadata for each virtual parameter in the global virtual parameter space based on the new parallel dimension partitioning information includes at least one of the following: determining the TP parameter distribution state of the virtual parameter based on TP partitioning information as TP parallel metadata; determining the new working node where the virtual parameter is located based on PP partitioning information as PP parallel metadata; and determining the new working node where the virtual parameter is located based on EP allocation information as EP parallel metadata. Traversing the global virtual parameter space based on the parallel metadata to determine the partitioning method and mapping scheme for each model parameter includes at least one of the following: reconstructing all virtual parameters according to their respective TP parallel metadata based on TP partitioning information to obtain the virtual parameter distribution state; aggregating each virtual parameter's PP parallel metadata layer by layer based on PP partitioning information to obtain the virtual parameter combinations on each new working node participating in layer computation; and partitioning virtual parameters as non-shared expert parameters based on EP allocation information and EP parallel metadata to obtain the expert parameter combinations on each new working node participating in expert model computation.

[0009] Optionally, the method further includes: in response to enabling local allocation of optimizer parameters, determining an optimizer state parameter distribution scheme based on the data parallel DP splitting information and the sharding method and mapping scheme; and migrating the optimizer state parameters to the corresponding new working nodes based on the optimizer state parameter distribution scheme.

[0010] Optionally, based on the sharding method and mapping scheme, migrating model parameters to a new working node includes: in response to determining that the new working node includes a legacy node, during the migration of model parameters to the new working node, performing a transmission at the model parameter granularity, and the transmission at the model parameter granularity includes: loading the parallel metadata of the current model parameters under the new parallel dimension partition; performing point-to-point transmission to complete the migration of the current model parameters from the old working node to the new working node; and releasing the buffer of the current model parameters in the old working node, wherein the legacy node is a working node that participates in parallel model training both before and after the parallel training strategy switch.

[0011] Optionally, the method further includes: in response to determining that the number of new working nodes is greater than the number of old working nodes and that the new working nodes include newly created nodes, performing the following operations in the newly created nodes: during the parallel training process performed on the old working nodes, creating a new process and performing an initialization operation, and initializing a local communication group and establishing a communication connection according to the new communication group information, wherein the newly created nodes are working nodes that did not participate in parallel model training before the parallel training strategy switch and have no training process existing on them.

[0012] Optionally, the method further includes: based on a cached unified dataset, reconstructing a data loader and a data iterator according to the new parallel dimension partitioning information and the new worker node topology information, to provide input data in parallel model training under the new parallel dimension partitioning, wherein the unified dataset is created and cached for the data format and structure of nonparametric data involved in the training of the deep learning model.

[0013] According to a second aspect of this disclosure, a deep learning model training system is proposed for online switching of parallel training strategies, and includes: a communication management module for determining new communication group information based on new parallel dimension partitioning information and new worker node topology information; and a reslicing and migration module for determining the slicing method of each model parameter of the deep learning model based on the new parallel dimension partitioning information and the new worker node topology information, further generating a mapping scheme for model parameter slicing migration based on the old parallel dimension partitioning information and the old worker node topology information, and migrating the model parameters to the corresponding new worker node based on the slicing method and the mapping scheme, wherein the new communication group information is used to enable the communication group, and the enabled communication group and the model parameters migrated to the new worker node are used by the new worker node to perform parallel model training under the new parallel dimension partitioning.

[0014] Optionally, the system further includes: a state management module for caching a unified dataset and reconstructing a data loader and a data iterator based on the unified dataset according to the new parallel dimension partitioning information and the new worker node topology information, in order to provide input data in the parallel model training under the new parallel dimension partitioning, wherein the unified dataset is created and cached for the data format and structure of the nonparametric data involved in the training of the deep learning model.

[0015] Optionally, the system further includes: a state management module, used to release the GPU memory requested for training on the old working node and load the training state information required for training on the new working node before the reslicing and migration module migrates the model parameters, and to request the GPU memory required for training on the new working node after the reslicing and migration module migrates the model parameters.

[0016] Therefore, this disclosure achieves efficient online parallel switching by abandoning disk read / write operations in traditional offline switching methods, and instead relies on parameter reslicing, migration, and direct communication between GPUs. It also supports complex parallel strategies and scaling capabilities, demonstrating strong adaptability to large-scale model training frameworks. This technology enables large-scale model training to complete parallel strategy switching within milliseconds, significantly improving training efficiency compared to the tens of minutes required by checkpointing methods. Attached Figure Description

[0017] The above and other objects, features and advantages of this disclosure will become more apparent from the more detailed description of exemplary embodiments thereof taken in conjunction with the accompanying drawings, wherein like reference numerals generally denote like parts.

[0018] Figure 1 A schematic flowchart of an online switching method for a parallel training strategy according to an embodiment of the present disclosure is shown.

[0019] Figure 2 A flowchart illustrating an inter-process scaling scenario according to an embodiment of the present disclosure is shown.

[0020] Figure 3 A schematic diagram of the structure of a deep learning model training system according to an embodiment of the present disclosure is shown.

[0021] Figure 4 A flowchart illustrating a state switching operation performed by a state management module according to an embodiment of the present disclosure is shown.

[0022] Figure 5 A schematic diagram of a deep learning model training system according to an embodiment of the present disclosure is shown. Detailed Implementation

[0023] Preferred embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0024] As used in the specification and appended claims of this application, the singular expressions “a,” “the,” “the,” and “the” are intended to also include expressions such as “one or more,” unless the context explicitly indicates otherwise. The term “comprising” and its variations, as used herein, indicate an open-ended inclusion, i.e., “including but not limited to.” Unless specifically stated otherwise, the term “or” means “and / or.” The term “according to” means “at least in part according to.” The terms “an example embodiment” and “an embodiment” mean “at least one example embodiment.” The term “another embodiment” means “at least one additional embodiment.” The terms “first,” “second,” etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0025] In training large-scale distributed deep learning models, checkpointing and local elastic switching mechanisms are typically relied upon to change parallel strategies.

[0026] Under the checkpoint mechanism, checkpoints are periodically saved. Upon detecting resource changes or failures, the current process is terminated, and the state is restored from the saved checkpoints before training is restarted. This mechanism relies on stable storage and the restart of the training process, resulting in high switching costs and failing to meet the real-time elastic adjustment requirements of dynamic GPU resource environments.

[0027] Local elastic switching mechanisms avoid training restarts without relying on checkpoints by pre-storing model parameters, optimizer states, and even some intermediate computation results in memory or backing them up between worker nodes (e.g., GPUs). Here, "elastic switching" refers to the ability to dynamically and seamlessly adjust parallel strategies during parallel training. However, this local elastic switching scheme only supports a small number of adjustments within a predetermined range and cannot achieve highly elastic switching of parallel strategies.

[0028] To address the challenge of parallel strategy switching in large-scale distributed deep learning model training, a highly resilient parallel strategy switching scheme with low cost is needed. To meet this requirement, this disclosure proposes a novel deep learning model training method, specifically, an online method for switching parallel training strategies. Since training frameworks like Megatron employ parallelism based on reconstructable communication groups and explicit parameter sharding rules, communication groups can be regenerated, parameter mappings reconstructed, and resharding performed at runtime with minimal cost, depending on the new GPU topology and task scale. The resharded model parameters can be transferred computationally (e.g., point-to-point transfer from the source worker node to the target worker node), and communication costs and memory usage remain manageable during the switching process.

[0029] Figure 1 A schematic flowchart illustrating an online switching method for parallel training strategies according to one embodiment of this disclosure is shown. This online switching method for parallel training strategies (hereinafter also referred to as the "parallel strategy switching method" or "online switching method") can be used for deep learning model training and is also part of a parallel training method for deep learning models. The deep learning model is typically a large-scale deep learning model, such as a large language model (LLM). In one embodiment, the online switching method can be incorporated into a mainstream training framework, such as Megatron, and provide it with the ability to elastically adjust the parallel strategy online for parallel training.

[0030] Here, the "parallel training strategy" (hereinafter also referred to as the "parallel strategy") is a high-level concept that determines how to allocate and partition computing resources, workload, and data in training rounds. Parallel dimension partitioning information and worker node topology information are the input information for formulating the parallel training strategy. The former mainly determines how to partition model parameters and data, while the latter determines how to configure computing resources and devices. Therefore, the "old parallel training strategy" is formulated based on the old parallel dimension partitioning information and the old worker node topology information, and is used to guide the parallel training of the model in previous rounds (one or more previous training rounds). Once the new parallel dimension partitioning information and the new worker node topology are known, a "new parallel training strategy" can be formulated accordingly for the new round of model training. The online switching method for parallel training strategies disclosed herein is a scheme for switching deep model training from an old parallel training strategy to a new parallel training strategy.

[0031] In step S110, a new communication group is determined based on the new parallel dimension partitioning information and the new worker node topology information. In this disclosure, a worker node refers to a parallel computing unit / node participating in parallel training, and the parallel computing unit can in particular be implemented as a GPU. The following description will primarily use a GPU as an example. It should be understood that in other implementations, other parallel computing units / nodes besides GPUs, such as TPUs, can also be used.

[0032] Here, the new parallel dimension partitioning information refers to the information on the new parallel dimension partitioning, corresponding to the parallel dimension partitioning information to be used after switching parallel training strategies; that is, the relevant information on the parallel dimension partitioning to be used by the new parallel training strategy. In parallel training, multiple parallel dimensions are typically used to improve computational efficiency. Parallel dimensions can include: TP (Tensor Parallelism), used to partition the model parameter tensors along a specified dimension and allocate them to different GPUs; PP (Pipeline Parallelism), used to divide the model into multiple stages and execute the computation of these stages on different GPUs, forming a pipelined training process; DP (Data Parallelism), which divides the training data into multiple batches and processes the same model copy in each batch in parallel on multiple GPUs; and EP (Expert Parallelism), commonly used in MoE (Mixture of Experts) models, allocating / distributing expert models to different GPUs for computation. These parallel dimensions can be combined to maximize the utilization of computational resources and accelerate the training process. The new parallel dimension partitioning information can refer to the set of parallel dimension partitioning information involved in the parallel strategy to be executed, and may include the splitting / allocation of a certain dimension, the granularity of the splitting, and the allocation range, etc. It should be understood that, depending on the specific implementation, each parallel strategy may include one or more of the above-mentioned parallel dimensions, or other parallel dimensions, such as those proposed in the future.

[0033] Here, the new worker node topology information refers to the topology information of the new worker node, including the number, composition, and connection method of the parallel computing units participating in parallel training in the new parallel training strategy. For example, the topology information may include 64 GPUs, evenly distributed across 8 devices, with GPUs within devices connected using a first communication method and GPUs between devices connected using a second communication method, etc. The new worker node corresponds to the worker node included in the topology information of the new worker node, that is, the parallel computing unit, such as a GPU.

[0034] Knowing the new parallel dimension partitioning information and the new GPU topology, the grouping scheme for the new communication groups can be determined. The new communication groups are those used when executing the new parallel training strategy, including a global communication group and multiple communication subgroups. The global communication group can include all GPUs under the new GPU topology, while each notification subgroup corresponds to a GPU participating in parallel training under each parallel dimension partition.

[0035] In one embodiment, the new parallel dimension partitioning information includes at least one of the following: TP partitioning information; PP partitioning information; and EP allocation information. Accordingly, enabling a new communication group based on the new parallel dimension partitioning information and the new worker node topology information may include: enabling a new global communication group based on the new worker node topology information; and enabling at least one of the following based on the new parallel dimension partitioning information and the new worker node topology information: TP sub-communication group, PP sub-communication group, and EP sub-communication group. The specific sub-communication groups enabled correspond to the parallel dimensions included in the new parallel dimension partitioning information.

[0036] In step S120, based on the new parallel dimension partitioning information and the new working node topology information, the sharding method of each model parameter of the deep learning model is determined, and further based on the old parallel dimension partitioning information and the old working node topology information, a mapping scheme for model parameter sharding migration is generated.

[0037] In contrast to the above, the old parallel dimension partitioning information corresponds to the parallel dimension partitioning information used before the parallel training strategy switch; that is, the information related to the parallel dimension partitioning used by the old parallel training strategy. The old worker node topology information is the topology information of the old worker nodes, including the number, composition, and connection method of the parallel computing units participating in parallel training in the old parallel training strategy. The old worker node corresponds to the worker nodes included in the topology information of that old worker node, i.e., the parallel computing units, such as GPUs. It is important to note that "old" and "new" here are concepts relative to before and after the parallel strategy switch. Therefore, a new worker node does not necessarily mean a completely new and different node, but rather refers to a node participating in training under the new parallel strategy. As will be described below, in practical application scenarios, the new parallel strategy can completely reuse all or some of the worker nodes from the old parallel strategy. Therefore, the sets of old and new worker nodes often overlap, or even one is a subset of the other.

[0038] Specifically, in step S120, the partitioning method of each model parameter in each dimension can be determined first according to the partitioning rules of the new parallel dimension. For example, for the TP dimension, the weight parameter is divided into multiple sub-tensor partitions according to the size of the model parameter and the selected partitioning dimension; for the PP dimension, the allocation of each parameter in different stages is determined according to the number of pipeline stages; for the EP dimension, the parameter is divided into different expert groups according to the expert ID of the expert parameter.

[0039] After determining how each model parameter should be sharded, a mapping scheme for migrating each model parameter shard from the old working nodes to the new working nodes can be further determined based on the old parallel dimension partitioning information and the old working node topology information. That is, based on the new parallel dimension partitioning information and the new working node topology information, the migration destinations (destination working nodes, also known as target working nodes) of these shards obtained based on the determined sharding method can be determined, and the migration sources (source working nodes) of these shards can be determined based on the old parallel dimension partitioning information and the old working node topology information, and a mapping scheme can be generated accordingly.

[0040] Subsequently, in step S130, the model parameters can be migrated to each new working node based on the sharding method and mapping scheme. Specifically, based on the aforementioned determined sharding method, each model parameter can be re-sharded in practice, that is, reconstructed from sharding under the old parallel strategy to sharding under the new parallel strategy. Then, based on the mapping scheme, the re-sharded model parameters can be migrated to migrate from the source GPU of the old parallel strategy to the target GPU under the new parallel strategy.

[0041] After completing the online switch between the old and new parallel training strategies according to steps S110-S130, a new round of model training can begin. The new communication group information can be used by each new worker node to enable the communication group. The enabled communication group and the model parameters migrated to the new worker node can then be used by the new worker node to perform parallel model training under the new parallel dimension partitioning, that is, enabling each new worker node to perform a new round of parallel training based on the new parallel training strategy. Specifically, each new worker node can establish a local communication group or select an existing local communication group based on the new communication group information determined in step S110, enable these communication groups, and complete the communication connection. Thus, after the model parameter migration is completed, the new worker node can use the migrated model parameters based on the enabled communication group to perform parallel model training under the new parallel dimension partitioning. That is, under the new parallel strategy, the GPU cluster corresponding to all new worker nodes begins to execute a new round of model training.

[0042] Therefore, the online switching method disclosed herein achieves rapid online switching of parallel training strategies by determining a new communication group, reconstructing the parameter sharding method, generating a mapping scheme, and performing actual model parameter migration accordingly. In one embodiment, this method can be implemented as an extension of a large-scale training framework (e.g., Megatron), adding online scaling and dynamic parallel reconfiguration capabilities. This allows training jobs to be scaled up and accommodated without restarting when GPU availability fluctuates, and adapts to changes in parallel strategies across various dimensions.

[0043] Since the switching of the parallel training strategy can only be performed after the previous iteration of training has ended, in one embodiment, steps S110-S130 can be executed between the previous iteration and the new iteration. In other embodiments, steps S110 and S120 can also be performed in parallel with the previous iteration. That is, only S130 is performed between the two training rounds. However, since steps S110 and S120 take very little time in practice (e.g., a few milliseconds), they can be performed together with step S130 between the two training rounds. Furthermore, since there is no dependency, steps S110 and S120 can also be executed in parallel, or step S120 can be executed before step S110. In the following description, for ease of explanation, steps S120 and S130 can be referred to as the "online resharding and migration pipeline".

[0044] In practice, the determination of the sharding method for each model parameter in step S120 and the generation of the mapping scheme for model parameter shard migration do not need to be performed on actual model parameters, but can be based on virtual model parameters that do not actually occupy a large amount of GPU memory. Therefore, determining the sharding method and mapping scheme through virtual parameters can achieve actual model parameter transfer / migration with less computational and GPU memory overhead. In one embodiment, the above operation can be implemented by constructing a global virtual parameter space. To this end, step S120 may include: extracting parallel metadata for each virtual parameter in the global virtual parameter space based on the new parallel dimension partitioning information; binding the physical shards of the model parameters held by each old worker node performing parallel model training under the old parallel dimension partitioning to the corresponding entries in the global virtual parameter space to obtain binding information; and traversing the global virtual parameter space based on the parallel metadata and binding information to determine the sharding method and the mapping scheme. The global virtual parameter space can be constructed as a parameter space composed of a set of virtual parameters. In some embodiments, this space is a unified logical view of all model parameters and can be replicated completely and consistently across all GPUs. This global virtual parameter space allows for the decoupling of parameter representations from their physical partitions and eliminates the isolation of local parameter views between worker nodes. In other embodiments, considering that other model parameters, such as bias parameters, are typically not fragmented and have explicit and simple transfer methods, the virtual parameters in the global virtual parameter space may correspond only to the weight parameters of the deep learning model.

[0045] The parallel metadata for each model parameter can be metadata describing the parallel dimensions of that parameter. Based on the new parallel dimension partitioning information, extracting the parallel metadata for each virtual parameter in the global virtual parameter space can include at least one of the following: determining the TP parameter distribution state of the virtual parameter based on TP partitioning information, as TP parallel metadata; determining the new working node where the virtual parameter resides based on PP partitioning information, as PP parallel metadata (corresponding to determining the PP rank for each virtual parameter as described below); determining the new working node where the virtual parameter resides based on EP allocation information, as EP parallel metadata (corresponding to determining the EP rank for virtual parameters that are non-shared model parameters as described below). Accordingly, traversing the global virtual parameter space based on the parallel metadata to determine the partitioning method and mapping scheme for each model parameter includes at least one of the following: reconstructing all virtual parameters according to their respective TP parallel metadata based on TP partitioning information to obtain the virtual parameter distribution state (i.e., how each virtual parameter should be partitioned under TP partitioning); aggregating each virtual parameter by layer according to the PP parallel metadata based on PP partitioning information to obtain the virtual parameter combinations on each new working node participating in the layer computation; and partitioning the virtual parameters as non-shared expert parameters according to EP allocation information based on EP parallel metadata to obtain the expert parameter combinations on each working node participating in the expert model computation.

[0046] Here, the virtual parameter distribution state refers to how each virtual parameter should be partitioned according to the TP parallel metadata under the new TP partitioning information, and how the partitioned virtual parameters (i.e., slices) will be distributed across the new worker nodes. Specifically, the TP partitioning information includes the partitioning method of each virtual parameter in the TP dimension (e.g., how many slices, what data each slice contains, etc.), and the virtual parameter distribution state describes the partitioning result of each virtual parameter based on this TP partitioning information, and specifies how these slices are distributed among the new worker nodes during training, for example, how they are distributed across different GPUs. The virtual parameter combination on each new worker node participating in layer computation refers to the virtual parameters distributed on each new worker node after the layer partitioning in the PP dimension. The expert parameter combination on each worker node participating in expert model computation refers to which virtual parameters should be distributed as expert parameters on each new worker node involved in expert model computation after the EP dimension partitioning.

[0047] Different parallel dimensions affect the sharding of model parameters across devices. Although these parallel dimensions interact in distributed training, their effects on parameter arrangement are orthogonal. Therefore, when extracting parallel metadata, the metadata for each parallel dimension can be extracted independently and then unified into a unified layout description in a global virtual parameter space. For each virtual parameter, parallel metadata describing its layout under different parallel strategies (e.g., tensor parallel sharding dimension, pipeline stage allocation, data parallel replica index) and the transmission requirements required for actual migration / transmission can be extracted. In one embodiment, for each model parameter, the parallel metadata can capture the following information: sharding granularity (e.g., tensor sharding dimension); position in the multidimensional parallel topology (e.g., pipeline stage, replica index); and logical range in the global virtual parameter space.

[0048] In some implementations, the old parallel strategy is also a "new" strategy compared to some even older parallel strategies. Therefore, parallel metadata of the model parameters under the old parallel strategy is also extracted in previous operations. In this case, by comparing the parallel metadata under the old and new strategies, the minimum set of data migrations required to reconcile the two configurations can be determined, and a mapping scheme can be generated accordingly.

[0049] In some embodiments, a metadata generation module can be used to generate parallel metadata. For the TP dimension, model parameters are segmented along a specified dimension. The metadata generation module can maintain TP segmentation information for each parameter, including whether it is segmented, the segmentation dimension, and the current segmentation granularity. Given any model information, the parameter distribution in different TP spaces can be obtained through the TP segmentation information. For the PP dimension, model parameters are segmented into multiple PPStages. The metadata generation module maintains layer information for each model parameter. Through this layer information, the PP rank (or sequence number) of any parameter in a given PP space can be obtained. The PP rank of each virtual parameter can serve as the PP parallel metadata for that parameter as described above. For the EP dimension, expert parameters are mapped to different EP ranks based on their expert IDs. The metadata generation module can maintain expert ID information for each expert parameter. Through this expert ID, the EP rank of any expert parameter in a given EP space can be obtained. The EP rank of each virtual parameter belonging to an expert parameter can serve as the EP parallel metadata for that parameter as described above.

[0050] It should be understood that in parallel training, the TP space, PP space, and EP space refer to the spaces where model parameters or computational tasks are mapped to different GPUs or devices in their respective parallel dimensions. Since different local TP partitioning information can exist, different TP spaces can exist within a single parallel strategy. Furthermore, rank is a unique identifier for each worker node (e.g., GPU) in distributed training. Each GPU is assigned a unique rank in a given parallel strategy of distributed training. For example, in a parallel round involving 64 GPUs, these 64 GPUs would be assigned ranks 0 through 63. The PP rank and EP rank are the identifiers of each GPU in the PP and EP parallel dimensions during parallel computation.

[0051] As mentioned earlier, the global virtual parameter space is visible to every GPU, which also means it is visible to every rank, and its initial state is consistent. Therefore, it can be considered as a parameter space where all parallel techniques are enabled. Given any valid parallel dimension partitioning information and worker node topology information, the global virtual parameter space can be mapped to the corresponding parameter distribution. Specifically, the steps to obtain the parameter distribution of the global virtual parameter space based on the above information are as follows: (1) First, consider the TP dimension. Given the TP Size (i.e., the number of tensors to be split in tensor parallelism), reshape all virtual parameters according to the TP attributes to obtain virtual parameters on any TP rank in the TP space. (Here, the TP dimension reconstruction is to re-slice the parameters according to the size of the TP dimension under the new parallel strategy, so as to adapt them to the new parallel training strategy without changing the essential content of the tensor.) (2) Next, consider the PP dimension. Given the PP Size (i.e., the number of pipeline stages in pipeline parallelism), aggregate according to the layer ID of each virtual parameter to obtain the combination of virtual parameters under different PP ranks. (3) Then consider the EP dimension. If the current virtual parameter belongs to the non-shared expert parameter, it is divided according to its expert ID to obtain the expert parameter combination under different EP ranks.

[0052] Given any two sets of valid information, two parameter distribution methods in the global virtual parameter space can be obtained. Therefore, for any parameter, the source and destination of each parameter fragment under the condition of parallel strategy switching can be determined, and thus the mapping scheme can be determined. Therefore, during the resharding process, by traversing the global virtual parameter space, the precise send / receive operations used to transform the parameter layout from the old parallel strategy to the new parallel strategy can be determined.

[0053] Common parallel dimensions in parallel training, besides TP, PP, and EP, include DP. DP divides the training data into multiple batches, with each batch computing the same model in parallel on different GPUs. Each GPU copies the complete model parameters, then each GPU computes the gradient of its allocated portion of the data. Finally, the gradients are aggregated across all GPUs via a communication method (e.g., All-Reduce), updating the model parameters. Regular DP does not involve sharding model parameters. For example, ZeRO-1 (Zero Redundancy Optimizer stage one) is an optimization technique for local allocation of optimizer parameters. When ZeRO-1 (or other local parameter allocation optimization techniques) is not enabled in DP, each DP rank has the same parameter space. In this case, the mapping scheme generation can disregard the DP dimension, and the parameter distribution requirements of the DP dimension can be met by simply copying model parameters between GPUs after the model parameter migration / transfer based on the mapping scheme is completed.

[0054] When ZeRO-1 optimization is enabled, the optimizer parameters are split across different DP ranks. The online switching method disclosed herein may further include: in response to enabling local allocation of optimizer parameters, determining the optimizer state parameter distribution scheme (i.e., determining how the optimizer state parameters are divided and allocated to each new working node) based on the data-parallel DP splitting information and the splitting method and mapping scheme; and migrating the optimizer state parameters to the corresponding new working nodes based on the optimizer state parameter distribution scheme. For example, when DP supports ZeRO-1 splitting and includes the CP (ContextParallelism) domain and the Group ZeRO domain, the parallel metadata extracted for model parameters may also include optimizer state splitting information at the splitting granularity. After obtaining the model parameters at any rank through steps (1) to (3) above, these parameters are used as input to simulate the method of constructing a model in a training framework (e.g., Megatron), and the ZeRO-1 optimizer parameter distribution is obtained based on the parallel metadata. Then, based on the obtained optimizer parameter distribution, an online switching scheme including optimizer parameter allocation can be implemented.

[0055] During online switching of parallel training strategies, each worker node typically corresponds to one of three states: a continuing node, a decommissioned node, or a newly used node. Here, a continuing node is a worker node that participates in parallel model training both before and after the strategy switch. In this case, the "old worker node" becomes a "new worker node" after the switch. A newly used node is a worker node newly added to the parallel training under the new strategy. A decommissioned node is a worker node that participated in training under the old strategy but not under the new strategy. For example, in an online parallel strategy switch where the worker node topology remains unchanged (e.g., both the old and new strategies use the same 64 GPUs, but one or more parallel dimensions change), all GPUs in the new GPU topology are continuing nodes. However, in an online scaling-up parallel strategy switch (e.g., the old strategy uses 64 GPUs, the new strategy requires 128 GPUs, requiring an additional 64 GPUs), 64 GPUs in the new GPU topology are continuing nodes, and the newly added 64 GPUs are newly used nodes. In the case of performing a parallel strategy switch for online scaling down (for example, the old strategy uses 64 GPUs to execute, and the new strategy requires 32 GPUs to execute), the 32 GPUs in the new GPU topology are all reused nodes, and the 32 GPUs that no longer participate in the training of the new strategy are exited nodes.

[0056] On the inherited nodes, both old and new training states coexist. Here, the training state may include model parameters, and in some embodiments, may also include parameters such as optimizers and learning rates. The old and new training states correspond to the training states required for parallel training of the inherited nodes under the old and new parallel strategies, respectively. Since both old and new training states coexist on the inherited nodes, in some embodiments, the parameter transfer process can be optimized to minimize memory shortages caused by the simultaneous existence of both old and new training states in the GPU memory. Therefore, when the new working node includes the inherited node, during the model parameter migration process, a transfer at the model parameter granularity is performed. This transfer at the model parameter granularity includes: loading the parallel metadata of the current model parameters under the new parallel dimension partition; performing point-to-point transfer to complete the migration of the current model parameters from the old working node to the new working node; and releasing the buffer of the current model parameters in the old working node. Since parameter migration involves releasing the buffer of the source node and loading the buffer of the target node, a transfer at the model parameter granularity needs to be performed on all inherited nodes included in the new working node topology.

[0057] Specifically, if GPU memory has already been allocated when requesting a new training state before parameter transfer, the GPU memory allocated for the new training state can be unloaded first. In embodiments that utilize the global virtual parameter space to determine the sharding method and mapping scheme, the above-mentioned transfer at the model parameter granularity can be implemented by sequentially transferring each parameter in the order of traversing the global virtual parameter space, and can be achieved using P2P (point-to-point) transfer. Before starting the transfer of a parameter, its parallel metadata GPU memory in the new training state is loaded; after the transfer of a parameter is completed, its parallel metadata GPU memory in the old training state is unloaded. This can significantly alleviate the significant increase in GPU memory peaks during online parallel strategy switching. It should be understood that the same parameter can be divided into multiple shards under both the old and new parallel strategies; therefore, the above-mentioned transfer of the parameter at the model parameter granularity may involve multiple source worker nodes and multiple target worker nodes.

[0058] Furthermore, after completing the model parameter transfer, the optimizer state can also be coordinated. Utilizing the aforementioned global virtual parameter space, the GPU location of each optimizer state slice can be precisely determined under any parallel strategy, and P2P transmission can also be used to transfer optimizer parameters between different GPUs to ensure consistency.

[0059] Specifically, by utilizing information from the global virtual parameter space, the distribution of each model parameter and its corresponding optimizer parameter on the GPU can be located. After the optimizer parameters are transferred, global training variables, such as hyperparameters like the learning rate, can be synchronized via lightweight broadcasting, ensuring that the training process remains consistent across all GPUs. Finally, the model weights can be regenerated from the updated optimizer state.

[0060] The model parameters and optimizer parameters mentioned above are all parameters, which are variables in the model that can be learned and updated during the training process. Correspondingly, data is the content input into the model for forward propagation, a signal passed through the training or inference process, and flows between the different layers of the model. Since the loading form of input data may change when switching parallel strategies, data loaders and data iterators under different parallel strategies can be constructed to ensure that the input data under different parallel strategies meets the training expectations. To this end, a unified dataset can be created and cached for the data format and structure of non-parametric data involved in the training of deep learning models. Non-parametric data includes input data and activation values ​​passed during training. In some embodiments, non-parametric data may also include label data. Accordingly, the online switching method of this disclosure may also include: reconstructing the data loader and data iterator based on the cached unified dataset, according to the new parallel dimension partitioning information and the new working node topology information, to provide input data in the parallel model training under the new parallel dimension partitioning executed on the new working node.

[0061] It should be understood that the execution and switching between the old and new parallel strategies are essentially completed by the training processes on the worker nodes. In a distributed training architecture, training processes and worker nodes are usually bound one-to-one, meaning each GPU corresponds to one training process; at the same time, in each round of training, each process is assigned a unique rank number to identify its position in the cluster.

[0062] Switching between parallel strategies not only involves adjusting the rules for partitioning parallel dimensions such as TP, PP, or EP, but also typically involves changes in the scale of computing resources used for training, i.e., scaling. Scaling up refers to increasing available computing resources, often manifested as an increase in the number of worker nodes used for training, such as increasing the number of GPUs used for training from 64 to 128; scaling down refers to reducing available computing resources, often manifested as a decrease in the number of worker nodes used for training, such as reducing the number of GPUs used for training from 64 to 32.

[0063] Depending on whether the maximum available GPU cluster size remains stable during model training, the switching of parallel strategies involves two different scenarios: intra-process reconfiguration and inter-process reconfiguration, each corresponding to different process scheduling logic.

[0064] In-process reconfiguration is suitable for situations where the maximum available GPU cluster size remains unchanged. It is particularly well-suited for handling latency-sensitive reinforcement learning (RL) workloads, allowing worker processes to switch parallel strategies without restarting existing processes. Specifically, in this scenario, even if some GPUs are idle during certain training rounds due to parallel strategy adjustments, their corresponding training processes remain alive and can directly participate in training when the strategy switches in subsequent rounds. For example, if the old parallel strategy requires a topology of 64 GPUs, while the new parallel strategy requires a topology of 128 GPUs, the additional 64 GPUs needed for the switch do not correspond to newly created training processes. These processes participated in training in earlier rounds of the current training task and were not destroyed; they were simply idle during the old strategy phase. Therefore, this parallel strategy switch is an internal switch within all existing processes and does not involve the creation of new training processes. In this scenario, the nodes that are used as described above are still considered used nodes, and the training processes that continue on them can be called used processes; the nodes that are newly used can be further called reused nodes (i.e., nodes that are reactivated to participate in training and whose training processes are reused), and the training processes that continue on them can be called reused processes; the nodes that are exited can be further called standby nodes (i.e., nodes that are not activated for training in the next round, but the training processes that continue on them can wait for future use), and the training processes that continue on them can be called standby processes.

[0065] In this mode, since model training usually switches between several known parallel strategies, the communication group manager can pre-build multiple candidate parallel strategy communication groups, and the operation of determining a new communication group in step S110 can correspond to selecting the candidate communication group that matches the new parallel dimension partitioning information and the new working node topology information from multiple candidate communication groups as the new communication group.

[0066] In one embodiment, the overhead of creating and destroying communication groups during training can be reduced by maintaining a communication group pool. For a communication group, its member list and backend are fixed. The communication group pool uses the member list and backend as cache keys to record and query the corresponding communication group, thereby reducing the overall number of communication groups. During model training, communication groups (e.g., NCCL communication groups) typically consume GPU memory resources. Switching between multiple parallel strategies within the same training process can lead to excessive NCCL GPU memory usage. Based on the communication group pool design, NCCL communication groups can be rebuilt, thus releasing the GPU memory space occupied by them.

[0067] During in-process reconfiguration, the training process on the new worker node switches communication groups, rebuilds the parallel topology, and initiates an online resharding and migration pipeline to migrate optimizer shards without restarting the process. Simultaneously, tensor lifecycles can be managed (e.g., using the training state management module below) to minimize memory spikes during model repartitioning.

[0068] Inter-process reconfiguration often corresponds to situations where the maximum available GPU cluster size cannot remain unchanged. For example, when training using a cloud platform, once a GPU does not participate in parallel training in the current iteration, its related training process must be destroyed. Similarly, taking the example of switching from 64 GPUs executing the old strategy to 128 GPUs executing the new strategy, if the new strategy continues to use the original 64 GPUs of the old strategy, an additional 64 GPUs are needed. However, there are no surviving training processes on these additional 64 GPUs. When creating new processes for them, the execution of the new parallel strategy needs to be implemented. Therefore, this kind of parallel strategy switching involves switching from existing processes to newly created processes, that is, it involves the creation of new processes. In this scenario, the nodes that are still in use are called in-use nodes, and the training processes that continue on them can be called in-use processes; the nodes that are newly used can be further called newly created nodes (i.e., working nodes that did not participate in parallel model training before the old parallel training strategy was switched and have no training processes continuing on them), and the training processes that are newly created on them can be called newly created processes; the nodes that are exited can be further called abandoned nodes (i.e., nodes that are not used for training in the new round and whose training processes are destroyed), and the training processes that no longer continue on them can be called abandoned processes.

[0069] It should be understood that inter-process reconfiguration can also involve situations where additional nodes need to be created because GPUs used in the previous round need to be reclaimed. For example, when scaling up from 64 GPUs executing the old strategy to 128 completely different GPUs executing the new strategy, there are no reused nodes; all 128 GPUs are newly created nodes, requiring the creation of a new process on each GPU—that is, creating 128 new processes. As another example, when scaling down from 64 GPUs executing the old strategy to 32 completely different GPUs executing the new strategy, there are no reused nodes; all 32 GPUs are newly created nodes, requiring the creation of both a joint process and a new process on each GPU to switch between the old and new parallel strategies.

[0070] To minimize online training interruptions caused by the switchover, during inter-process reconfiguration, the process creation, re-rendezvous, and communication group initialization of the new process can be performed in parallel with the previous iteration's operations based on the old parallel strategy (i.e., forward / backward computation in the previous iteration). The new process independently performs the bootstrap operation, while the training process of the old worker node continues computation and periodically polls for the ready signal. Once the new process has completed the above operations and is ready to participate in the online switchover of the parallel strategy, the training process of the old worker node can enter a brief safe point to perform online resharding and migration pipelines, and then resume training under the new configuration, i.e., a new round of parallel training is performed by the new worker node.

[0071] To mask the overhead of starting new processes during inter-process scaling and reconfiguration, the online switching method of this disclosure may further include: in response to determining that the number of new worker nodes is greater than the number of old worker nodes (which can be determined based on the topology information of the old and new worker nodes) and that the new worker nodes include newly created nodes, the following operations are performed on the newly created nodes: during the parallel training process performed on the old worker nodes (i.e., when performing the previous round of parallel training based on the old parallel strategy), a new process is created and an initialization operation is performed, and the local communication group is initialized and a communication connection is established (hereinafter referred to as "connection") according to the new communication group information. That is, the creation of the new process and the activation and connection of the local communication group are performed in parallel with the previous round of iterative training.

[0072] For ease of understanding, Figure 2 A flowchart illustrating an inter-process scaling scenario according to an embodiment of this disclosure is shown. In the diagram, "old process" refers to the training process that participated in training under the old parallel strategy, and regardless of whether the "old process" is reused or deprecated in the new parallel strategy, it will coexist with the new process before the old process transmits training-related data to the new process.

[0073] like Figure 2As shown, the inter-process scaling operation begins upon receiving the scaling signal. While the old process continues training, a new process is started and a switch to not allocate model memory is completed (e.g., a switch in meta-device state). At this time, the communication group of the new process completes initialization and actual connection establishment. After the initialization and communication group connection of the new process are completed, the old process is notified that the new process has completed communication group connection and can perform data transmission. After receiving the data transmission signal, the old process performs the transmission of training-related data (e.g., transmission via IPC) before the start of a new round of training iterations. If the old process participates in parallel training under the new parallel strategy, it subsequently becomes a reused process, participating in online resharding and migration pipeline operations and starting a new round of training together with the new process as a training process on the new worker node. However, if the old process does not participate in parallel training under the new parallel strategy (i.e., it is not a reused process), it is subsequently deprecated and exits. After receiving IPC data, the newly created process completes the state switch (from the meta-device state where it does not request model memory usage to the GPU-device state where it requests model memory usage) and begins to participate in a new round of parallel training.

[0074] In inter-process scaling down scenarios, in addition to creating a new process on the new node, a connection process is also needed for scaling down and partitioning. Similar to inter-process scaling up scenarios, the old process is responsible for maintaining training until IPC data transmission, while the new process is responsible for receiving IPC data and starting training based on the new parallel strategy. The additional connection process is responsible for handling the switchover. Upon receiving the scaling down signal, the connection process and the new process are started synchronously on the new node. The connection process completes the meta-device state switch, realizing the actual connection of the communication group, and then performs scaling down and partitioning by receiving IPC data from the old process. After completing the partitioning, the connection process transmits data to the new process via IPC technology and then exits the connection process.

[0075] By performing the above operations, it can be ensured that the GPU survives in scenarios where processes are switched, interrupting training only when real model splitting occurs, and performing model training at other times.

[0076] This disclosure can also be implemented as a deep learning model training system. Figure 3 A schematic diagram of a deep learning model training system according to an embodiment of the present disclosure is shown. The deep learning model training system of the present disclosure is used to manage the distributed training of deep learning models, especially large-scale deep learning models, and is particularly useful for online switching of parallel training strategies.

[0077] like Figure 3 As shown, system 300 includes a communication management module 310 and a reslicing and migration module 320.

[0078] The communication management module 310 is used to determine the new communication group information based on the new parallel dimension partitioning information and the new working node topology information.

[0079] The reslicing and migration module 320 is used to determine the slicing method of each model parameter of the deep learning model based on the new parallel dimension partitioning information and the new working node topology information, further generate a mapping scheme for model parameter slicing migration based on the old parallel dimension partitioning information and the old working node topology information, and migrate the model parameters to the corresponding new working nodes based on the slicing method and the mapping scheme. Therefore, the new working node can enable the corresponding communication group based on the new communication group information, and the enabled communication group and the model parameters migrated to the new working node can be used by the new working node to perform parallel model training under the new parallel dimension partition.

[0080] System 300 may also include a state management module 330. The state management module 330 manages the states during the training process, with each parallel strategy corresponding to a set of states. In situations where switching between multiple parallel strategies is required during model training (e.g., during the training of a large-scale reinforcement learning model), this state management module can maintain multiple sets of states. For example, it can store state information describing each set of states, ensure that only one set of states is activated, and prevent inactive states from consuming GPU memory.

[0081] In one embodiment, the state management module 330 can be used to cache a unified dataset and, based on the unified dataset, reconstruct the data loader and data iterator according to the new parallel dimension partitioning information and the new worker node topology information, to provide input data for parallel model training under the new parallel dimension partitioning performed on the new worker node. The unified dataset is created and cached for the data format and structure of the non-parametric data involved in the deep learning model training process. State management of the data can be, for example, as follows: Figure 5 Other state management submodules 533 in the process are executed.

[0082] In one embodiment, the state management module 330 can be used for memory management of the training state. For example, it may include the training state management submodule 531 as described below, thereby releasing the memory requested for training of the old working node and loading the training state information required for training of the new working node before the reslicing and transfer module 320 transfers the model parameters, and requesting the memory required for training of the new working node after the reslicing and transfer module 320 transfers the weight parameters.

[0083] Figure 4 A flowchart illustrating a state switching operation performed by a state management module according to an embodiment of the present disclosure is shown.

[0084] exist Figure 4 In the example, the states during the parallel training of a large model can include: (1) the training state, which includes the model, optimizer, and learning rate; (2) the communication group state, which includes all communication group variables used under a given parallel strategy; and (3) other states, such as the training data loader. Accordingly, the state management module can perform training state management, communication group state management, and other state management.

[0085] The training state is not only the object updated during model training, but also the object of parameter partitioning. A complete set of training states can be maintained for each parallel strategy. For example... Figure 4 As shown, when starting an online policy switch, the model state needs to be checked. The training state management submodule is primarily responsible for managing the GPU memory usage of the training state. This submodule first releases the GPU memory occupied by the model in the old training state; then it loads the new training state, without allocating GPU memory during the loading process. After parameter resharding and migration are completed, the optimizer's GPU memory space in the old training state is released, and the optimizer's GPU memory space in the new training state is loaded. The new training state allocates its corresponding model GPU memory space only after the optimizer's GPU memory is fully loaded. The training state management module ensures that the peak GPU memory usage does not increase when two sets of training states coexist during parallel policy switching. Furthermore, the training state management submodule can generate not only GPU-device training states but also meta-device training states. The latter does not occupy GPU memory space but retains information such as model structure and parameter attributes, which can be used to optimize GPU memory usage during policy switching. Meanwhile, as... Figure 4 As shown in the green box, peak memory usage is optimized by releasing memory first and then allocating it later.

[0086] The communication group state management submodule is responsible for communication group state management. In one embodiment, the communication group state can be a series of communication-related global variables maintained by the MPU (Model Parallelism Unit) in the training framework, including communication group variables and communication group member list variables. When registering a new parallel strategy, its communication group state must first be registered. After switching to the current global communication group, the communication group state management submodule executes the MPU initialization process and saves the global variables in the MPU. When a parallel strategy is activated, the communication group state management module applies its corresponding global variables to the MPU module.

[0087] Other state management submodules are responsible for other state management tasks. For example, changing the parallel strategy may alter the way input data is loaded. These submodules can cache a global dataset; for instance, they can create and cache a unified dataset for the non-parametric data format and structure involved in the training of the deep learning model to minimize redundant dataset construction. Furthermore, this dataset can be used to construct data loaders and iterators for different parallel strategies, ensuring that the input data under different parallel strategies meets training expectations.

[0088] Thus, the activation state is updated through the parallel operation of the three sub-modules, and the state switching in the online policy switching operation is completed.

[0089] and Figure 4 The detailed operations involved in the state management module are similar, and the communication management module and the re-slicing and migration module also involve more detailed operations. Figure 5 A schematic diagram of a deep learning model training system according to an embodiment of the present disclosure is shown. Figure 3 Similarly, system 500 includes a communication management module 510, a reslicing and migration module 520, and a state management module 530, but the internal composition of each module is shown in more detail.

[0090] The communication management module 510 is responsible for communication management during parallel strategy switching and scaling up / down processes. Depending on the application scenario, it is further divided into an intra-process communication management sub-module 511 and an inter-process communication management sub-module 512.

[0091] The in-process communication management submodule 511 is used for communication management in the "in-process reconfiguration" scenario described above. To enable switching between different parallel strategies and training nodes within the same training process, submodule 511 needs to manage a dynamic global communication group. This means maintaining a dynamic global communication group and ensuring that the original global communication group (e.g., torch.distributed.GroupMember.WORLD) is not destroyed or rebuilt. Submodule 511 can generate different dynamic global communication groups based on different worker node topology information and parallel partitioning information, and replace them when, for example, PyTorch communication methods use the global communication group as a default parameter.

[0092] Furthermore, submodule 511 also performs communication group pool cache management to reduce the overhead of communication group creation and destruction during training by maintaining the communication group pool. For a communication group, its communication member list and communication backend are fixed. The communication group pool uses the communication member list and communication backend as cache keys to record and query the corresponding communication group, thereby reducing the overall number of communication groups. Additionally, during model training, NCCL communication groups typically consume GPU memory resources. Switching between multiple parallel strategies in the same training process may lead to excessive NCCL GPU memory usage. Therefore, submodule 511 performs NCCL GPU memory optimization management. Submodule 511 can rebuild NCCL communication groups based on the communication group pool design, thereby releasing the GPU memory space occupied by the NCCL communication groups.

[0093] The inter-process communication management submodule 512 is used for communication management in the "inter-process reconfiguration" scenario described above.

[0094] In inter-process reconfiguration scenarios, the startup and initialization of new processes can take a long time. To mask this time consumption and allow surviving GPUs to continue training, submodule 512 provides a mechanism to mask the process startup overhead.

[0095] Specifically, in expanding communication management, the following can be executed: Figure 2 The method shown is as described above; however, in scaling down communication management, scaling down can be performed through the connection process described above. In both scaling up and scaling down communication management, the IPC data transfer mechanism can be used to transfer training-related data to the newly created process, and then the process exits. Therefore, submodule 512 can ensure that in scenarios involving process switching, the surviving GPU only interrupts training when a real model split occurs, and performs model training at other times.

[0096] The reslicing and transfer module 520 is responsible for reslicing and transferring model parameters from the old strategy to the new strategy, and can support complex parallel slicing, namely TP, PP, DP, and EP and their combinations. DP also supports ZeRO-1 slicing and includes CP and Group ZeRO domains. During model training, the effects of different parallel techniques on parameter distribution are orthogonal. Therefore, we can first analyze a single parallel dimension to obtain its corresponding metadata, and then complete the mapping of parameter distribution by constructing a global virtual parameter space; finally, we perform parameter transfer. When performing parameter transfer, for the inherited nodes, we can perform the memory optimization processing described above. For example, we can perform actual data transfer of model parameters at the granularity of a single parameter (e.g., point-to-point transfer from the source worker node to the target worker node), which effectively alleviates the problem of excessively high memory peaks, thereby ensuring the smooth operation of the training process.

[0097] The state management module 530 may include a training state management submodule 531, a communication group state management submodule 532, and other state management submodules 533. The specific operations of each submodule are integrated... Figure 4 The details will not be elaborated here.

[0098] The online switching method and deep learning model training system according to the parallel training strategy of this disclosure have been described in detail above with reference to the accompanying drawings. By enabling the required communication group, reconstructing the parameter mapping, and performing resharding and transfer with minimal cost at runtime, online fast switching based on the new parallel training strategy can be achieved.

[0099] Parallel strategy switching in related technologies often employs offline methods, requiring disk read / write operations through checkpoints. This not only results in high communication latency but also causes GPU idleness during switching, reducing overall computational efficiency. The online switching method disclosed in this paper avoids disk read / write dependencies through direct communication between GPUs, significantly improving communication efficiency. Online switching also reduces or even eliminates the overhead of process switching. In intra-process reconfiguration scenarios, process reuse essentially completely avoids the additional overhead of process startup and destruction; in inter-process reconfiguration scenarios, parallel processing with the previous training round effectively masks the latency of starting and switching processes, thereby improving GPU utilization efficiency.

[0100] Furthermore, this disclosure supports online switching of complex parallel strategies, including multiple dimensions such as TP, PP, EP, and DP (which may include the CP domain). It can accurately solve the distribution of model parameters under these complex parallel strategies by constructing a global virtual parameter space, thereby effectively solving the problem that a single GPU cannot obtain the parameter distribution of other GPUs in traditional methods, making parallel switching in large-scale distributed training more efficient.

[0101] This disclosure is highly flexible, supports scaling up and down as well as in-situ parallel strategy switching, and is compatible with general training frameworks (such as Megatron). By incorporating the function of dynamically adjusting the parallel strategy at runtime, the training framework can more quickly and flexibly switch parallel strategies when performing training tasks, adapting to training tasks of different scales and needs.

[0102] To optimize GPU memory usage, this disclosure also introduces training state management to ensure that multiple training states can be saved simultaneously within the same training process without significantly increasing GPU memory usage. By finely controlling the allocation and release of GPU memory, the system can perform data transfer at the granularity of a single parameter during parallel strategy switching, effectively alleviating the problem of excessively high GPU memory peaks and thus ensuring the smooth operation of the training process.

[0103] This disclosure achieves efficient online parallel switching by eliminating disk read / write operations in traditional offline switching methods and instead relying on direct communication between GPUs. It also supports complex parallel strategies and scaling capabilities, exhibiting strong adaptability to large-scale model training frameworks such as Megatron. This technology enables large-scale model training to complete parallel strategy switching within milliseconds. For example, in a 64-card training environment with Llama2-70B, the overhead of a single switch is only 300ms, significantly improving training efficiency compared to the tens of minutes required by checkpointing methods.

[0104] In one embodiment, this disclosure can also be implemented as a non-transitory machine-readable storage medium storing executable code that, when executed by a processor of an electronic device, causes the processor to perform the method described above.

[0105] In one embodiment, this disclosure can also be implemented as a computer program product, including computer program instructions that, when executed by a processor, implement the method described above.

[0106] Those skilled in the art will also understand that the various exemplary logic blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both.

[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems and methods according to various embodiments of the present invention. For this purpose, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0108] Various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. An online switching method for parallel training strategies, the method being used for deep learning model training and comprising: Based on the new parallel dimension partitioning information and the new working node topology information, determine the new communication group information; Based on the new parallel dimension partitioning information and the new working node topology information, the sharding method of each model parameter of the deep learning model is determined, and further based on the old parallel dimension partitioning information and the old working node topology information, a mapping scheme for model parameter sharding migration is generated. as well as Based on the aforementioned sharding method and mapping scheme, the model parameters are migrated to the new working node; Wherein, in response to determining that the number of new working nodes is less than the number of old working nodes and that the new working nodes include newly created nodes, the migration of model parameters to new working nodes based on the sharding method and mapping scheme includes: While the old process, which executes the old parallel training strategy on the old working node, is maintaining training, the connection establishment process and the new process are simultaneously started on the new node. The connection establishment process completes the switch without requesting model memory usage and realizes the real connection establishment of the communication group; After the connection establishment process completes the actual connection establishment of the communication group, the old process transmits training-related data to the connection establishment process through inter-process communication before the start of a new round of training iteration. The connection establishment process receives the training-related data and performs scaling down and segmentation on the training-related data; The connection-establishing process transmits the training-related data that has completed the scaling-up and splitting to the newly created process through inter-process communication, and then exits the connection-establishing process; The newly created process is used to execute a new round of parallel model training under the new parallel dimension partition based on the training-related data that has been shrunk and split. The new communication group information is used to enable the communication group. The enabled communication group and the model parameters migrated to the new working node are used by the new working node to execute parallel model training under the new parallel dimension partition. The newly created node is a working node that did not participate in parallel model training before the parallel training strategy was switched and has no training process existing on it. The old process is the training process that executed the old parallel training strategy on the old working node.

2. The method according to claim 1, wherein, The new parallel dimension partitioning information includes at least one of the following: Tensor parallel TP segmentation information Pipeline parallel PP segmentation information, and Expert parallel EP allocation information; Based on the new parallel dimension partitioning information and the new worker node topology information, the new communication group information is determined to include: A new global communication group is determined based on the new working node topology information, and Based on the new parallel dimension partitioning information and the new working node topology information, at least one of the following is determined: TP subgroup, PP subgroup, and EP sub-communication group.

3. The method according to claim 1 or 2, wherein, Based on the new parallel dimension partitioning information and the new worker node topology information, the sharding method for each model parameter of the deep learning model is determined, and further, based on the old parallel dimension partitioning information and the old worker node topology information, a mapping scheme for model parameter sharding migration is generated, including: Based on the new parallel dimension partitioning information, the parallel metadata of each virtual parameter in the global virtual parameter space is extracted, wherein the global virtual parameter space is constructed as a unified logical view of the model parameters of the deep learning model. The physical slices of model parameters held by the old worker nodes that perform parallel model training under the old parallel dimension are bound to the corresponding virtual parameter entries in the global virtual parameter space to obtain binding information. The global virtual parameter space is traversed based on the parallel metadata and the binding information to determine the sharding method and mapping scheme.

4. The method of claim 3, wherein, Based on the new parallel dimension partitioning information, the parallel metadata extracted for each virtual parameter in the global virtual parameter space includes at least one of the following: The TP parameter distribution state of this virtual parameter is determined based on the TP segmentation information, and is used as TP parallel metadata. The new working node where the virtual parameter resides is determined based on the PP splitting information, and this node serves as the PP parallel metadata. The new working node where the virtual parameter is located is determined based on the EP allocation information, and is used as EP parallel metadata. Furthermore, based on the parallel metadata, the global virtual parameter space is traversed to determine the partitioning method of each model parameter and the mapping scheme, including at least one of the following: Based on the TP segmentation information, all virtual parameters are reconstructed according to their respective TP parallel metadata dimensions to obtain the virtual parameter distribution status. Based on the PP partitioning information, the PP parallel metadata for each virtual parameter is aggregated layer by layer to obtain the virtual parameter combinations on each new working node participating in the layer computation, and Based on the EP allocation information, the virtual parameters, which are non-shared expert parameters, are divided according to the EP parallel metadata to obtain the expert parameter combinations on each new working node participating in the expert model calculation.

5. The method of claim 1 or 2, further comprising: In response to enabling local allocation of optimizer parameters, the optimizer state parameter distribution scheme is determined based on the data parallel DP splitting information and the splitting method and mapping scheme. as well as Based on the optimizer state parameter distribution scheme, the optimizer state parameters are migrated to the corresponding new working nodes.

6. The method according to claim 1 or 2, wherein, Based on the aforementioned sharding method and mapping scheme, migrating model parameters to the new working node includes: In response to determining that the new working node includes the inherited node, during the migration of model parameters to the new working node, a transfer at the model parameter granularity is performed, and The transmission at the model parameter granularity includes: Load the parallel metadata of the current model parameters under the new parallel dimension partition; Perform point-to-point transmission to complete the migration of the current model parameters from the old working node to the new working node; and Release the buffer of the current model parameters in the old worker node. The "reused node" refers to a working node that participates in parallel model training both before and after the parallel training strategy switch.

7. The method according to claim 1 or 2, further comprising: In response to determining that the number of new working nodes is greater than the number of old working nodes and that the new working nodes include newly created nodes, the following operation is performed on the newly created nodes: During parallel training on the old working node, a new process is created and initialization is performed. The local communication group is initialized and a communication connection is established based on the new communication group information.

8. The method according to claim 1 or 2, further comprising: A unified dataset based on caching is used to reconstruct the data loader and data iterator according to the new parallel dimension partitioning information and the new worker node topology information, in order to provide input data in the parallel model training under the new parallel dimension partitioning. The unified dataset is created and cached for the data format and structure of nonparametric data involved in the training of the deep learning model.

9. A deep learning model training system for online switching of parallel training strategies, and comprising: The communication management module is used to determine new communication group information based on the new parallel dimension partitioning information and the new working node topology information; as well as The reslicing and migration module is used to determine the slicing method for each model parameter of the deep learning model based on the new parallel dimension partitioning information and the new working node topology information, further generate a mapping scheme for model parameter slicing migration based on the old parallel dimension partitioning information and the old working node topology information, and migrate the model parameters to the corresponding new working nodes based on the slicing method and the mapping scheme. Wherein, in response to determining that the number of new working nodes is less than the number of old working nodes and that the new working nodes include newly created nodes, the migration of model parameters to new working nodes based on the sharding method and mapping scheme includes: While the old process, which executes the old parallel training strategy on the old working node, is maintaining training, the connection establishment process and the new process are simultaneously started on the new node. The connection establishment process completes the switch without requesting model memory usage and realizes the real connection establishment of the communication group; After the connection establishment process completes the actual connection establishment of the communication group, the old process transmits training-related data to the connection establishment process through inter-process communication before the start of a new round of training iteration. The connection establishment process receives the training-related data and performs scaling down and segmentation on the training-related data; The connection-establishing process transmits the training-related data that has completed the scaling-up and splitting to the newly created process through inter-process communication, and then exits the connection-establishing process; The newly created process is used to perform a new round of parallel model training under the new parallel dimension partitioning, based on the training-related data that has undergone the shrinking and splitting process. The new communication group information is used to enable the communication group. The enabled communication group and the model parameters migrated to the new working node are used by the new working node to perform parallel model training under the new parallel dimension partition. The newly created node is a working node that did not participate in parallel model training before the parallel training strategy was switched and has no training process existing on it. The old process is the training process that executes the old parallel training strategy on the old working node.

10. The system of claim 9, further comprising: The state management module is used to cache a unified dataset and reconstruct the data loader and data iterator based on the unified dataset according to the new parallel dimension partitioning information and the new worker node topology information, so as to provide input data in the parallel model training under the new parallel dimension partitioning. The unified dataset is created and cached for the data format and structure of nonparametric data involved in the training of the deep learning model.

11. The system of claim 9 or 10, further comprising: The state management module is used to release the GPU memory requested for training on the old working node and load the training state information required for training on the new working node before the reslicing and migration module migrates the model parameters, and to request the GPU memory required for training on the new working node after the reslicing and migration module migrates the model parameters.