Virtual machine idle computing power-based distributed large model training method and system
Through the distributed large-model training system, a central dispatch server is used to uniformly manage the global virtual machine computing power, solving the problems of high centralized training costs and low resource utilization, and achieving efficient and stable large-model training.
Patent Information
- Application Number
- CN202510727205.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-26
AI Technical Summary
Existing large-scale model training relies on centralized data centers, which has high costs, insufficient resource utilization, fault tolerance and reliability issues, and high complexity in multi-GPU cluster training.
By building a distributed large-scale model training system, using a central scheduling server to uniformly manage global virtual machine computing power, dynamically allocating subtasks, setting timeouts and retry mechanisms, efficient resource utilization and environmental consistency are achieved, and Docker technology is used to ensure a unified operating environment.
It reduces hardware investment and maintenance costs, improves resource utilization and system stability, provides an elastically scalable computing platform, and ensures model consistency and computing efficiency.
Smart Images

Figure CN120704865A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of model training technology, and in particular to a training method and system for a distributed large model based on the idle computing power of a virtual machine. Background Art
[0002] Currently, large-scale model training mainly relies on centralized data centers and computing power leasing services. These systems have the advantages of high performance, low latency, and unified management, but also have the following disadvantages:
[0003] High cost: The investment and maintenance costs of centralized high-performance servers and GPU clusters are high.
[0004] Insufficient resource utilization: Idle computing resources are difficult to fully utilize, and scalability is limited by the size of the data center.
[0005] Fault tolerance and reliability issues: Node disconnection and high failure rates affect overall task stability and computing efficiency.
[0006] Currently, the training of large models in the market, such as ChatGPT, Tongyi Qianwen, and Wenxin Yiyan, usually relies on multi-card GPU clusters (for example, high-end graphics cards). It may take weeks to months to train a large model on a single high-end graphics card, and the coordination of multiple cards requires efficient synchronization, which further increases the complexity of system construction and operation and maintenance.
[0007] Therefore, it is necessary to provide a new technical solution to improve one or more problems existing in the above solutions.
[0008] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0009] The purpose of this application is to provide a training method and system for distributed large models based on the idle computing power of virtual machines, thereby overcoming one or more problems caused by the limitations and defects of related technologies to at least a certain extent.
[0010] According to a first aspect of an embodiment of the present application, a method for training a distributed large model based on idle computing power of a virtual machine is provided. The method is applied to a distributed large model training system, wherein the training system includes a central scheduling server, one or more cloud servers, and a data storage center. The central scheduling server is connected to the one or more cloud servers and the data storage center, respectively. Each cloud server includes multiple computing nodes, each computing node being a virtual machine. The data storage center is used to store training data and calculation results. The method includes:
[0011] After the cloud server starts each computing node, the central scheduling server registers each computing node based on the computing node registration request sent by the cloud server, and builds a unified operating environment for each computing node;
[0012] The central scheduling server splits the overall training task of the distributed large model into multiple subtasks;
[0013] The central scheduling server dynamically assigns the plurality of subtasks to the corresponding computing nodes based on the comprehensive score of each computing node sent by the cloud server, and sets a timeout period and a retry mechanism for the subtasks; wherein the comprehensive score of each computing node is calculated based on the idle computing power and network status of each computing node;
[0014] The central scheduling server distributes the initialization parameters of the distributed large model to the computing nodes that receive the subtasks, so that the computing nodes that receive the subtasks train the received subtasks through the training data, and calculate the calculation results of the computing nodes to obtain calculation results;
[0015] The central scheduling server receives the calculation results sent by the cloud server, aggregates the calculation results of all the computing power nodes, and updates the global model parameters;
[0016] The central scheduling server feeds back the aggregated global model parameters to each of the computing nodes, so that each of the computing nodes performs training based on the latest global model parameters in the next round of calculation until the distributed large model converges and the training ends.
[0017] In an embodiment of the present application, the step of calculating the comprehensive score of each computing power node based on the idle computing power and network status of each computing power node includes:
[0018] Normalizing the GPU floating-point computing capability, the number of CPU cores, and the available memory of each computing power node by the cloud server to obtain the normalized GPU floating-point computing capability, the normalized number of CPU cores, and the normalized available memory of each computing power node;
[0019] The cloud server calculates the idle computing power score of each computing power node according to the normalized GPU floating-point computing power, the normalized number of CPU cores, and the normalized available memory of each computing power node, as well as their corresponding weights; wherein the idle computing power score of each computing power node is used to quantify the idle computing power of each computing power node;
[0020] Normalizing the average bandwidth, delay, and packet loss rate of each computing power node by the cloud server to obtain the normalized average bandwidth, normalized delay, and normalized packet loss rate of each computing power node;
[0021] Calculating, by the cloud server, a network status score for each computing power node based on the normalized average bandwidth, normalized latency, and normalized packet loss rate of each computing power node, as well as their corresponding weights; wherein the network status of each computing power node is quantified by the network status score of each computing power node;
[0022] The cloud server calculates the comprehensive score of each computing power node based on the idle computing power score and network status score of each computing power node, as well as their respective corresponding weights, and sends the comprehensive score of each computing power node to the central scheduling server.
[0023] In an embodiment of the present application, the central scheduling server dynamically assigns the plurality of subtasks to the respective corresponding computing nodes based on the comprehensive score of each computing node sent by the cloud server, and sets the timeout time and retry mechanism of the subtasks, further comprising:
[0024] The central scheduling server sorts all the subtasks to be assigned based on the priority of each computing node sent by the cloud server; wherein the priority of each subtask is calculated according to the data volume and model block of each computing node;
[0025] The central scheduling server allocates all the subtasks to be allocated to their respective appropriate computing nodes according to the allocation strategy; wherein the allocation strategy is: the subtasks with high priority are allocated to the computing nodes with high comprehensive scores, the subtasks with medium priority are allocated to the computing nodes with medium comprehensive scores, and the subtasks with low priority are allocated to the computing nodes with low comprehensive scores.
[0026] In an embodiment of the present application, the step of calculating the priority of each subtask based on the data volume and model block of each subtask includes:
[0027] Normalizing the data volume and model block of each subtask by the cloud server to obtain the normalized data volume and normalized model block of each subtask;
[0028] The cloud server calculates the priority score of each subtask based on the normalized data volume, normalized model block and corresponding weight of each subtask; wherein the priority of each subtask is quantified by the priority score of each subtask.
[0029] In an embodiment of the present application, the step of setting the timeout period and retry mechanism of the subtask includes:
[0030] The timeout period of each subtask is determined by the estimated execution time of each subtask and a safety factor;
[0031] The retry mechanism is that the computing power node recalculates the subtask it has received within a preset number of retries.
[0032] In an embodiment of the present application, if a computing node fails to upload its calculation result within the timeout period, the central scheduling server reallocates the subtask to another computing node;
[0033] If a subtask is not completed by the corresponding computing power node within the timeout period, a retry mechanism is triggered.
[0034] In an embodiment of the present application, the step of the central scheduling server reallocating the subtask to other computing nodes includes:
[0035] The central scheduling server selects a backup computing node to calculate the subtask based on the calculation progress of the completed subtask and the comprehensive score of the current computing power node, so that the backup computing power node continues to execute the same subtask; wherein, the calculation progress of the subtask is that each computing power node uploads the progress of the corresponding subtask to the central scheduling server within a preset time.
[0036] In the embodiment of the present application, the operation results include key model parameters and intermediate results;
[0037] The step of updating the global model parameters includes:
[0038] The key model parameters are updated in a synchronous updating manner;
[0039] Some of the intermediate results are updated in an asynchronous manner.
[0040] In an embodiment of the present application, the step of sending the calculation result of the computing power node to the central scheduling server includes:
[0041] The calculation results of all the computing nodes are encrypted and transmitted in segments to the central scheduling server through the cloud server.
[0042] According to a second aspect of an embodiment of the present application, a distributed large model training system based on idle computing power of virtual machines is provided, the system comprising:
[0043] A startup registration module is configured to register each computing node and perform environment configuration for each computing node based on a registration request for registering computing nodes sent by the cloud server to the central scheduling server after the cloud server starts each computing node;
[0044] A splitting module, configured to split the overall training task of the distributed large model into multiple subtasks through the central scheduling server;
[0045] an allocation module, configured for the central scheduling server to dynamically allocate the plurality of subtasks to the respective corresponding computing nodes based on the comprehensive score of each computing node sent by the cloud server, and to set a timeout period and a retry mechanism for the subtasks; wherein the comprehensive score of each computing node is calculated based on the idle computing power and network status of each computing node;
[0046] a distribution module, configured for the central scheduling server to distribute the initialization parameters of the distributed large model to the computing nodes that receive the subtasks, so that the computing nodes that receive the subtasks perform training on the received subtasks using the training data, and to calculate the calculation results of the computing nodes to obtain calculation results;
[0047] An aggregation and updating module is used for the central scheduling server to receive the calculation results sent by the cloud server, aggregate the calculation results of all the computing power nodes, and update the global model parameters;
[0048] A feedback module is used for the central scheduling server to feed back the aggregated global model parameters to each of the computing power nodes, so that each of the computing power nodes can be trained based on the latest global model parameters in the next round of calculation until the distributed large model converges and the training is completed.
[0049] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:
[0050] In one embodiment of the present application, through the above method, by constructing a distributed large model training system, the virtual machine computing power on the cloud servers with idle computing power worldwide can be integrated into the training system architecture of the same platform, and the central scheduling server uniformly manages each computing power node, which can achieve efficient allocation and utilization of resources; building a unified operating environment for each computing power node can ensure that the software and dependent environment of each computing power node are consistent, thereby reducing the system instability caused by environmental differences. Splitting the overall training task of the distributed large model into multiple subtasks not only facilitates the training of the distributed large model, but also ensures the consistency of the model; at the same time, through environmental unification, reasonable task decomposition and dynamic scheduling, it is possible to achieve efficient utilization of the virtual machine computing power distributed on cloud terminal servers in different regions, providing a flexible expansion computing platform for distributed large model training. By setting a reasonable timeout time and retry mechanism, the stability of the system can be improved in the environment of computing power node failure and network delay. Through this application, it is possible to fully utilize the idle computing resources around the world and reduce the hardware investment and maintenance costs required for large-scale model training.
[0051] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0053] Figure 1 A flowchart schematically illustrates a method for training a distributed large model based on idle computing power of a virtual machine in an exemplary embodiment of the present application;
[0054] Figure 2 A schematic diagram schematically illustrates a distributed large model training system in an exemplary embodiment of the present application;
[0055] Figure 3 A block diagram of a distributed large model training system based on idle computing power of virtual machines in an exemplary embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0056] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0057] In addition, the accompanying drawings are merely schematic illustrations of the present application and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0058] This example implementation first provides a distributed large model training method based on the idle computing power of virtual machines. This method is applied to a distributed large model training system. The training system includes a central scheduling server, one or more cloud servers, and a data storage center. The central scheduling server is connected to one or more cloud servers and the data storage center. Each cloud server contains multiple computing nodes, each computing node is a virtual machine. The data storage center is used to store training data and calculation results. Figure 1 and Figure 2 As shown in , the method may include: steps S101 to S106.
[0059] Among them, step S101: after the cloud server starts each computing power node, the central scheduling server registers each computing power node based on the registration request for registering the computing power node sent by the cloud server, and builds a unified operating environment for each computing power node.
[0060] Step S102: The central scheduling server splits the overall training task of the distributed large model into multiple subtasks.
[0061] Step S103: The central scheduling server dynamically allocates multiple subtasks to their corresponding computing nodes based on the comprehensive score of each computing node sent by the cloud server, and sets the timeout time and retry mechanism of the subtasks; wherein, the comprehensive score of each computing node is calculated by the idle computing power and network status of each computing node.
[0062] Step S104: The initialization parameters of the distributed large model are distributed to the computing power nodes that receive the subtasks through the central scheduling server, so that the computing power nodes that receive the subtasks can train the received subtasks through the training data, and calculate the calculation results of the computing power nodes to obtain the calculation results, and send the calculation results to the central scheduling server.
[0063] Step S105: The central scheduling server receives the calculation results sent by the cloud server, aggregates the calculation results of all computing nodes, and updates the global model parameters.
[0064] Step S106: The central scheduling server feeds back the aggregated global model parameters to each computing node, so that each computing node performs training based on the latest global model parameters in the next round of calculation until the distributed large model converges and the training ends.
[0065] In one embodiment of the present application, through the above method, by constructing a distributed large model training system, the virtual machine computing power on cloud servers with idle computing power worldwide can be integrated into the training system of the same platform, and the central scheduling server uniformly manages each computing power node, which can achieve efficient allocation and utilization of resources; building a unified operating environment for each computing power node can ensure that the software and dependent environment of each computing power node are consistent, thereby reducing the system instability caused by environmental differences. Splitting the overall training task of the distributed large model into multiple subtasks not only facilitates the training of the distributed large model, but also ensures the consistency of the model; at the same time, through environmental unification, reasonable task decomposition and dynamic scheduling, it is possible to achieve efficient utilization of the virtual machine computing power distributed on cloud terminal servers in different regions, providing a flexible expansion computing platform for distributed large model training. By setting a reasonable timeout time and retry mechanism, the stability of the system can be improved in the environment of computing power node failure and network delay. Through this application, it is possible to fully utilize the idle computing resources around the world and reduce the hardware investment and maintenance costs required for large-scale model training.
[0066] Below, we will refer to Figures 1 to 2 Each step of the above method in this exemplary embodiment is described in more detail.
[0067] Before discussing the training method of distributed large models based on the idle computing power of virtual machines, the training system is first explained.
[0068] The training system architecture includes a central scheduling server, all idle virtual machines (VMs), and a data storage center. The central scheduling server manages the overall training architecture, decomposing tasks, scheduling nodes, aggregating computational results, and distributing incentives to encourage idle VMs to contribute computing power. The central scheduling server utilizes a high-speed network to maintain communication with each VM, monitoring node status and task progress in real time.
[0069] All computing nodes are located on one or more cloud servers. A computing node is a virtual machine. Distributed computing nodes are composed of virtual machines on the participants' cloud servers. Each computing node is pre-installed with a unified containerized operating environment to support local computing for distributed large-scale model training tasks.
[0070] The data storage center is used to store training data and calculation results, and is also responsible for data backup and secure storage.
[0071] This application integrates the virtual machine computing power on cloud servers with idle computing power worldwide into the same platform architecture, and the central scheduling server scheduler uniformly manages each computing power node to achieve efficient allocation and utilization of virtual machine computing power resources.
[0072] In step S101, after startup, each computing node sends a registration request to the central scheduling server through the dedicated client program pre-installed on each computing node, so as to send information such as the hardware configuration, network status and current load of each computing node to the central scheduling server. Subsequently, each computing node realizes data interaction processing with the central scheduling server through its dedicated client program.
[0073] Build a unified operating environment for each computing node: Use Docker technology to ensure that all computing nodes run the same version of deep learning frameworks (such as PyTorch, TensorFlow) and dependent libraries, and verify the environment through a unified configuration file. Among them, PyTorch is an open source deep learning framework. TensorFlow is also an open source deep learning framework. Use Docker containers to unify the operating environment of each computing node and solve problems caused by different operating systems and dependency versions. Docker is an open source application container engine that allows developers to package their applications and their dependent packages into a portable container, and then publish it to any popular Linux machine. It can also achieve virtualization. Containers use a complete sandbox mechanism and there will be no interface between them.
[0074] Initialization of the computing power node incentive mechanism: The system architecture assigns a unique identifier to each computing power node, initializes the points or token account, and records the subsequent computing power contribution of each computing power node.
[0075] The incentive mechanism is designed using blockchain or distributed ledger technology to ensure that the contribution records of computing power nodes cannot be tampered with, while sandbox technology is used to prevent malicious code intrusion and system risks.
[0076] In step S102, the overall training task of the distributed large model is preprocessed by the central scheduling server, and the task decomposition algorithm is used to split the overall training task into multiple subtasks, each subtask corresponding to a local data set, a model substructure or a gradient update task.
[0077] In step S103, each computing power node sends its corresponding GPU floating-point computing power, number of CPU cores, available memory, average bandwidth, latency, and packet loss rate to the cloud server so that the cloud server calculates the idle computing power and network status of each computing power node, and further calculates the comprehensive score of each computing power node.
[0078] The central scheduling server uses a dynamic scheduling algorithm to assign multiple subtasks to appropriate computing nodes based on the comprehensive score of each computing node sent by the cloud server, and sets the timeout and retry mechanism for the subtasks.
[0079] In one embodiment, the step of calculating the comprehensive score of each computing power node based on the idle computing power and network status of each computing power node includes:
[0080] The cloud server normalizes the GPU floating-point computing power, CPU core number, and available memory of each computing power node to obtain the normalized GPU floating-point computing power, normalized CPU core number, and normalized available memory of each computing power node.
[0081] The cloud server calculates the idle computing power score of each computing power node based on the normalized GPU floating-point computing power, normalized number of CPU cores, normalized available memory, and their corresponding weights. The idle computing power score of each computing power node is used to quantify the idle computing power of each computing power node.
[0082] The cloud server normalizes the average bandwidth, latency, and packet loss rate of each computing power node to obtain the normalized average bandwidth, normalized latency, and normalized packet loss rate of each computing power node.
[0083] The cloud server calculates the network status score of each computing power node based on the normalized average bandwidth, normalized latency, and normalized packet loss rate of each computing power node, as well as their corresponding weights. The network status score of each computing power node quantifies the network status of each computing power node.
[0084] The cloud server calculates the comprehensive score of each computing power node based on the idle computing power score and network status score of each computing power node, as well as their respective weights.
[0085] Assuming the idle computing power score of computing power node i is Ci, a comprehensive evaluation can be performed based on information such as GPU floating-point computing power, number of CPU cores, and available memory. Normalization is performed so that the idle computing power scores of all computing power nodes are within the range of [0, 1].
[0086] Assume that the GPU floating-point computing power of computing power node i is F_i, the number of CPU cores is C_i, and the available memory is Mi_i. At the same time, the maximum GPU floating-point computing power F_max, the minimum GPU floating-point computing power F_min, the maximum number of CPU cores C_max, the minimum number of CPU cores C_min, the maximum available memory M_min, and the minimum available memory M_max of all computing power nodes are pre-calculated.
[0087] Normalize the GPU floating-point computing power, number of CPU cores, and available memory of each computing power node to obtain the normalized GPU floating-point computing power, normalized number of CPU cores, and normalized available memory of each computing power node. The calculation formulas for the normalized GPU floating-point computing power S_gpu, the normalized number of CPU cores S_cpu, and the normalized available memory S_mem of each computing power node are as follows:
[0088] S_gpu=(F_i–F_min) / (F_max–F_min)(1)
[0089] S_cpu=(C_i–C_min) / (C_max–C_min)(2)
[0090] S_mem=(M_i–M_min) / (M_max–M_min)(3)
[0091] The normalized GPU floating-point computing power S_gpu of each computing power node can be calculated by formula (1), the normalized number of CPU cores S_cpu of each computing power node can be calculated by formula (2), and the normalized available memory S_mem of each computing power node can be calculated by formula (3).
[0092] Set the weight w_gpu corresponding to the normalized GPU floating-point computing capability, the weight w_cpu corresponding to the normalized number of CPU cores, and the weight w_mem corresponding to the normalized available memory size, w_gpu + w_cpu + w_mem = 1.
[0093] The idle computing power score of each computing power node is calculated based on the normalized GPU floating point computing power, normalized CPU core number and normalized available memory of each computing power node, as well as their corresponding weights. The calculation formula for the idle computing power score Ci of each computing power node is as follows:
[0094] Ci=w_gpu·S_gpu+w_cpu·S_cpu+w_mem·S_mem(4).
[0095] The idle computing power score Ci of each computing power node can be calculated using formula (4).
[0096] Assuming that the network status score of computing power node i is Ni, we can comprehensively refer to the average bandwidth, latency, and packet loss rate of computing power node i. Similarly, normalization is performed to make the score within the range of [0,1].
[0097] Assume that the average bandwidth of computing power node i is B_i, the delay is L_i, and the packet loss rate is P_i. Count the average maximum bandwidth B_max, the average minimum bandwidth B_min, the maximum delay L_max, the minimum delay L_min, the maximum packet loss rate P_max, and the minimum packet loss rate P_min of all computing power nodes.
[0098] The average bandwidth, latency, and packet loss rate of each computing power node are normalized to obtain the normalized average bandwidth, normalized latency, and normalized packet loss rate of each computing power node. The calculation formulas for the normalized average bandwidth S_bandwidth, the normalized latency S_latency, and the normalized packet loss rate S_loss of each computing power node are as follows:
[0099] S_bandwidth=(B_i–B_min) / (B_max–B_min)(5)
[0100] S_latency=1–(L_i–L_min) / (L_max–L_min)(6)
[0101] S_loss=1–(P_i–P_min) / (P_max–P_min)(7)
[0102] The normalized average bandwidth of each computing power node can be calculated by formula (5), the normalized delay of each computing power node can be calculated by formula (6), and the normalized packet loss rate of each computing power node can be calculated by formula (7).
[0103] Set the weight w_bandwidth corresponding to the normalized average bandwidth, the weight w_latency corresponding to the normalized delay, and the weight w_loss corresponding to the normalized packet loss rate, w_bandwidth + w_latency + w_loss = 1.
[0104] The network status score of each computing power node is calculated based on the normalized average bandwidth, normalized latency, and normalized packet loss rate of each computing power node, as well as their corresponding weights. The calculation formula for the network status score of each computing power node is as follows:
[0105] N_i=w_bandwidth·S_bandwidth+w_latency·S_latency+w_loss·S_loss (8).
[0106] The network status score of each computing power node can be calculated using formula (8).
[0107] Furthermore, the comprehensive score of each computing node is calculated based on the idle computing power score and network status score of each computing node, as well as their corresponding weights. The calculation formula for the comprehensive score Si of each computing node i is as follows:
[0108] Si=α×Ci+β×Ni(9)
[0109] In the formula, α represents the weight corresponding to the idle computing power score, β represents the weight corresponding to the network status score, and α+β=1.
[0110] If the idle computing power of the computing nodes may be required to be higher when training a large model, the weight corresponding to the idle computing power score can be set to a larger value, for example, α = 0.7, β = 0.3.
[0111] After calculating the comprehensive score of the computing power node, the central scheduling server will assign the subtask to the computing power node or computing power node group with the highest comprehensive score.
[0112] It should be noted that as training progresses, computing nodes may experience overheating, frequency reduction, or network congestion, requiring periodic or event-driven updates to the node's comprehensive score. The central scheduling server assigns tasks based on the latest comprehensive score before assigning each subtask.
[0113] In one embodiment, the central scheduling server dynamically assigns multiple subtasks to their respective corresponding computing nodes based on the comprehensive score of each computing node sent by the cloud server, and sets the timeout period and retry mechanism of the subtasks, further comprising:
[0114] The central scheduling server sorts all assigned subtasks by priority based on the priority of each computing node sent by the cloud server. The priority of each subtask is calculated based on the data volume and model block of each computing node.
[0115] The central dispatch server assigns all subtasks to their respective appropriate computing nodes according to the allocation strategy; the allocation strategy is: high-priority subtasks are assigned to computing nodes with high comprehensive scores, medium-priority subtasks are assigned to computing nodes with medium comprehensive scores, and low-priority subtasks are assigned to computing nodes with low comprehensive scores.
[0116] It can be understood that by sorting all subtasks to be assigned according to priority, and assigning all subtasks to be assigned to their respective appropriate computing nodes according to the allocation strategy, the central scheduling server can assign high-priority subtasks to computing nodes with high comprehensive scores, medium-priority subtasks to computing nodes with medium comprehensive scores, and low-priority subtasks to computing nodes with low comprehensive scores, thereby achieving efficient allocation and utilization of computing node resources.
[0117] Furthermore, the step of calculating the priority of each subtask based on the data volume and model block of each subtask includes:
[0118] The cloud server normalizes the data volume and model block of each subtask to obtain the normalized data volume and normalized model block of each subtask;
[0119] The cloud server calculates the priority score of each subtask based on the normalized data volume, normalized model block and corresponding weight of each subtask; wherein, the priority of each subtask is quantified by the priority score of each subtask.
[0120] It can be understood that the calculation formulas for the normalized data volume of each subtask j and the calculation formulas for the normalized model block of each subtask j are as follows:
[0121] S_data=(D_j–D_min) / (D_max–D_min)(10)
[0122] S_model=(M_j–M_min) / (M_max–M_min)(11)
[0123] Where j represents the subtask, S_data represents the normalized data size of each subtask, and S_model represents the normalized model block of each subtask.
[0124] The normalized data volume of each subtask can be calculated by formula (10), and the normalized model block of each subtask can be calculated by formula (11)
[0125] The calculation formula for the priority score P_j of each subtask is as follows:
[0126] P_j=w_data×S_data+w_model×S_model(12)
[0127] Where w_data represents the weight corresponding to the normalized data volume, w_model represents the weight corresponding to the normalized model block, and w_data + w_model = 1.
[0128] The priority score of each subtask can be calculated by formula (12). The higher the priority score of each subtask, the larger the data volume or model block of the corresponding subtask, and the higher the score, the higher the priority score will be assigned to the computing power node.
[0129] In one embodiment, the step of setting a timeout period and a retry mechanism for a subtask includes:
[0130] The timeout for each subtask is determined by the estimated execution time of each subtask and the safety factor;
[0131] The retry mechanism is that the computing power node recalculates the subtasks it receives within the preset number of retries.
[0132] Furthermore, if a computing node fails to upload its computation results within the timeout period, the central scheduling server will reallocate the subtask to other computing nodes;
[0133] If a subtask is not completed by the corresponding computing node within the timeout period, the retry mechanism will be triggered.
[0134] It is understandable that if a computing power node fails to upload the results on time due to a failure or network problem, that is, a computing power node has exceeded the timeout period and has not yet uploaded its calculation results, the central scheduling server will assign the subtasks it has received to other computing power nodes to ensure that the overall training process is not interrupted.
[0135] The preset number of retries is three. If a subtask is not completed by the corresponding computing node within the timeout period, a retry mechanism is triggered, allowing the computing node to continue computing the subtask three more times. If the computing node still has not completed the corresponding subtask within three times, the central scheduling server will assign the subtask to another computing node. The preset number of retries can be set based on actual conditions and is not restricted by this application.
[0136] In one embodiment, the step of the central scheduling server reallocating subtasks to other computing nodes includes:
[0137] The central scheduling server selects a backup computing node to calculate the subtask based on the calculation progress of the completed subtask and the comprehensive score of the current computing power node, so that the backup computing power node can continue to execute the same subtask; the calculation progress of the subtask is that each computing power node uploads the progress of the corresponding subtask to the central scheduling server within a preset time.
[0138] It is understandable that the central dispatch server may set a timeout of 10 minutes for uploading computation results. If computing node A still has not uploaded the computation results after 10 minutes, the central dispatch server will deem that computing node A has failed. The timeout period can be set based on actual conditions, and this application will not elaborate on this.
[0139] Each computing node will report the computational progress of the corresponding subtask to the central scheduling server within a preset time, so that the central scheduling server can select a backup computing node (such as computing node B) to take over the task based on the computational progress of the completed subtask, the idle computing power of the current computing node, and the network status. That is, the backup computing node B continues to execute the same subtask as computing node A to ensure that training is not affected.
[0140] Once computing node B completes the subtask originally executed by computing node A, computing node B will upload the new calculation results to the central scheduling server.
[0141] In step S104 and step S105, after the subtask is assigned to the corresponding computing node, the central scheduling server will distribute the initialization parameters of the distributed large model to the computing node that receives the subtask, so that the computing node that receives the subtask can train the subtask, and perform calculations on the calculation results of the computing node to obtain the calculation results, and send the calculation results to the central scheduling server.
[0142] It is understandable that before training the subtasks, the central scheduling server will distribute the initialization parameters of the distributed large model to the computing nodes that receive the subtasks, so that the computing nodes that receive the subtasks can train the subtasks.
[0143] The computing power node that receives the subtask uses its computing power to train the received subtask through the training data of the data storage center, and calculates the calculation results of the computing power node to obtain the calculation results, and then sends the calculation results of all computing power nodes to the central scheduling server, so that the central scheduling server can aggregate the calculation results of all computing power nodes and update the global model parameters.
[0144] It is understood that the calculation results include key model parameters and intermediate results. The computing power node that receives the subtask uses local computing power to train the subtask based on the training data, calculates the key model parameters and intermediate results of the computing power node, and then sends the calculated model parameters and intermediate results to the central scheduling server. The central scheduling server aggregates the key model parameters and intermediate results and updates the global model parameters.
[0145] Furthermore, the calculation results include key model parameters and intermediate results;
[0146] The steps for updating global model parameters include:
[0147] Key model parameters are updated using synchronous updating;
[0148] Some intermediate results are updated asynchronously.
[0149] It is understandable that the synchronous update method is used for key model parameters to ensure the consistency of the calculation results of each computing power node. For some intermediate results, asynchronous updates are used to reduce the congestion caused by the performance differences of computing power nodes and network delays.
[0150] In distributed neural network training, key model parameters typically refer to the core weights and biases that determine the model's ultimate predictive performance, such as parameters of convolutional layers, fully connected layers, and normalization layers. These parameters must be strictly synchronized across all computing nodes. Intermediate results, including activation values, local feature maps, and local gradient statistics generated during the forward propagation process, can be processed asynchronously, reducing overall congestion caused by performance differences between computing nodes and network latency. Key model parameters also include gradients.
[0151] In distributed training, synchronous updating means that after all computing nodes complete computation in the same iteration or batch, they aggregate their gradients or parameters to obtain new global parameters before continuing with the next round of training. Each computing node uses the backpropagation algorithm to calculate the gradient (i.e., partial derivative) of the loss function with respect to the local model parameters. Then, using methods such as AllReduce or Parameter Server, the local gradients from all computing nodes are aggregated (typically averaged) to obtain a global average gradient. Next, the gradient descent update formula is applied: θ(t+1) = θ(t) – η × (average gradient), where θ(t) represents the gradient at time t, θ(t+1) represents the gradient at time t+1, and η represents the learning rate. The global average gradient is applied to the current model parameters, completing the global parameter update. A11Reduce is a common communication operation in parallel and distributed computing, used to aggregate data between multiple computing nodes (such as GPUs or CPUs). It is commonly used in distributed training in deep learning to integrate gradient information from various computing nodes for parameter updates. Parameter Server is a programming framework that facilitates the writing of distributed parallel programs, with a focus on supporting the distributed storage and coordination of large-scale parameters. The word "iteration" refers to iteration, and the word "batch" refers to batches.
[0152] The following are two common synchronization update methods:
[0153] (1) Synchronous update
[0154] Computing nodes include central computing nodes and worker computing nodes. During synchronous updates, central computing nodes store and maintain global model parameters. Worker computing nodes calculate local gradients and send them to the central scheduling server (PS).
[0155] Training Process
[0156] Initialization: PS distributes the initial model parameters to all working computing nodes.
[0157] Local calculation: Each computing node uses local data to calculate gradients.
[0158] Gradient aggregation: Each computing power node sends the calculated gradient to the PS, which accumulates or averages all the calculated gradients.
[0159] Update parameters: PS updates the global model parameters according to the aggregated gradients.
[0160] Broadcast updated parameters: PS sends the new global model parameters to all computing nodes and starts the next round of calculation.
[0161] Synchronization mechanism
[0162] In strict synchronization mode, PS will only update and broadcast parameters when all computing power nodes have uploaded gradients; if a computing power node is delayed or fails, the entire process will be blocked until it times out or is rescheduled.
[0163] Advantages: The model converges consistently and can theoretically achieve the same accuracy as single-machine training.
[0164] Disadvantages: If there are a large number of computing nodes, the delay of any one computing node may cause the overall training speed to decrease. However, this disadvantage is avoided in the subsequent fault tolerance process.
[0165] (2) Asynchronous Update
[0166] During asynchronous updates, the central computing node stores the global model parameters and updates them immediately upon receiving the gradients from any computing node. Worker computing nodes independently calculate the gradients and then asynchronously upload them to the PS.
[0167] Training Process
[0168] Initialization: PS distributes the initial model parameters to all working computing nodes.
[0169] Asynchronous computing: Each computing node performs forward and backward propagation based on the current (possibly asynchronous) model parameters to obtain local gradients.
[0170] Asynchronous upload of gradients: The computing power node sends the gradients to the PS without waiting for other computing power nodes.
[0171] PS updates parameters in real time: PS updates global model parameters immediately after receiving gradients.
[0172] Computing nodes obtain the latest parameters: Computing nodes can pull the latest parameters from the PS before the next iteration starts; updates can also be pushed by the PS (both push and pull mechanisms are possible).
[0173] Example of asynchronous update of intermediate results
[0174] Model segmentation: Divide the model into different layers or different parameter blocks. Each computing node may be responsible for the calculation of part of the layers. After completing the calculation for which it is responsible, the computing node can pass the intermediate gradient or parameter increment back to the PS.
[0175] Distributed cache: PS uses distributed cache or message queues to store intermediate results. Computing nodes can read the latest intermediate parameters at any time, or pull the latest global parameters again before the next training step.
[0176] Delay tolerance: In asynchronous mode, the model parameters used by different computing nodes may not be completely consistent (because they obtain the latest parameters at different times), which may, to a certain extent, lead to slower training convergence or a slight decrease in final accuracy. However, in scenarios with large heterogeneity in the network and computing nodes, asynchronous updates can significantly improve overall throughput and resource utilization.
[0177] Through asynchronous updates, the entire training process will not be blocked by delays or failures of individual computing nodes, and higher resource utilization can be achieved in environments with large differences in computing node performance.
[0178] In step S106, the central scheduling server feeds back the aggregated global model parameters to each computing node, so that each computing node performs training based on the latest global model parameters in the next round of calculation until the distributed large model converges and the training ends.
[0179] It is understandable that the central scheduling server sends the updated global model parameters to all computing nodes and starts the next round of calculation until the distributed large model converges and the training ends.
[0180] In one embodiment, the step of sending the computation results of all computing nodes to the central scheduling server includes:
[0181] The calculation results of all computing nodes are encrypted and fragmented and transmitted to the central scheduling server through the cloud server.
[0182] It's understood that after receiving the corresponding subtask, the computing node uses its local GPU / CPU to calculate the subtask and run the training program in a container environment. After the computing node completes the calculation, it performs operations on the results to obtain the operation results, and uploads the local calculation results of key model parameters (such as gradients and local model parameters) to the central scheduling server. To reduce communication overhead, the system uses data compression, fragmented transmission, and encrypted transmission.
[0183] In this application, different encryption strategies are selected based on the data type, transmission security requirements and performance requirements.
[0184] For model computing power node parameters: Use AES algorithm for symmetric encryption. Because most of the encrypted data transmission volume is large, AES algorithm has good encryption performance.
[0185] For the reward points key: use RSA asymmetric encryption to ensure security.
[0186] For the use of customer personal information: Encrypted using the AES-256 encryption algorithm.
[0187] When sharding data, it is necessary to perform data sharding. When generating data (such as parameters, gradients, or training sets of training models), the data is divided into multiple smaller parts according to certain rules. For example, the matrix of model parameters can be divided into multiple blocks, and each block size is a predetermined value. The specific size can be set according to the actual situation, and this application does not elaborate on this.
[0188] These data shards can use fixed-size buffers or be dynamically split based on data content. During shard transmission, each data shard is transmitted over the network to the corresponding computing power node. To ensure the order of transmission, a sequence or identifier can be attached to each shard.
[0189] After the transmission is complete, the receiver reassembles all the fragments to restore the complete data. If a fragment is lost or damaged during transmission, the receiver can request retransmission of that fragment. Checksums (such as CRC) can be used to verify data integrity and ensure the correctness of each fragment.
[0190] It should be noted that you can also use multiple threads or multiple connections to transfer multiple shards simultaneously to further improve the transfer progress. For example, different shards can be transmitted through different network paths to avoid the bottleneck of a single path.
[0191] It's important to note that efficient communication protocols based on TCP / IP or gRPC, combined with SSL (Secure Socket Layer) / TLS (Transport Layer Security) encryption technology, ensure low latency and secure data transmission between computing nodes. TCP / IP stands for Transmission Control Protocol / Internet Protocol. gRPC is a high-performance, general-purpose remote procedure call framework open sourced by Google.
[0192] It should be noted that although the steps of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in this specific order, or that all steps must be performed to achieve the desired results. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into a single step, and / or a single step may be decomposed into multiple steps. In addition, it is also easy to understand that these steps may be executed synchronously or asynchronously, for example, in multiple modules / processes / threads.
[0193] Furthermore, in this example implementation, a distributed large model training system based on idle computing power of virtual machines is also provided. Figure 3 As shown in , the system 200 may include a startup registration module 210, a splitting module 220, an allocation module 230, a distribution module 240, an aggregation update module 250 and a feedback module 260. Among them: the startup registration module 210 is used for the central scheduling server to register each computing power node based on the registration request for registering the computing power node sent by the cloud server after the cloud server starts each computing power node, and to build a unified operating environment for each computing power node; the splitting module 220 is used for the central scheduling server to split the overall training task of the distributed large model into multiple subtasks; the allocation module 230 is used for the central scheduling server to dynamically allocate multiple subtasks to their respective corresponding computing power nodes based on the comprehensive score of each computing power node sent by the cloud server, and set the timeout time and retry mechanism of the subtasks; wherein the comprehensive score of each computing power node is calculated by the idle computing power and network status of each computing power node; the distribution module Block 240 is used for the central scheduling server to distribute the initialization parameters of the distributed large model to the computing power nodes that receive the subtasks, so that the computing power nodes that receive the subtasks can train the received subtasks through the training data, and perform calculations on the calculation results of the computing power nodes to obtain calculation results; the aggregation update module 250 is used for the central scheduling server to receive the calculation results sent by the cloud server, aggregate the calculation results of all computing power nodes, and update the global model parameters; the feedback module 260 is used for the central scheduling server to feed back the aggregated global model parameters to each computing power node, so that each computing power node can be trained based on the latest global model parameters in the next round of calculations until the distributed large model converges and the training is completed.
[0194] Regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0195] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the implementation mode of the present application, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of a module or unit described above can be further divided into multiple modules or units for concretization. The components displayed as modules or units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present application scheme. Those of ordinary skill in the art can understand and implement it without paying any creative work.
[0196] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.
Claims
1. A distributed large model training method based on the idle computing power of virtual machines, characterized in that: The method is applied to a distributed large-scale model training system, which includes a central scheduling server, one or more cloud servers, and a data storage center. The central scheduling server is connected to the one or more cloud servers and the data storage center, respectively. Each cloud server includes multiple computing nodes, each computing node being a virtual machine. The data storage center is used to store training data and calculation results. The method includes: After the cloud server starts each computing node, the central scheduling server registers each computing node based on the registration request for registering the computing node sent by the cloud server, and builds a unified operating environment for each computing node; The central scheduling server splits the overall training task of the distributed large model into multiple subtasks; The central scheduling server dynamically assigns the plurality of subtasks to the corresponding computing nodes based on the comprehensive score of each computing node sent by the cloud server, and sets a timeout period and a retry mechanism for the subtasks; wherein the comprehensive score of each computing node is calculated based on the idle computing power and network status of each computing node; The central scheduling server distributes the initialization parameters of the distributed large model to the computing nodes that receive the subtasks, so that the computing nodes that receive the subtasks train the received subtasks through the training data, and calculates the calculation results of the computing nodes to obtain calculation results, and sends the calculation results to the central scheduling server; The central scheduling server receives the calculation results sent by the cloud server, aggregates the calculation results of all the computing power nodes, and updates the global model parameters; The central scheduling server feeds back the aggregated global model parameters to each of the computing nodes, so that each of the computing nodes performs training based on the latest global model parameters in the next round of calculation until the distributed large model converges and the training ends.
2. The distributed large model training method based on idle computing power of virtual machines according to claim 1 is characterized in that: The step of calculating the comprehensive score of each computing power node based on the idle computing power and network status of each computing power node includes: Normalizing the GPU floating-point computing capability, the number of CPU cores, and the available memory of each computing power node by the cloud server to obtain the normalized GPU floating-point computing capability, the normalized number of CPU cores, and the normalized available memory of each computing power node; The cloud server calculates the idle computing power score of each computing power node according to the normalized GPU floating-point computing power, the normalized number of CPU cores, and the normalized available memory of each computing power node, as well as their corresponding weights; wherein the idle computing power score of each computing power node is used to quantify the idle computing power of each computing power node; Normalizing the average bandwidth, delay, and packet loss rate of each computing power node by the cloud server to obtain the normalized average bandwidth, normalized delay, and normalized packet loss rate of each computing power node; Calculating, by the cloud server, a network status score for each computing power node based on the normalized average bandwidth, normalized latency, and normalized packet loss rate of each computing power node, as well as their corresponding weights; wherein the network status of each computing power node is quantified by the network status score of each computing power node; The cloud server calculates the comprehensive score of each computing power node based on the idle computing power score and network status score of each computing power node, as well as their respective corresponding weights, and sends the comprehensive score of each computing power node to the central scheduling server.
3. The distributed large model training method based on idle computing power of virtual machines according to claim 2 is characterized in that: The step of the central scheduling server dynamically allocating the plurality of subtasks to the respective corresponding computing nodes based on the comprehensive score of each computing node sent by the cloud server, and setting the timeout time and retry mechanism of the subtasks further includes: The central scheduling server sorts all the subtasks to be assigned based on the priority of each computing node sent by the cloud server; wherein the priority of each subtask is calculated according to the data volume and model block of each computing node; The central scheduling server allocates all the subtasks to be allocated to their respective appropriate computing nodes according to the allocation strategy; wherein the allocation strategy is: the subtasks with high priority are allocated to the computing nodes with high comprehensive scores, the subtasks with medium priority are allocated to the computing nodes with medium comprehensive scores, and the subtasks with low priority are allocated to the computing nodes with low comprehensive scores.
4. The distributed large model training method based on idle computing power of virtual machines according to claim 3 is characterized in that: The step of calculating the priority of each subtask according to the data volume and model block of each subtask includes: Normalizing the data volume and model block of each subtask by the cloud server to obtain the normalized data volume and normalized model block of each subtask; The cloud server calculates the priority score of each subtask based on the normalized data volume, normalized model block and corresponding weight of each subtask; wherein the priority of each subtask is quantified by the priority score of each subtask.
5. The distributed large model training method based on idle computing power of virtual machines according to claim 1 is characterized in that: The step of setting the timeout period and retry mechanism of the subtask includes: The timeout period of each subtask is determined by the estimated execution time of each subtask and a safety factor; The retry mechanism is that the computing power node recalculates the subtask it has received within a preset number of retries.
6. The distributed large model training method based on idle computing power of virtual machines according to claim 5 is characterized in that: If a computing node fails to upload its calculation result within the timeout period, the central scheduling server reallocates the subtask to another computing node; If a subtask is not completed by the corresponding computing power node within the timeout period, a retry mechanism is triggered.
7. The distributed large model training method based on idle computing power of virtual machines according to claim 6 is characterized in that: The step of the central scheduling server reallocating the subtask to other computing nodes includes: The central scheduling server selects a backup computing node to calculate the subtask based on the calculation progress of the completed subtask and the comprehensive score of the current computing power node, so that the backup computing power node continues to execute the same subtask; wherein, the calculation progress of the subtask is that each computing power node uploads the progress of the corresponding subtask to the central scheduling server within a preset time.
8. The distributed large model training method based on idle computing power of virtual machines according to claim 1 is characterized in that: The calculation results include key model parameters and intermediate results; The step of updating the global model parameters includes: The key model parameters are updated in a synchronous updating manner; Some of the intermediate results are updated in an asynchronous manner.
9. The distributed large model training method based on idle computing power of virtual machines according to claim 1 is characterized in that: The step of sending the calculation result of the computing power node to the central scheduling server includes: The calculation results of all the computing nodes are encrypted and transmitted in segments to the central scheduling server through the cloud server.
10. A distributed large model training system based on idle computing power of virtual machines, characterized in that: The system comprises: A startup registration module is configured to register each computing node based on a registration request sent by the cloud server to the central scheduling server after the cloud server starts each computing node, and to establish a unified operating environment for each computing node; A splitting module, used for the central scheduling server to split the overall training task of the distributed large model into multiple subtasks; an allocation module, configured for the central scheduling server to dynamically allocate the plurality of subtasks to the respective corresponding computing nodes based on the comprehensive score of each computing node sent by the cloud server, and to set a timeout period and a retry mechanism for the subtasks; wherein the comprehensive score of each computing node is calculated based on the idle computing power and network status of each computing node; a distribution module, configured for the central scheduling server to distribute the initialization parameters of the distributed large model to the computing nodes that receive the subtasks, so that the computing nodes that receive the subtasks perform training on the received subtasks using the training data, and to calculate the calculation results of the computing nodes to obtain calculation results; An aggregation and updating module is used for the central scheduling server to receive the calculation results sent by the cloud server, aggregate the calculation results of all the computing power nodes, and update the global model parameters; A feedback module is used for the central scheduling server to feed back the aggregated global model parameters to each of the computing power nodes, so that each of the computing power nodes can be trained based on the latest global model parameters in the next round of calculation until the distributed large model converges and the training is completed.
Citation Information
Cited By
Heterogeneous environment prediction model training and deployment method and system
CN121957827A