Distributed scheduling training method for AI model training

Through the distributed scheduling training method, the resource allocation and number of nodes of AI model training are monitored and dynamically adjusted in real time, which solves the problems of idle resources, load fluctuations and failure interruptions in AI model training, and achieves efficient resource utilization and training efficiency improvement.

CN120216205AInactive Publication Date: 2025-06-27HANGZHOU SHENGHENG TECH CO LTD

Patent Information

Application Number
CN202510657624.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the training of AI model, the existing technology has problems such as idle resources, inability to respond to training node load fluctuations in real time, training interruptions caused by node failure, and scheduling difficulties caused by inconsistent performance of different types of GPUs.

Method used

Provide a distributed scheduling training method, including API gateway, resource monitoring module, dynamic scheduling module, node training module, and result saving and fault tolerance processing module. Through the dynamic scheduling module, the node load and resource status are monitored in real time, and the model sharding strategy and number of nodes are dynamically adjusted to achieve efficient resource utilization and automatic recovery of failures.

Benefits of technology

It realizes dynamic resource awareness and allocation of model training, supports dynamic node addition and removal, and automatically schedules to solve node failures. It can effectively utilize a variety of GPU resources in heterogeneous environments, significantly improving training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216205A_ABST
    Figure CN120216205A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence training, and belongs to a distributed scheduling training method for AI model training, which comprises an API gateway, a resource monitoring module, a dynamic scheduling module, a node training module and a result storage and fault-tolerant processing module. According to the method, dynamic resource perception and dynamic allocation of training nodes of model training can be realized, model fragmentation can dynamically adjust a fragmentation strategy according to a node load, an aggregation algorithm can reduce communication overhead, especially processing in a heterogeneous environment, and more efficient training is realized and addition and removal of dynamic nodes are supported through combination of the two. And moreover, the fragment number deviation of node processing caused by a poor network can be corrected, updating inconsistency caused by model fragmentation is inhibited, and other nodes can be automatically scheduled to continue training due to training failure caused by seamless connection node faults supporting elastic capacity expansion and contraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence training, and belongs to a distributed scheduling training method for AI model training. Background Art

[0002] The traditional way of model training usually uses static resource allocation, but there is a problem of resource idleness in this allocation method. Although container orchestration systems such as Kubernetes have a scheduling system, they cannot respond to the load fluctuations of training nodes in real time, and there are great limitations. In addition, if a node fails and causes the training to be interrupted, it generally needs to be manually processed to resume the training. Moreover, there are now many different models of GPUs, and the performance of each GPU is different. How to uniformly schedule and train multiple different types of GPUs is also an urgent problem to be solved. Summary of the Invention

[0003] In view of the above technical problems, the present invention provides a distributed scheduling training method for AI model training.

[0004] To achieve the above object, the present invention provides the following technical solutions: Provide a distributed scheduling training method for AI model training, including an API gateway, a resource monitoring module, a dynamic scheduling module, a node training module, and a result saving and fault tolerance processing module; The API gateway receives a client training request and transfers the information to the dynamic scheduling module; The dynamic scheduling module performs model sharding, hands the model shards to the corresponding training nodes, and allocates training tasks to the node training module; The node training module executes specific training tasks and returns the number of model shards processed per unit time to the dynamic scheduling module. The dynamic scheduling module dynamically adjusts the sharding strategy according to the returned data; The resource monitoring module real-time collects the hardware metrics and network status of computing nodes, reports the node load to the dynamic scheduling module, and the dynamic scheduling module adjusts the number of nodes according to the reported data; After the training is completed, each shard is aggregated into a final aggregated model that has completed training, and the final aggregated model is transferred to the result saving and fault tolerance processing module. The result saving and fault tolerance processing module saves the training result and senses whether the status of the training node is abnormal. If there is an abnormality, the training task is migrated to other nodes to continue training.

[0005] Preferably, the dynamic scheduling module includes an expansion unit and a load prediction unit; the expansion unit increases or decreases the number of physical nodes according to the relationship between the average node load reported by the resource monitoring module and the threshold value. When the average node load is greater than the threshold value, capacity expansion is triggered; when the average node load is less than the threshold value, capacity reduction is triggered; the load prediction unit predicts future resource requirements based on historical data, and when there is a large deviation between the predicted value and the current actual node resources, it triggers the increase or decrease of the number of physical nodes.

[0006] Preferably, when the dynamic scheduling module allocates model shards to the node training module, it selects a number of training nodes from the currently available nodes, and allocates the model shards to the selected nodes according to the node computing capabilities of each training node.

[0007] Preferably, the dynamic scheduling module introduces a gradient residual compensation algorithm for error correction and convergence acceleration.

[0008] Preferably, the model shards are optimized in batches proportionally by the Lagrange multiplier method.

[0009] Preferably, the aggregated model is completed by gradient aggregation.

[0010] Preferably, the result saving and fault tolerance processing module adopts an adaptive gradient synchronization algorithm.

[0011] Compared with the prior art, the present invention provides a distributed scheduling training method for AI model training, having the following beneficial effects: 1. It can realize dynamic resource perception of model training, dynamically allocate training nodes, the model shards can dynamically adjust the sharding strategy according to the node load, the aggregation algorithm can reduce the communication overhead, especially for processing in a heterogeneous environment. The combination of the two achieves more efficient training and supports the addition and removal of dynamic nodes.

[0012] 2. Training failures caused by node failures can automatically schedule other nodes to continue training. Through the quantization index-driven strategy for dynamic adjustment, the precise matching of communication efficiency and computing resources is realized, and the training efficiency in a heterogeneous environment is significantly improved compared with the prior art. The core innovation lies in forming a closed-loop optimization system of real-time performance perception, policy decision-making model and gradient compensation mechanism, breaking through the limitations of traditional fixed strategies.

[0013] 3. It can simultaneously use different types of GPUs for unified scheduling training. The model is first sharded and then trained. Different types of GPUs, according to the sharding strategy, each undertake different numbers of shards and jointly train the same model.

[0014] 4. It can correct the deviation of the number of shards processed by nodes caused by a poor network, suppress the update inconsistency caused by model sharding, and support the seamless connection of elastic scaling.

[0015] The features and advantages of the present invention will be described in detail through embodiments in conjunction with the accompanying drawings. Description of the Drawings

[0016] Figure 1 It is an architecture diagram of a distributed AI training scheduling system for a distributed scheduling training method for AI model training according to the present invention. Detailed Embodiments

[0017] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below through the accompanying drawings and embodiments. However, it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessarily confusing the concepts of the present invention.

[0018] Embodiment 1: Description of relevant modules included in the implementation solution of the present invention: (1) API gateway: responsible for receiving and processing requests; (2) Resource monitoring module: real-time collection of hardware metrics and network status of computing nodes; (3) Dynamic scheduling module: automatically adjusts the composition of the computing node cluster according to preset policies; (4) Node training module: executes specific training tasks; (5) Result saving and fault tolerance processing module: realizes hot migration of training status and resuming training from breakpoint, and adopts a differential data compression and transmission protocol.

[0019] Refer to Figure 1 , the present invention provides a distributed scheduling training method for AI model training, including an API gateway, a resource monitoring module, a dynamic scheduling module, a node training module, and a result saving and fault tolerance processing module; the API gateway receives a client training request and transfers the information to the dynamic scheduling module; the dynamic scheduling module performs model sharding, hands the model shards to the corresponding training nodes, and assigns training tasks to the node training module; the node training module executes specific training tasks and returns the number of model shards processed per unit time to the dynamic scheduling module, and the dynamic scheduling module dynamically adjusts the sharding strategy according to the returned data; the resource monitoring module real-time collects the hardware metrics and network status of computing nodes, reports the node load to the dynamic scheduling module, and the dynamic scheduling module adjusts the number of nodes according to the reported data; after training is completed, each shard is aggregated into a final aggregated model that has completed training, and the final aggregated model is transferred to the result saving and fault tolerance processing module, and the result saving and fault tolerance processing module saves the training result and senses whether the status of the training node is abnormal. If there is an abnormality, the training task is migrated to other nodes to continue training.

[0020] Specifically, the dynamic scheduling module includes an expansion unit and a load prediction unit; the expansion unit increases or decreases the number of physical nodes according to the magnitude relationship between the average node load reported by the resource monitoring module and the threshold. When the average node load is greater than the threshold, expansion is triggered; when the average node load is less than the threshold, contraction is triggered; the load prediction unit predicts future resource requirements based on historical data, and when there is a large deviation between the predicted value and the current actual node resources, it triggers an increase or decrease in the number of physical nodes.

[0021] Specifically, when the dynamic scheduling module allocates model shards to the node training module, it selects several training nodes from the currently available nodes and distributes the several model shards to the selected nodes according to the computing power of each training node.

[0022] Specifically, the dynamic scheduling module introduces a gradient residual compensation algorithm for error correction and convergence acceleration.

[0023] Specifically, the model shards are optimized in batches proportionally by the Lagrange multiplier method.

[0024] Specifically, the aggregated model is completed by gradient aggregation.

[0025] Specifically, the result saving and fault tolerance processing module adopts an adaptive gradient synchronization algorithm.

[0026] Embodiment 2: Resource monitoring module: Real-time collection of node performance parameters (CPU / GPU utilization rate, memory occupancy, network bandwidth, storage IOPS); Dynamic scheduling module: Includes an expansion unit and a load prediction unit, Expansion unit, increases or decreases the number of physical nodes. For example, when the average node load > threshold T1, expansion is triggered. For GPU nodes, set T1 = 75%, or T1 = 85%. For edge nodes, set T1 = 60%; Load prediction unit, predicts future resource requirements based on historical data. If there is a large deviation between the predicted value and the current actual nodes, it triggers an increase or decrease in the number of physical nodes, such as the resource usage after 5 minutes, or 10 minutes; The dynamic scheduling also introduces a gradient residual compensation algorithm, which plays a dual role of error correction and convergence acceleration; Node training module: Virtualized resource pool: Provides fine-grained resource allocation such as GPU slices and video memory blocks; Result saving and fault tolerance processing module: Saves the training results. When it senses that the status of a training node is abnormal, it migrates the training task to other nodes to continue training.

[0027] Embodiment 3: The dynamic scheduling module is based on the formula: ∝ L k The number of model shards allocated to node k, c k The computing power of the node, S param The video memory occupancy of each model shard, it is concluded that the number of model shards that a node can allocate is positively correlated with the computing power of the node and negatively correlated with the video memory occupancy of each model shard. When S param is the same (for example, 2 Gb per shard), the number of shards of k for each node is determined by the computing power of the node.

[0028] Node computing power: After the model training scheduling service of the present invention is started, each training node registers with the resource monitoring center, and the resource monitoring center collects the performance parameters (CPU / GPU utilization rate, memory occupancy, network bandwidth, storage IOPS) of each training node.

[0029] When the scheduling service of the present invention is initialized and started, the dynamic scheduling module will preset the node computing power c k as a relatively small initial value. During subsequent training, the node computing power c k is recalculated. If it is greater than the initially set c k initial value, then change the c k of the current node to the new value. If the recalculated node computing power c k is not greater than the initial value, then c does not need to be updated k . The node computing power c k is recalculated every once in a while. The best recommended value is to calculate it once every 5 seconds or 10 seconds; When the API gateway of the model training scheduling service receives a model training request, it will call the dynamic scheduling module. The dynamic scheduling module estimates that the total video memory required for the model is P, and the video memory of each shard in the present invention is set as S param , and the total number of shards required is F: F = ; Take a certain number of target training nodes n from the current available nodes. The computing power of each node is c k , and allocate the F shards to the selected nodes according to the ratio of the computing power of each node. When allocating, it is ensured that the total number of shards allocated to each node * the video memory occupancy of each model shard <= the total video memory of the node. If the condition is not met, then increase the number of target training nodes n until the condition is met; if the number of target training nodes n is expanded to all the current available nodes and still no nodes that meet the above rules are found, then expand the number of current available nodes. If it still cannot meet the above rules after expanding to the maximum, then end the current training.

[0030] After the shards of the model are assigned to the corresponding training nodes for training, after training is completed, each shard is aggregated and then transmitted to the result saving and fault tolerance processing module. This aggregation method is also called gradient aggregation. When aggregating, the nodes used for training are divided into M synchronization groups, and each synchronization group contains one or more model shards. When aggregating, first perform in-group aggregation. After the in-group aggregation is completed, then asynchronously aggregate between groups into the result saving and fault tolerance processing module, and finally form the finally trained model.

[0031] For the node computing power c k : The dynamic scheduling module will recalculate the node computing power c of each node every once in a while according to the number of shards processed during this period k , so as to ensure that the model shards can dynamically adjust the sharding strategy according to the node load. However, when some training nodes have poor network, the number of shards processed by the node during this period may be inaccurately counted, which will cause the value of the node computing power c k to also deviate, which will in turn affect the best way of model sharding. Therefore, a gradient residual compensation mechanism is introduced. The residual Rt can be obtained through real-time monitoring of the time window, or prediction compensation of historical fluctuations, or communication feedback mechanism, etc. When the residual Rt>0, mark this node as a delayed node, then when calculating the node computing power c k of this delayed node, take the average node computing power c k1 of several historical cycles of this delayed node, multiply it by a compensation coefficient, and then add it to the current node computing power c k2 to obtain the final node computing power of the current node.

[0032] The resource monitoring module detects the load of each training node in the node training module every once in a while. When reporting the load situation to the dynamic scheduling module, when the load situation is greater than a certain threshold, perform node expansion and add nodes. When the load situation is less than the threshold, perform node contraction and reduce nodes. For example: when the average value of the loads of all nodes exceeds 85% within 10 consecutive minutes, trigger node expansion. When the average value of the loads of all nodes exceeds 10% within 10 consecutive minutes and there are nodes without training tasks, remove 1 to 2 nodes.

[0033] Example 4: The calculation steps of model sharding and gradient aggregation are as follows: Optimize the sharding ratio through the Lagrange multiplier method, and its purpose is to: under the condition of meeting the computing resource constraints, find the model sharding scheme that makes the overall training efficiency the highest (or the training time the shortest).

[0034] The optimized formula for minimizing the training time T is: Among them, the constraints are as follows: · Node computing power limit (the number of model shards processed per unit time), · Memory capacity limit (the video memory occupied by the model shard), Among them, Wk is the shard ratio of the node k ; it can be obtained that: the shard ratio is proportional to the computing power of the node ( ), and the memory capacity ( ). is directly proportional.

[0035] Suppose a model needs to be allocated to three nodes ( K = 3), and their computing powers are c 1 = 100, c 2 = 80, c 3 = 60 (unit: allocated samples / second), and the memory capacities are M 1 = 24GB, M 2 = 16GB, M 3 = 12GB.

[0036] All shards need to meet the conditions: Formula 1: Formula 2: Among them: · : the number of model shards allocated to node k; · : the total number of model shards; · : the memory capacity limit; · : the video memory occupancy of each model shard; Finally, it can be obtained that: .

[0037] Suppose is set to 2GB; Node resources: Node 1: c 1 = 100, M 1 = 24GB → at most 12 layers; Node 2: c 2 = 80, M 2 = 16GB → at most 8 layers; Node 3: c 3 = 60, M 3 = 12GB → at most 6 layers; Nodes 1, 2, and 3 are different GPU resources, so their computing capabilities and video memory resources are different; a specific usage example 7 provides a GPU virtualization slicing solution.

[0038] Optimization results: Ideal allocation (without memory constraints): According to the computing power ratio of 5:4:3 → Node 1: 5 layers, Node 2: 4 layers, Node 3: 3 layers.

[0039] Actual allocation (constrained by memory): Node 1 is allocated 10 layers (close to the memory limit), Node 2 is allocated 2 layers, and Node 3 is allocated 0 layers (due to low computing power and insufficient memory).

[0040] Adjustment strategy: If the memory of Node 3 is insufficient, its sharding ratio can be reduced, and some layers can be dynamically migrated to Node 1 or 2.

[0041] Example 5: I. Gradient aggregation optimization: Apply a hierarchical aggregation protocol: Divide the nodes into synchronization groups , and define two-stage updates: Intra-group aggregation: : The result of the synchronization group aggregation K: The k-th shard within the group : The result of the k-th shard aggregation Inter-group asynchronous update: : The result of the (t + 1)-th asynchronous group aggregation : The result of the t-th asynchronous group aggregation : Learning rate : The sharding ratio coefficient of the m-th synchronization group : The result of the synchronization group aggregation for the t-th asynchronous group and the The following is the convergence analysis: Under the following conditions, the algorithm converges at a rate of : Learning rate ηt satisfies The upper bound of the inter-group delay is τ, the delay coefficient is , and II. Comparison of Communication Complexity: Among them is the number of synchronization groups, and the optimal communication-computation balance point is reached when .

[0042] III. Experimental Verification Data Compare the gradient consistency error during the training of the ResNet-152 model: IV. Core Functions 1. Model Sharding Technology Function for the present invention: Dynamic sharding strategy: Adjust the sharding strategy in real time according to the node computing power, and increase or decrease nodes according to the node load (GPU utilization rate, network bandwidth) feedback by the resource monitoring module.

[0043] Patent innovation point: Propose a non-linear mapping relationship between sharding weights and node computing power Technical effect: Automatically reduce the sharding ratio of high-load nodes to prevent overload; Support mixed training of heterogeneous devices (such as the scenario of mixing V100 and A100).

[0044] 2. Gradient Aggregation Technology Technical effect: Maintain the connectivity of the aggregation topology when nodes dynamically join / leave; Reduce the cross-region communication volume by 63% (measured data).

[0045] V. Comparison of Performance Improvement Example VI: I. Implementation steps of the gradient residual compensation mechanism: Gradient Residual ( Rk ): Refers to the amount of gradient calculation that the node k has not completed on time in the current training cycle, that is: Rk = Expected gradient amount to be completed - Actual gradient amount completed.

[0046] Method for the dynamic scheduling module to detect gradient residuals: 1. Real-time monitoring based on time window Index collection: The dynamic scheduling module periodically (such as every 1 second) collects the following data: The gradient generation rate c (unit: samples / second) of node k ; Nodek The number of processed samples Nprocessed(k); Global synchronization timestamp T sync (all nodes need to submit gradients before this time).

[0047] Residual calculation: At the synchronization deadline T sync, if the node k has not completed the expected sample size N expected( k ), then the residual is: Rk = max( ⋅(Tsync - Tstart) - Nprocessed(k), 0); where Tstart: the start time of the current training cycle.

[0048] Physical meaning: Predict the unfinished sample size according to the node's computing power and the remaining time.

[0049] 2. Prediction compensation based on historical fluctuations Sliding window statistics: The dynamic scheduling module maintains the historical residual mean k of the node μk and the variance of the past cycles σk , such as the data of the past 5 cycles.

[0050] If the current residual Rk exceeds μk +2 σk , it is determined as an abnormal delay and an emergency compensation is triggered.

[0051] 3. Communication feedback mechanism Active reporting: The node sends its status to the scheduling module in the following situations: (1) When it has completed 50% and 75% of the gradient calculation; (2) When it detects that the local computing rate has dropped by more than the threshold (such as 20%).

[0052] Passive detection: The scheduling module sends a lightweight probe (ping) to the nodes that have not responded in time, and estimates the remaining computing amount through the response time.

[0053] II. Implementation process of residual compensation 1. Residual detection: At T the sync time point, the scheduling module calculates the Rk of all nodes.

[0054] IfRk > 0, the marked node k is a "delay node".

[0055] 2. Residual filling: For the delay node k , extract the gradients of the most recent m cycles from the historical gradient cache pool { gk ( t -1), gk ( t -2),..., gk ( t - m )}.

[0056] Generate a compensation gradient using interpolation or weighted average: , where the weight ωi can be designed to decay exponentially (e.g., , α = 0.8) 3. Dynamic weight adjustment: Correct the weight of the compensation gradient according to the residual ratio: g final = (1 - β ) g real + βg ^ k , β = min( N expected( k ) Rk , 1) where g real: the partial gradient actually completed by node k For example: The computing power of node A = 100, and it needs to process expected( N ) = 500 samples in a certain cycle, and the cycle duration A ) = 500 samples, and the cycle duration T sync - T start = 5 seconds The actual processing N processed( A ) = 400 samples; The residual RA = 100 × 5 - 400 = 100 samples; Compensation gradient generation: Extract the gradients of the most recent 3 cycles from the cache gA ( t -1),gA ( t -2), gA ( t -3).

[0057] Use the weight ω =[0.6, 0.3, 0.1] for weighted average to obtain g ^ A .

[0058] Final gradient: g final = 0.7 g real + 0.3 g ^ A (assuming β = 0.3).

[0059] The dynamic scheduling module determines the gradient residual by monitoring the computing progress of nodes in real time, historical data analysis, and active / passive feedback mechanisms. The specific steps include: Step 1, Residual prediction based on computing power: Use Sc ( k ) and the remaining time to estimate the unfinished amount.

[0060] Step 2, Historical fluctuation analysis: Identify abnormal delays and trigger compensation.

[0061] Step 3, Gradient interpolation compensation: Fill the residual with historical gradients to avoid training stagnation.

[0062] The core of this mechanism is to balance real-time performance and accuracy, quickly respond to changes in node performance while minimizing communication overhead, and ensure the efficiency of distributed training.

[0063] III. The role of the gradient residual compensation mechanism for the present invention 1. Correct the gradient deviation when the network is poor Technical scenario: In cross-regional distributed training, there are significant differences in network bandwidth and latency between different nodes (such as the mixed deployment of edge nodes and cloud nodes), resulting in timing deviations during gradient aggregation.

[0064] Mechanism of action: Residual accumulation: Record the part of the historical gradient that was not timely participated in aggregation due to transmission delay.

[0065] 2. Suppress the update inconsistency caused by model sharding Technical scenario: Dynamic sharding causes the dimension of the gradient tensor calculated by different nodes to change dynamically (such as the sharding granularity changing from hierarchical to parameter block level), and traditional aggregation algorithms will introduce tensor alignment errors.

[0066] 3. Support seamless connection for elastic scaling Technical scenario: When nodes dynamically join or exit, the shard migration and aggregation topology reconstruction cause some gradient updates to be lost.

[0067] Compensation implementation: Before a node exits, write its unfinished gradient increment to the global cache pool.

[0068] Example 7: 1. Steps for the result saving and fault tolerance processing module to implement the adaptive gradient synchronization algorithm Step 1: Dynamically collect node performance characteristics 1.1 Calculation ability measurement: Count the number of training samples processed by each node within a fixed time window (typical value 5 - 15 seconds) to calculate the gradient generation rate. 1.2 Network efficiency evaluation: Record the start and end timestamps of node gradient data transmission to calculate the effective transmission bandwidth. 1.3 Calculation of heterogeneity index: Comprehensively calculate the node performance score based on the calculation ability and network efficiency. Step 2: Dynamically make communication strategy decisions 2.1 Judgment of difference degree threshold: Calculate the difference degree of gradient generation rate between nodes. 2.2 Network state classification: Divide the network state levels according to the cluster average transmission bandwidth. Low bandwidth state: ; High bandwidth state: ; 2.3 Policy selection logic: When ΔSc > 30% and in the low bandwidth state, enable the delayed synchronization policy. When ΔSc < 15% and in the high bandwidth state, enable the synchronous AllReduce policy. In other cases, enable the hierarchical aggregation policy.

[0069] Step 3: Construct the hierarchical aggregation topology 3.1 Optimization of node grouping: Dynamically cluster and group based on the performance score Adopt the improved K-Means algorithm, and the objective function is to minimize the intra-group performance difference.

[0070] 3.2 Double-layer aggregation architecture: Intra-group aggregation: Each member node in the group performs gradient averaging calculation; Inter-group update: The representative nodes of each group perform weighted gradient fusion.

[0071] Step 4: Delay synchronization compensation mechanism 4.1 Historical gradient caching: Maintain a circular buffer to store the gradient parameters of the last N times (the typical value of N = 5).

[0072] 4.2 Gradient interpolation compensation: When it is detected that the node gradient delay exceeds the threshold, generate a compensation gradient.

[0073] Where: is the compensation gradient, M is the number of gradient of the delayed node, is the compensation coefficient, is the delay gradient between the t-th gradient and the i-th gradient.

[0074] 4.3 Dynamic weight adjustment: Exponentially decay the compensation coefficient according to the delay time Where: T is the time decay constant, e is the natural constant, and are the gradient delay times of the i-th gradient and the j-th gradient respectively Step 5: Smooth transition of policy switching 5.1 State pre-synchronization: Broadcast metadata (including grouping structure, compression parameters, etc.) before the new policy takes effect; 5.2 Gradient format conversion: Achieve lossless conversion of gradient tensors of different policies through a double-buffer mechanism; 5.3 Hybrid mode transition: Perform weighted output of the old and new policies within the switching window period (the typical value is 3-5 iteration steps), Where: Weighted hybrid gradient, The amount of the old compensation gradient, The amount of the new compensation gradient, Weighting coefficient, .

[0075] Step 6: Abnormal state self-recovery 6.1 Heartbeat detection mechanism: Verify the node survival status at a fixed period (the typical value is 2 seconds); 6.2 Gradient integrity verification: Verify the integrity of the transmitted gradient data through the CRC32 check code; 6.3 Resume interrupted transfer processing: Trigger gradient recalculation for failed nodes: Wherein: is the recovery gradient from the failed node, K is the number of failed nodes, is the delay gradient of the k-th failed node. When updating the topology, retain the original node shard replicas for at least Tretain time (typical value 60 seconds) Wherein, 1. Dynamic performance evaluation method: Node performance evaluation includes a linear combination of time window sliding average calculation and network effective bandwidth estimation, where the time window length is adaptively adjusted according to the cluster scale: Wherein: is the time window length, is the total number of shards to the model, The number of shards currently running to.

[0076] 2. Hierarchical aggregation topology generation algorithm: The group partitioning method of the hierarchical aggregation strategy optimizes the clustering process by introducing node physical topology constraints. The constraint conditions are: Nodes within the same rack cannot be assigned across groups; The difference in communication latency between node groups across switches does not exceed 20%.

[0077] 3. Delay synchronization compensation mechanism: The delay gradient compensation adopts the exponential decay weight assignment method, and the decay coefficient τ is inversely proportional to the historical performance stability of the node: Wherein: represents the number of gradients synchronized per unit time, represents the variance of, represents the average value of.

[0078] 4. Comparison of technical effects Through the dynamic adjustment of the quantization index-driven strategy, the precise matching of communication efficiency and computing resources is achieved, and the training efficiency in heterogeneous environments is significantly improved compared with the existing technology. Form a closed-loop optimization system with real-time performance perception, policy decision-making model and gradient compensation mechanism, breaking through the limitations of traditional fixed strategies.

[0079] Example Seven: GPU virtualization slicing scheme: Inner layer channel: Based on the dynamic confusion encryption of model parameters, # Example of the core logic of the virtualized device management module class GPUSlicer: def __init__(self, physical_gpu): self.memory_total = physical_gpu.memory_capacity self.compute_units = physical_gpu.cuda_cores def create_slice(self, req_mem, req_compute): # The memory block uses non - contiguous address mapping technology memory_map = self._allocate_noncontiguous_mem(req_mem) # The compute units use time - division multiplexing scheduling compute_scheduler = TimeDivisionScheduler(req_compute) return VirtualGPU(memory_map, compute_scheduler) # Usage example physical_gpu = NVIDIA_A100() slicer = GPUSlicer(physical_gpu) vGPU1 = slicer.create_slice(10GB, 50%) # Create a virtual GPU with 10GB memory + 50% computing power vGPU2 = slicer.create_slice(5GB, 30%) # Create a virtual GPU with 5GB memory + 30% computing power Fault recovery process Health check interval: Perform a heartbeat detection every 10 seconds; Fault determination criterion: Three consecutive heartbeat losses and GPU utilization is 0; Status saving mechanism: Adopt gradient snapshot + model parameter differential saving.

[0080] The above - mentioned are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, or improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A distributed scheduling training method for AI model training, characterized in that: Including API gateway, resource monitoring module, dynamic scheduling module, node training module, and result storage and fault tolerance processing module; The API gateway receives the client training request and passes the information to the dynamic scheduling module; The dynamic scheduling module performs model sharding, delivers the model shards to the corresponding training nodes, and assigns training tasks to the node training modules; The node training module executes specific training tasks and returns the number of model shards processed per unit time to the dynamic scheduling module. The dynamic scheduling module dynamically adjusts the sharding strategy based on the returned data. The resource monitoring module collects the hardware indicators and network status of computing nodes in real time, and reports the node load to the dynamic scheduling module, which adjusts the number of nodes based on the reported data; After the training is completed, each shard is aggregated into an aggregate model that has completed the training, and the final aggregate model is passed to the result preservation and fault-tolerant processing module. The result preservation and fault-tolerant processing module saves the training results and senses whether the training node status is abnormal. If there is an abnormality, the training task will be migrated to other nodes to continue training.

2. A distributed scheduling training method for AI model training according to claim 1, characterized in that: The dynamic scheduling module includes an expansion unit and a load prediction unit; The expansion unit increases or decreases the number of physical nodes according to the relationship between the average node load reported by the resource monitoring module and the threshold. When the average node load is greater than the threshold, expansion is triggered, and when the average node load is less than the threshold, reduction is triggered. The load prediction unit predicts future resource requirements based on historical data, and increases or decreases the number of physical nodes when there is a large deviation between the predicted value and the current actual node resources.

3. A distributed scheduling training method for AI model training according to claim 1, characterized in that: When the dynamic scheduling module allocates model shards to the node training module, several training nodes are taken out from the currently available nodes, and several model shards are allocated to the taken out nodes according to the node computing capacity of each training node.

4. A distributed scheduling training method for AI model training according to claim 1, characterized in that: The dynamic scheduling module introduces a gradient residual compensation algorithm to perform error correction and convergence acceleration.

5. A distributed scheduling training method for AI model training according to claim 1, characterized in that: The node training module sets up a virtualized resource pool, and the virtualized resource pool performs fine-grained resource allocation.

6. A distributed scheduling training method for AI model training according to claim 1, characterized in that: The model sharding is optimized in batches proportionally via the Lagrange multiplier method.

7. A distributed scheduling training method for AI model training according to claim 1, characterized in that: The polymerization model completes polymerization through a gradient polymerization method.

8. A distributed scheduling training method for AI model training according to claim 1, characterized in that: The result preservation and fault-tolerant processing module adopts an adaptive gradient synchronization algorithm.

Citation Information

Patent Citations

  • Asynchronous distributed deep learning training method, device and system

    CN110245743A

  • Distributed training method and device for deep learning model

    CN112000473A

  • AI model training method and device, computing equipment and storage medium

    CN114154641A

  • Distributed training method, device and equipment based on end-to-end self-adaption

    CN114169427A

  • Intelligent model training method and device

    CN115600681A

Cited By

  • Node detection method and device, equipment, storage medium and computer program product

    CN121785876A

  • Node detection method, apparatus, device, storage medium, and computer program product

    CN121785876B