Classification model training method based on big data distributed computing
By dynamically adjusting the batch size and resource allocation between distributed computing nodes, the problem of inefficient synchronization caused by communication delay in distributed computing systems is solved, and more efficient model training and resource utilization are achieved.
Patent Information
- Application Number
- CN202510716453.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-30
AI Technical Summary
In a big data environment, the distributed computing system has reduced synchronization efficiency due to unbalanced network communication delay between nodes, and the model training time is extended and resource utilization is reduced.
By obtaining communication delay data between distributed computing nodes, analyzing the delay characteristics to generate regulation parameters, dynamically adjusting the batch size of each training batch, ensuring the optimal synchronization efficiency, and dynamically allocating resources based on the node communication efficiency and calculation pressure.
It improves the accuracy of training scheduling and real-time adaptability, avoids resource waste and training blockage under traditional static batch strategies, improves parallel efficiency and task matching, and enhances model training speed and global collaboration capabilities.
Smart Images

Figure CN120234158A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing and machine learning, and particularly relates to a classification model training method based on big data distributed computing. Background Art
[0002] In the existing big data environment, the training of machine learning models usually relies on a distributed computing architecture to improve the training speed and the ability to process ultra-large-scale data sets. The distributed computing system distributes data and computing tasks to multiple computing nodes and coordinates and synchronizes parameters through communication between nodes to achieve joint training of the model. Among them, distributed synchronization mechanisms, such as synchronous SGD, are widely used in the training process to ensure the consistency of model parameters among nodes. However, due to the unevenness of network communication delays between nodes, especially when the number of nodes is large or widely distributed, the synchronization efficiency drops significantly, and the overall training process is easily affected by the slow node effect, resulting in an extended model training time and reduced resource utilization. To address the synchronization problem caused by communication delays, various optimization strategies have been proposed in the prior art, such as reducing the synchronization frequency by increasing the batch size, or adopting an asynchronous update mechanism, such as asynchronous SGD, to relieve the synchronization pressure. However, these methods still have limitations: the adjustment of the batch size is usually a static configuration and it is difficult to dynamically adapt to the real-time communication conditions between nodes, resulting in a waiting bottleneck still existing in high-latency scenarios; although asynchronous updates can improve the training parallelism, they will introduce the problem of decreased parameter consistency, thus affecting the convergence speed and accuracy of the final model. Summary of the Invention
[0003] The purpose of the present invention is to provide a classification model training method based on big data distributed computing, which adjusts the batch size according to the communication delay between nodes to solve the problem of low distributed synchronization efficiency.
[0004] To achieve the above purpose, the present invention provides the following technical solution: A classification model training method based on big data distributed computing, the method comprising: S1. Obtain the communication delay data between distributed computing nodes, and analyze the delay characteristics to generate regulation parameters, including collecting network delay data of each node once every fixed time, setting a delay threshold, accumulating the sample periods with delays less than the threshold as the effective communication time, and calculating the communication efficiency of each node; S2. Dynamically determine the batch size of each training batch based on the regulation parameters to ensure the optimization of synchronization efficiency, including calculating the actual training batch size of each node according to the communication efficiency of each node, discarding or filling the non-integer part, and writing the adjusted result into the distributed training management scheduler for task distribution; S3. Allocate the adjusted batch size to each node and execute the forward and backward propagation calculation processes of the model, including collecting the average response time of the previous round of training for each node, calculating the computational pressure corresponding to the current batch task, and allocating the thread concurrency number and I / O waiting queue length according to the computational pressure corresponding to the current batch task. S4. Collect the calculation results of each node for parameter aggregation and update to complete the iterative training of the classification model, including each node uploading a local copy of the model parameters, calculating the node gradient volatility, and then taking its reciprocal to define the weight, aggregating the final global parameters, synchronizing the aggregated model update parameters to all nodes, and starting the next round of iteration.
[0005] Preferably, calculating the communication efficiency of each node in S1 includes collecting network communication delay data from each distributed computing node at fixed time intervals in sequence. Within each time interval, judge all the collected delay samples, set a delay threshold, filter out the samples smaller than the set threshold, and count the total observation time of the current time window, that is, the total time period experienced from the start of collection to the end. Divide the effective communication time by the total observation time to obtain the communication efficiency value of the node, and use the communication efficiency value as the communication regulation parameter of the node.
[0006] Preferably, calculating the actual training batch size of each node in S2 includes receiving the calculated communication efficiency values of each node, setting a standard batch size as a reference benchmark for the initial training task volume, taking the communication efficiency value of each node as the weight, dividing it by the sum of the weight of the communication efficiency value of each node and one, and then multiplying by the standard batch size to obtain the preliminary training batch size of the node.
[0007] Preferably, calculating the computational pressure corresponding to the current batch task in S3 specifically includes collecting the average response time experienced by each distributed computing node in the previous round of training, combining the size of the training batch task currently assigned to the node, and performing a multiplication operation with the average response time to obtain the overall computational pressure of the node's current round of task.
[0008] Preferably, S3 further includes automatically allocating corresponding hardware execution resources to each node according to the calculated pressure value, including the number of threads that can be used in parallel, the length of the input / output buffer queue, and the task prefetch parameter.
[0009] Preferably, S3 further includes immediately starting the training calculation process after the node completes the resource configuration, including immediately starting the training calculation process after the node completes the resource configuration, including performing forward propagation for the predictive calculation of the model output, performing error calculation for measuring the difference between the output and the actual target, performing backward propagation for updating the internal parameters of the model, and outputting the corresponding gradient information or model update result.
[0010] Preferably, the aggregated final global parameters in S4 specifically include receiving the model parameter copies uploaded from each distributed computing node, calculating the fluctuation degree of the parameter gradients, using the fluctuation degree as the reference basis for weights, adjusting it by taking its reciprocal, and integrating the model parameters of all nodes in a weighted fusion manner to calculate a final global parameter representing the collaborative training results of all nodes, and synchronously distributing the integrated global model parameters to all nodes for use as the initial parameters for the next round of training.
[0011] Preferably, calculating the fluctuation degree of the parameter gradients in S4 includes calculating the standard deviation of each parameter item in the model parameters uploaded by each node in the current training round, and using this standard deviation as the basis for measuring the stability of the model update of this node.
[0012] Preferably, the weighted fusion method in S4 is to sum the products of the weights of each node and the model parameter values, and then divide by the sum of the weights of all nodes to obtain the final global model parameters.
[0013] Preferably, the synchronous distribution of the integrated global model parameters in S4 adopts a broadcast mechanism, and uniformly distributes the final aggregation result to all distributed computing nodes at the same timestamp to ensure the consistency and synchronization of parameter updates.
[0014] As can be seen from the above technical solutions, the present invention has the following beneficial effects: This classification model training method based on big data distributed computing obtains the communication delay data between distributed computing nodes, analyzes the delay characteristics to generate regulation parameters, dynamically determines the batch size of each training batch based on the regulation parameters to ensure the optimization of the synchronization efficiency, distributes the adjusted batch size to each node, executes the forward and backward propagation calculation processes of the model, collects the calculation results of each node for parameter aggregation and update, completes the iterative training of the classification model, realizes the quantitative modeling of the node network state, improves the accuracy and real-time adaptability of training scheduling, avoids the resource waste and training blockage caused by the traditional static batch strategy in high-delay or heterogeneous computing environments, improves the parallel efficiency and task matching degree in the training process, improves the flexibility of node-level training scheduling and the utilization efficiency of the overall system resources, gives greater weights to the nodes with high training stability in the parameter fusion process, realizes the enhancement of the robustness of model parameter updates and the improvement of the convergence quality, realizes the adaptive parallel training regulation in heterogeneous network environments, effectively alleviates the impact of slow node effects on the training cycle, improves the model training speed and global collaboration ability, can significantly improve the model training efficiency, resource utilization rate and parameter aggregation quality in complex and dynamically changing distributed environments, and has good engineering practical value and promotion prospects. Description of the Drawings
[0015] Figure 1 This is the flowchart of the method of the present invention. Detailed implementation manners
[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0017] As Figure 1 shown, the present invention provides a technical solution: a classification model training method based on big data distributed computing, and the method includes: S1. Obtain the communication delay data between distributed computing nodes, and analyze the delay characteristics to generate regulation parameters, including collecting the network delay data of each node once every fixed time, setting a delay threshold, accumulating the sample time periods with delays less than the threshold as the effective communication time, and calculating the communication efficiency of each node; S2. Dynamically determine the batch size of each training batch based on the regulation parameters to ensure the optimization of the synchronization efficiency, including calculating the actual training batch size of each node according to the communication efficiency of each node, discarding or filling the non-integer part, and writing the adjusted result into the distributed training management scheduler for task distribution; S3. Allocate the adjusted batch size to each node, and execute the forward and backward propagation calculation processes of the model, including collecting the average response time of the previous round of training of each node, calculating the calculation pressure corresponding to the current batch task, and allocating the thread concurrency number and the I / O waiting queue length according to the calculation pressure corresponding to the current batch task; S4. Collect the calculation results of each node for parameter aggregation and update to complete the iterative training of the classification model, including each node uploading a local model parameter copy, calculating the node gradient volatility, then taking its reciprocal as the weight, aggregating the final global parameters, synchronizing the aggregated model update parameters to all nodes, and starting the next round of iteration.
[0018] The classification model training method based on big data distributed computing proposed in the present invention adopts an adaptive optimization mechanism for dynamic network environment and heterogeneous computing resources, covering key links such as communication delay evaluation, batch size control, computing load management and parameter aggregation. Its core idea is to dynamically perceive the changes in communication and computing performance between nodes, and adjust the training resource configuration and collaboration strategy in real time to achieve high efficiency and convergence stability in the distributed classification model training process. In step S1, the system samples the network delay of all participating nodes at a set time interval to obtain real-time delay information of the cross-node communication link. After the delay data is processed by threshold comparison, only the low-delay effective communication period is retained for statistical analysis, and then the effective ratio of each node that can be used for synchronous communication per unit time is calculated and defined as the communication efficiency parameter. This parameter not only reflects the network state, but also indirectly characterizes the ability of the node to participate in global synchronization, and is an important input for subsequent control strategies. In step S2, using the above communication efficiency, the system can dynamically allocate the optimal training batch size for each node. Through the positive correlation between batch size and communication efficiency, it is ensured that nodes with strong communication performance undertake larger training tasks, while nodes with weak communication capabilities are burdened with less burden, thereby shortening the overall synchronization waiting time. In order to maintain computational consistency and data integrity, the non-integer part of the batch size is discretized by discarding or minimum padding algorithms, and instructions are uniformly issued by the scheduling module to realize the distributed scheduling of training tasks. In step S3, each node performs forward propagation and back propagation operations according to the assigned training tasks. In order to further balance the computational pressure between nodes, the system monitors the average response time of each node in the previous round, and calculates the predicted computational pressure in combination with the current task volume. The number of concurrent threads and the length of the I / O queue are dynamically adjusted according to the pressure value to match the node processing capacity with the task load, improve local execution efficiency and avoid resource blocking. In step S4, after all nodes complete the calculation, they upload their own local model parameter copies. The master node or the designated aggregation node is responsible for collecting all parameters, and calculates the gradient volatility according to the gradient fluctuation of each node, and then defines it as the parameter weight by taking its inverse. This strategy can automatically improve the contribution of nodes with small differences in training stability, suppress the influence of nodes with severe fluctuations, and effectively prevent training instability caused by gradient oscillation. Finally, parameter aggregation is achieved through weighted summation, new global model parameters are constructed, an iterative update is completed and synchronized to each computing node, and preparations are made for the next round of training. Overall, this method opens up the feedback loop among the three levels of data communication, resource scheduling and model synchronization, realizes the collaborative optimization of the training process in a distributed environment, and ensures the high speed, robustness and adaptability of model training.
[0019] This method can dynamically adapt to differences in network communication and computing resources in a distributed training environment, and optimize model training efficiency and stability. Specifically, the communication efficiency is used to guide batch size allocation, which effectively reduces waiting delays between nodes and improves synchronization efficiency. The resource scheduling mechanism driven by computing pressure ensures load balancing and system response speed. The aggregation mechanism based on the inverse weight of gradient volatility improves model convergence stability and accuracy. At the same time, this method has good scalability and environmental adaptability, and is suitable for a variety of large-scale machine learning scenarios.
[0020] The calculation of the communication efficiency of each node in S1 includes collecting network communication delay data from each distributed computing node in turn at fixed time intervals. In each time interval, all collected delay samples are judged, a delay threshold is set, and samples less than the set threshold are screened out. The total observation time of the current time window, that is, the total time period from the start of collection to the end, is counted. The communication efficiency value of the node is obtained by dividing the effective communication time by the total observation time, and the communication efficiency value is used as the communication control parameter of the node.
[0021] The communication efficiency evaluation mechanism proposed in this embodiment is to quantitatively analyze the network status of distributed nodes in units of time windows. Within a fixed time interval, the system polls all computing nodes in turn and collects network delay data from each node once. These data can be measured by the response time between the node and the central scheduler. In each collection cycle, the system centrally processes the collected delay samples, first setting a predetermined delay upper limit value as the baseline for judging the communication quality. Then the collected delay samples are judged one by one, and those samples with delay times less than the set upper limit are screened out, and the duration period occupied by these samples in the current time window is accumulated. The accumulated time is the effective communication time of the node in this cycle, which is used to indicate the total time it can maintain in a good communication state. At the same time, the total time experienced from the start to the end of the collection cycle is defined as the total observation time. Subsequently, the system compares the effective communication time with the total observation time to obtain the communication efficiency of the node in this time window. Specifically, if a node experiences a total of ten minutes in this time window, and the time in which the communication state is good is six minutes, it can be inferred that the communication efficiency of the node is the ratio of six minutes to ten minutes. This ratio is the communication efficiency value, which indicates the ability of the node to maintain high-quality communication during the observed period. The communication efficiency value finally obtained will be stored as the node's control parameter, which is used to guide the batch task allocation, computing resource scheduling, and synchronization frequency tuning strategies in the subsequent training process, so as to achieve dynamic matching of node resources and network status, and ensure the synchronization efficiency and global stability of model training.
[0022] Through the refined communication efficiency calculation process, this embodiment can more accurately distinguish the differences in network communication capabilities between different nodes, making the distributed training task scheduling more scientific and targeted. The specific advantages include: First, it avoids the overall communication performance evaluation accuracy being dragged down by extreme delay samples; second, it adopts a sliding time window mechanism to enhance real-time performance and self-adaptability; third, it realizes the automatic dynamic update of node control parameters, improving the communication efficiency and resource utilization rate during the model training stage. Compared with the coarse-grained network detection method, this method is more suitable for the stable training requirements in complex network environments.
[0023] In S2, calculating the actual training batch size of each node includes receiving the calculated communication efficiency values of each node, setting a standard batch size as a reference benchmark for the initial training task volume, taking the communication efficiency value of each node as a weight, dividing it by the sum of the communication efficiency value of each node as a weight and one, and then multiplying by the standard batch size to obtain the preliminary training batch size of the node.
[0024] This embodiment proposes a method for determining the batch size of training tasks weighted by communication efficiency in response to the impact of node network capacity differences on synchronous training efficiency in distributed computing. Its core idea is to dynamically adjust the task load based on the node communication efficiency value, enabling each node to undertake a workload matching its communication ability during training and avoiding training delays caused by communication bottlenecks. During the specific operation process, the system first receives the communication efficiency parameters calculated by each distributed node in the previous stage. Subsequently, the system sets a standard batch size, which serves as a benchmark for measuring the total amount of training tasks in this round. On this basis, taking the communication efficiency of each node as a weight factor, first sum the communication efficiency values of all nodes to obtain a total reference value of communication capacity. Then, for each node, the system takes its communication efficiency value as its own contribution ratio and forms an overall weight ratio together with other nodes. By allocating the standard batch size according to the relative proportion of node communication efficiency, the preliminary training batch size of each node can be obtained. This batch size is proportional to the network communication ability of the node. Nodes with smoother communication will receive more training tasks, while nodes with poorer communication will have a relatively reduced task volume. Since the communication efficiency value may be a decimal or non-integer value, the initially calculated batch size may also result in a decimal. To maintain the integrity of the training data structure and system consistency, appropriate rounding or interpolation processing can be performed on the non-integer part, and finally, the adjusted value is used as the training data task volume for the current iteration of the node and input into the scheduler for distribution to each node for execution.
[0025] Through this method of calculating the batch size weighted by communication efficiency, the system can more precisely match the training load with the node capabilities, significantly improving the synchronization efficiency and data throughput rate during the overall distributed training process. It effectively reduces the drag of slow communication nodes on the entire training progress and reduces the waste of computing resources caused by waiting. At the same time, this mechanism has good scalability and generality, and can adjust the standard batch reference value and weight composition according to different training objectives and network conditions, enhancing the system's adaptive optimization ability.
[0026] Calculating the computational pressure corresponding to the current batch task in S3 specifically includes collecting the average response time experienced by each distributed computing node in the previous round of training, combining the size of the training batch tasks currently assigned to the node, and performing a multiplication operation with the average response time to obtain the overall computational pressure of the node for this round of tasks.
[0027] This embodiment provides a method for evaluating the computational pressure of nodes by combining historical response characteristics and the current task scale, aiming to achieve dynamic perception of the usage status of computing resources and load regulation. During each round of model training, the system records the average response time required for each distributed computing node to complete its assigned training tasks. This response time can be obtained by dividing the total duration from the start of the training task to the node submitting the result by the number of batches, reflecting the execution efficiency and operating status of the node. When entering the next round of training task allocation phase, the system reads the size of the training batches already assigned to the current node, combines it with the average response time of the node in the previous round, and conducts a quantitative analysis. By multiplying the current training batch size by the average response time, a comprehensive index can be obtained, indicating the level of computational pressure that the node needs to bear in the current training round. This computational pressure value not only reflects the quantity of the task load but also the rate at which the node processes tasks, and is a key parameter for judging whether there is an overload risk for the node and whether resource compensation or scheduling adjustment is needed. The system can adjust resource parameters such as thread concurrency configuration and I / O queue length based on this pressure value, optimize the node processing path, balance the internal resource distribution of the system, and dynamically reduce tasks or reallocate them when necessary.
[0028] Through the computational pressure evaluation mechanism formed by integrating historical response time and the current task scale, it is possible to achieve quantitative modeling of the computational load of each node, making resource scheduling more precise and dynamic. This method avoids the allocation errors caused by simply relying on static hardware metrics, can reflect changes in the operating state in real time, and effectively prevents problems such as computational overload or resource idleness. It further improves the stability, execution efficiency, and fault tolerance of the training system, enhancing the reliability and flexibility of model training in a large-scale distributed environment.
[0029] S3 also includes automatically allocating corresponding hardware execution resources to each node according to the calculated pressure value, including the number of threads that can be used in parallel, the length of the input / output buffer queue, and the task prefetch parameter.
[0030] In this embodiment, to further improve the node execution efficiency and system resource utilization rate during the distributed model training process, based on the currently calculated computing pressure values of each node, the system introduces an automated hardware resource configuration mechanism. This mechanism takes the computing pressure as input and dynamically adjusts the hardware operating parameters of the node according to its magnitude, enabling the resource configuration to precisely match the task load. When the system calculates the training batch computing pressure of a certain node, it will search for the optimal resource allocation strategy in the preset resource scheduling mapping rules based on this pressure value. For example, for a node with a higher pressure, the system will allocate more parallel computing threads to it to accelerate the processing rate; at the same time, expand the length of its input / output buffer to reduce the blocking waiting caused by the I / O bottleneck; and also increase the task prefetch depth so that the task can be pre-loaded into the execution queue before execution, thereby reducing the scheduling delay. On the contrary, for a node with a lower pressure, the system can tighten the resource configuration to avoid resource waste and transfer the surplus resources to high-load nodes for use. This configuration process is automatically completed by the distributed training scheduler without manual intervention, and can achieve millisecond-level response adjustment to cope with the sudden imbalance problem caused by the load change between nodes. After the resource allocation is completed, all adjusted parameters will take effect immediately to adapt to the execution of the next round of model training tasks. The system will re-evaluate the computing pressure and update the resource strategy in the next training round, forming a closed-loop optimization control mechanism.
[0031] This embodiment realizes the dynamic matching between the node computing power and the task load by introducing a hardware resource automatic configuration mechanism driven by computing pressure, effectively improving the overall computing efficiency and parallel performance of the system. It avoids the inefficient problems of node resource overload or idleness under static resource configuration, and significantly improves the response speed and stability of the training process. At the same time, this method requires no manual intervention and has high self-adaptability, especially suitable for large-scale distributed environments with a large number of nodes and large fluctuations in computing load.
[0032] S3 also includes immediately starting the training calculation process after the node completes the resource configuration, including immediately starting the training calculation process after the node completes the resource configuration, including performing forward propagation for predictive calculation of the model output, performing error calculation for measuring the difference between the output and the actual target, performing backpropagation for updating the internal parameters of the model, and outputting the corresponding gradient information or model update result.
[0033] This embodiment further expands the specific training operation process after the calculation of pressure assessment and resource allocation is completed, and proposes a design pattern in which the execution of training tasks should be closely coupled with the resource allocation process. After the above steps complete node resource scheduling and hardware parameter adjustment, each computing node immediately starts the core computing stage of a new round of model training to ensure the continuity of resource use and the immediacy of scheduling response. The training process includes three key steps. First, the node starts the forward propagation process based on the current batch of data. In this process, the model sequentially performs neural computations for each layer, mapping the input feature data to output prediction values. This output result is the prediction behavior of the model in the current state and is used to compare with the true label. Subsequently, the system enters the error calculation stage. By comparing the prediction result generated by forward propagation with the actual target label, the prediction deviation is quantified, and this error value is used to evaluate the current training effect of the model. The error calculation method can be selected according to the model type. For example, the cross-entropy is used for classification models, and the mean squared error is used for regression models, etc. Immediately afterwards, the backpropagation operation is performed. This process takes the error as the input, backpropagates layer by layer and calculates the gradient values, and gradually updates the various parameters in the model, thereby realizing the optimization of the model weights. During this process, each layer adjusts its weights according to the gradient information of the previous round and the learning rate strategy, and outputs the corresponding gradient information or intermediate model update results for subsequent aggregation or local caching. The entire computing process is completed locally at the node. After the calculation is completed, the relevant results are sent back to the scheduling center or participate in the parameter aggregation process to realize the actions of the next stage of the training system.
[0034] By immediately starting the training process after resource configuration, this embodiment shortens the resource idle time, improves the response efficiency and execution compactness of the training task. Integrating forward propagation, error calculation, and backpropagation for unified management makes the model training process continuous and smooth, optimizes the use of computing resources and the data flow structure, further reduces system latency, and enhances the overall throughput capacity of the system. The training results are directly output as aggregable gradients or model update information, which is convenient for subsequent unified processing.
[0035] Specifically, aggregating the final global parameters in S4 includes receiving the model parameter copies uploaded from each distributed computing node, statistically analyzing the degree of fluctuation of the parameter gradients, using the degree of fluctuation as a reference basis for the weights, and adjusting them by taking the reciprocal. Using the weighted fusion method, the model parameters of all nodes are integrated, and a final global parameter representing the collaborative training results of all nodes is calculated. The integrated global model parameters are synchronously distributed to all nodes and used as the initial parameters for the next round of training.
[0036] This embodiment proposes a model parameter aggregation strategy based on adaptive weighting of gradient volatility, aiming to overcome the problem of aggregation instability caused by node performance differences or data heterogeneity in multi-node distributed training. After a training round is completed, the system sequentially receives the local model parameter copies uploaded from all distributed nodes, and each parameter copy is the weight result independently trained by the current node. To ensure the stability and representativeness of the aggregation result, the system needs to first evaluate the gradient fluctuation degree of each parameter copy. Specifically, the system statistically calculates the change amplitude of each weight in the model parameters among nodes during multiple training rounds to measure the stability of the parameter update of this node. If the model parameters of a certain node change violently, it indicates that there may be noise, data deviation, or abnormal learning rate in its training process; while the node with gentle changes may be trained on a more stable data set and has higher convergence stability. The system uses the gradient fluctuation degree as a reference index for evaluating reliability. Next, the aggregation weight of each node is constructed by taking the reciprocal of the fluctuation degree, so that the nodes with smaller fluctuations account for a larger proportion in the global parameter fusion, and vice versa. Subsequently, the system fuses all parameter copies in a weighted manner to generate a new global model parameter. This fusion process can take into account the training effect and node reliability, and effectively prevent individual abnormal nodes from having a negative interference on model convergence. Finally, the system synchronously pushes the fused global parameters to all computing nodes as the starting model for the next round of training, ensuring that all nodes continue to train under the same initial state, maintaining model consistency and collaborative optimization.
[0037] The global parameter aggregation method using the reciprocal weighting of gradient volatility significantly improves the robustness and convergence quality of distributed training. By suppressing the influence of nodes with large volatility, it effectively inhibits the perturbation caused by abnormal updates and improves the stability and accuracy of the aggregated model. At the same time, this method can adaptively reflect the node operation status and training stability, effectively enhancing the robustness and fault tolerance of the training system, and is particularly suitable for complex distributed environments with data imbalance, node heterogeneity, or high training instability.
[0038] The statistical calculation of the gradient fluctuation degree in S4 includes calculating the standard deviation of each parameter item in the model parameters uploaded by each node in the current training round, and using this standard deviation as the basis for measuring the model update stability of this node.
[0039] In this embodiment, the "gradient volatility" in global parameter aggregation is further quantified. Using the standard deviation as a metric, a more operable model update stability evaluation system is constructed. After each round of training tasks is completed, each distributed node uploads its local model parameter copy obtained from the current training to the aggregation node. This parameter copy is usually a set of weight values, corresponding to each neuron connection or computing unit in the model structure. After the aggregation node receives the model parameter copies of all nodes, for each model parameter item (i.e., each specific weight value position), it statistically analyzes the set of parameter values at this position from all nodes. The system regards this set of values of the same parameter item from different nodes as a distribution and calculates its standard deviation. This standard deviation reflects the degree of deviation of this parameter in the training results of each node, and further reflects the degree of consistency of the model in this parameter dimension. For each node, the system performs induction or weighted average processing on the standard deviations corresponding to all its parameter items, so as to form an overall standard deviation index, which is used as the quantitative basis for the stability of the node's model update in this round. If the standard deviation is small, it indicates that the training process of this node is relatively stable and the update result is more inclined to the global average; if the standard deviation is large, it may indicate that it is affected by factors such as data perturbation, learning rate oscillation, or model non-convergence. This standard deviation index is finally used to construct the weights in weighted aggregation. By taking the reciprocal of the standard deviation or constructing a proportionality factor, the system gives nodes with small fluctuations higher weights, so as to highlight their leading role in global parameter integration and suppress the contributions of nodes that may introduce instability.
[0040] By introducing the standard deviation as a measure of model parameter fluctuations, this embodiment can more accurately and objectively measure the consistency and reliability of node training results. This method is easy to implement, has low computational cost, and can be extended to large-scale parameter spaces, with good applicability. Its weighted strategy improves the robustness and model convergence accuracy of the aggregation process, reduces the risk of error amplification caused by individual abnormal nodes, and helps to achieve a higher-quality global model training effect.
[0041] The weighted fusion method in S4 is to sum the products of the weights of each node and the model parameter values, and then divide by the sum of the weights of all nodes to obtain the final global model parameters.
[0042] This embodiment specifically defines a weighted fusion strategy for global model parameters, adopting a weighted average mechanism based on the sum of the products of node weights and corresponding parameter values followed by normalization. This mechanism aims to achieve precise control in the parameter aggregation process, enabling the aggregation result to more fully reflect the differences in the training stability and contribution degree of each node model. After each round of model training, all nodes in the distributed system upload their locally updated copies of model parameters to the central aggregation module. The model parameters of each node include multiple scalar values, with each scalar corresponding to a certain weight term in the model structure. At the same time, the system has obtained the weight coefficient corresponding to each node according to the aforementioned standard deviation calculation method or other stability evaluation mechanisms, which is used to indicate its influence in the global aggregation. For each dimension of the model parameter item, the system multiplies the parameter value of this item of all nodes by their corresponding node weights respectively, and then adds up all the product results to form the weighted accumulation value of this parameter item. Subsequently, the accumulation value is divided by the sum of all node weights to obtain the fusion value of the current parameter item in the global model. This operation is repeated for each model parameter item, and finally a global model parameter set containing all parameter items is generated, that is, the unified model representing the training result of the entire system after the current training round. This fusion mechanism ensures that nodes with small fluctuations and stable contributions have a higher proportion in the final model, while effectively alleviating the deviation risk introduced by abnormal nodes.
[0043] By adopting the aggregation strategy of weighted average of the product of node weights and parameter values, this embodiment realizes the logical unity between training stability and aggregation weight on the basis of maintaining the mathematical rigor of parameter fusion. This method has a clear structure and simple operations, is suitable for parallel computing in a large-scale parameter space, and is particularly applicable to the distributed training framework of deep learning models. Its result has good numerical stability and global representation ability, which can effectively improve the model convergence quality and overall performance.
[0044] In S4, the synchronized distribution of the integrated global model parameters adopts a broadcast mechanism, and the final aggregation result is uniformly sent to all distributed computing nodes at the same timestamp to ensure the consistency and synchronization of parameter updates.
[0045] This embodiment focuses on the distribution mechanism of the last step in model training - global parameter update, and proposes to control the distribution in a unified time through the broadcast method to ensure that all distributed nodes receive the latest model parameters synchronously at the same logical time point. The core goal of this design is to prevent the model state from being inconsistent due to nodes receiving parameters asynchronously, thereby affecting the training stability and convergence efficiency. After completing the weighted fusion of the global model parameters, the system generates a new global parameter set. This set represents the unified expression of the training results of each node in the current training round, and it is necessary to ensure that all subsequent training operations are based on this result. Therefore, the system uses the broadcast mechanism to send this parameter set to all nodes. The broadcast mechanism is initiated by the scheduling center, and all receiving nodes are in a listening state. Once the broadcast signal is received, the local model state can be updated immediately. To ensure logical synchronization, the system attaches a unified timestamp to the broadcast message, which marks the training number or time sequence index of the current round. After receiving the broadcast parameters, all nodes need to verify the timestamp consistency. Only when the timestamp matches the current round identifier, the parameters are loaded into the local model. If there are delayed or out-of-step nodes, the system can set a retry mechanism or retransmission strategy to ensure that all nodes finally successfully synchronize this version of the parameters. This synchronization process is triggered immediately after the parameter aggregation is completed, and the next round of training starts after all nodes confirm the reception, so as to ensure that all nodes are always in the same training state during the model iteration process, and avoid accuracy deviation or training errors caused by inconsistent model versions.
[0046] By adopting the broadcast mechanism with timestamp for parameter synchronization, this embodiment greatly improves the consistency control level in the distributed model training system. This mechanism ensures that all nodes are consistent when the global state is updated, and avoids model divergence caused by asynchronous updates or network delays. At the same time, it improves the overall synchronization and convergence stability of the training system, and is especially suitable for tasks with high consistency requirements, such as federated learning, multi-task concurrent training and other scenarios.
[0047] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made therein without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A classification model training method based on big data distributed computing, characterized in that, The method includes: S1. Obtain the communication delay data between distributed computing nodes, and analyze the delay characteristics to generate regulation parameters, including collecting network delay data of each node once every fixed time, setting a delay threshold, accumulating the sample periods with delays less than the threshold as the effective communication time, and calculating the communication efficiency of each node; S2. Dynamically determine the batch size of each training batch based on the regulation parameters to ensure the optimization of synchronization efficiency, including calculating the actual training batch size of each node according to the communication efficiency of each node, discarding or filling the non-integer part, and writing the adjusted result into the distributed training management scheduler for task distribution; S3. Allocate the adjusted batch size to each node and execute the forward and backward propagation calculation processes of the model, including collecting the average response time of the previous round of training of each node, calculating the computational pressure corresponding to the current batch task, and allocating the number of concurrent threads and the length of the I / O waiting queue according to the computational pressure corresponding to the current batch task; S4. Collect the calculation results of each node for parameter aggregation and update to complete the iterative training of the classification model, including each node uploading a local copy of the model parameters, calculating the node gradient volatility, then taking its reciprocal as the weight, aggregating the final global parameters, synchronizing the aggregated model update parameters to all nodes, and starting the next round of iteration.
2. The classification model training method based on big data distributed computing according to claim 1, wherein: Calculating the communication efficiency of each node in S1 includes collecting network communication delay data from each distributed computing node in turn at fixed time intervals. In each time interval, judge all the collected delay samples, set a delay threshold, filter out the samples smaller than the set threshold, and count the total observation time of the current time window, that is, the total time period from the start of collection to the end. Divide the effective communication time by the total observation time to obtain the communication efficiency value of the node, and use the communication efficiency value as the communication regulation parameter of the node.
3. A classification model training method based on big data distributed computing according to claim 1, characterized in that: Calculating the actual training batch size of each node in S2 includes receiving the calculated communication efficiency values of each node, setting a standard batch size as a reference benchmark for the initial training task volume, using the communication efficiency value of each node as the weight, dividing it by the sum of the weight of the communication efficiency value of each node and one, and then multiplying by the standard batch size to obtain the preliminary training batch size of the node.
4. A classification model training method based on big data distributed computing according to claim 1, characterized in that: Calculating the computational pressure corresponding to the current batch task in S3 specifically includes collecting the average response time experienced by each distributed computing node in the previous round of training, combining the size of the training batch task currently assigned to the node, and performing a multiplication operation with the average response time to obtain the overall computational pressure of the node for this round of task.
5. A classification model training method based on big data distributed computing according to claim 4, characterized in that: S3 also includes automatically allocating corresponding hardware execution resources to each node according to the calculated pressure value, including the number of threads that can be used in parallel, the length of the input / output buffer queue, and the task prefetch parameters.
6. The classification model training method based on big data distributed computing according to claim 5, wherein: S3 also includes immediately starting the training calculation process after the node completes resource configuration, including performing forward propagation for predictive calculation of model output, performing error calculation to measure the difference between the output and the actual target, performing backpropagation to update the internal parameters of the model, and outputting the corresponding gradient information or model update result.
7. A classification model training method based on big data distributed computing according to claim 1, characterized in that: In S4, aggregating the final global parameters specifically includes receiving the model parameter copies uploaded from each distributed computing node, statistically calculating the fluctuation degree of the parameter gradients, using the fluctuation degree as a reference basis for weights, and adjusting it by taking its reciprocal. Then, using the weighted fusion method, integrating the model parameters of all nodes, calculating a final global parameter representing the collaborative training results of all nodes, and synchronously distributing the integrated global model parameters to all nodes as the initial parameters for the next round of training.
8. A classification model training method based on big data distributed computing according to claim 7, characterized in that: In S4, statistically calculating the fluctuation degree of the parameter gradients includes calculating the standard deviation of each parameter item in the model parameters uploaded by each node in the current training round, and using this standard deviation as a measure of the model update stability of the node.
9. A classification model training method based on big data distributed computing according to claim 7, characterized in that: The weighted fusion method in S4 is to sum the products of the weights of each node and the model parameter values, and then divide by the sum of the weights of all nodes to obtain the final global model parameter.
10. A classification model training method based on big data distributed computing according to claim 7, characterized in that: In S4, the synchronous distribution of the integrated global model parameters uses a broadcast mechanism to uniformly send the final aggregation result to all distributed computing nodes at the same timestamp, ensuring the consistency and synchronization of parameter updates.
Citation Information
Patent Citations
Distributed training micro-batch data determination method and device, equipment and medium
CN118709752A
Data synchronization method and system for communication software
CN119835280A
Strong-adaptation distributed data distribution method supporting dynamic expansion
CN119960991A
Model training-based communication method and apparatus, and system
US20220360539A1
Cited By
Large language model GPT-2 remote training method in storage and calculation separation scene
CN120806176A
Method for carrying out model training based on k8s
CN120909809A
Method for model training based on k8s
CN120909809B
Heterogeneous environment-oriented asynchronous batch data parallel training method
CN122220121A