Data processing method, device, electronic device and storage medium
By optimizing the order of training samples, the computational and communication void problems caused by uneven sample lengths in distributed model training are solved, thereby improving hardware utilization and training efficiency.
Patent Information
- Application Number
- CN202510983254.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-17
AI Technical Summary
In distributed model training, the different lengths of training samples lead to uneven processing of computing devices, resulting in computing bubbles and communication bubbles, which reduces hardware utilization and overall training efficiency.
By obtaining the order of training samples, determining the cluster performance loss and model effect loss, and using sample reorganization to update the training sample order, the comprehensive loss is minimized, ensuring the model training effect while reducing computing/communication cavitation problems.
It improves hardware utilization and overall training efficiency, and is especially suitable for hybrid parallel distributed model training.
Smart Images

Figure CN120492177B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a data processing method, device, electronic device, and storage medium. Background Art
[0002] With the rapid development of artificial intelligence (AI), large-scale deep learning models have achieved remarkable results in tasks such as natural language processing. Currently, large-scale deep learning models are typically trained using distributed model training, which primarily includes data parallelism, pipeline parallelism, and hybrid parallelism (data parallelism + pipeline parallelism).
[0003] In distributed model training, due to the varying lengths of training samples, multiple shorter samples are often combined (packaged) into a single packaged sample with a length close to the target maximum length. However, this sample packaging method results in different effective lengths (number of tokens) of the packaged samples processed by different devices (such as GPUs) in the cluster executing distributed model training. This causes faster devices to wait for slower devices to complete their own computations after completing their own computations, resulting in idle device resources (referred to as computation bubbles). Furthermore, while waiting for all devices to complete computations for collective communication (such as gradient synchronization), idle waiting time on the communication link or devices (referred to as communication bubbles) is also incurred. These computational and communication bubbles reduce hardware utilization and overall training efficiency. Summary of the Invention
[0004] The present disclosure provides a data processing method, apparatus, electronic device, and storage medium to at least address the related art issue of low hardware utilization and overall training efficiency in distributed model training due to the packaging of training samples of different lengths. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a data processing method, including:
[0006] Obtaining a training sample sequence for distributed model training, the training sample sequence including a plurality of training sample sets corresponding one-to-one to a plurality of iterations, the distributed model training including training of the plurality of iterations;
[0007] Determining a cluster performance loss for executing the distributed model training based on computing power consumption of the training samples corresponding to each iteration in the training sample sequence;
[0008] Predicting the model effect loss of a model obtained by sequentially executing the distributed model training based on the training samples;
[0009] A comprehensive loss is determined based on the cluster performance loss and the model effect loss. With the goal of minimizing the comprehensive loss, the training sample sequence is updated by reorganizing samples between training sample sets of different iterations until a preset end condition is met, thereby obtaining a target training sample sequence; the target training sample sequence is used to perform the distributed model training on the model in the computing device cluster.
[0010] In an exemplary embodiment, determining the cluster performance loss of executing the distributed model training based on the computing power consumption of the training samples corresponding to each iteration in the training sample sequence includes:
[0011] For each of the iterations, determining a load ratio of the iteration based on the computing power consumption of the training samples corresponding to the iteration in the training sample sequence;
[0012] determining a relative deviation of the load ratio from a target load ratio for each of said iterations;
[0013] The relative deviations corresponding to the iterations are accumulated to obtain a cluster performance loss for executing the distributed model training.
[0014] In an exemplary embodiment, the predicting the model effect loss of the model obtained by sequentially executing the distributed model training based on the training samples includes:
[0015] Obtaining reference model parameters and reference model loss for each iteration; the reference model parameters are obtained by successfully performing training on the corresponding iteration based on the reference training samples, and the reference model loss is used to characterize the difference between the predicted value of the validation sample in the validation sample set based on the reference model parameters and the corresponding true label value;
[0016] Based on the reference model parameters of each iteration, parameter trajectory prediction is performed on the training sample sequence to obtain prediction model parameters corresponding to the training sample sequence in each iteration;
[0017] Determining a prediction model loss for each iteration based on the prediction model parameters of each iteration; the prediction model loss is used to characterize the difference between a predicted value of a validation sample in the validation sample set based on the prediction model parameters and a corresponding true label value;
[0018] The model effect loss is determined based on the difference between the prediction model loss and the reference model loss of each iteration.
[0019] In an exemplary embodiment, determining the model effect loss based on the difference between the prediction model loss and the reference model loss of each iteration includes:
[0020] Determining a confidence level corresponding to each iteration based on a deviation between the prediction model parameters of each iteration and the reference model parameters of the iteration, wherein the confidence level represents a degree of credibility of the prediction model parameters of the corresponding iteration;
[0021] Based on the confidence of each iteration, the difference between the prediction model loss and the reference model loss of each iteration is accumulated to obtain the model effect loss.
[0022] In an exemplary embodiment, determining the comprehensive loss based on the cluster performance loss and the model effect loss includes:
[0023] Determine the fusion coefficient according to the application scenario of the model;
[0024] The cluster performance loss and the model effect loss are linearly fused based on the fusion coefficient to obtain the comprehensive loss.
[0025] In an exemplary embodiment, updating the training sample sequence includes:
[0026] Traversing the plurality of iterations, and determining, for a current iteration traversed, a load of the current iteration based on computing power consumption of training samples corresponding to the current iteration in the training sample sequence;
[0027] Determining a relative load deviation based on a difference between the load of the current iteration and the target load; determining a global batch adjustment coefficient based on a product of the relative load deviation and an adjustment rate coefficient;
[0028] Adjusting the number of training samples corresponding to the current iteration in the training sample sequence based on the global batch adjustment coefficient to obtain global batch information for the next iteration, where the global batch information represents the number of training samples corresponding to the next iteration;
[0029] Based on the global batch information of the next iteration, sample reorganization is performed between training sample sets of different iterations until the next iteration is the last iteration of the multiple iterations.
[0030] In an exemplary embodiment, adjusting the number of training samples corresponding to the current iteration in the training sample sequence based on the global batch adjustment coefficient to obtain global batch information for the next iteration includes:
[0031] Determine the product of the global batch adjustment coefficient and the number of training samples corresponding to the current iteration in the training sample sequence to obtain candidate global batch information for the next iteration;
[0032] In a case where the number indicated by the candidate global batch information is less than the number indicated by the basic global batch information, using the basic global batch information as the global batch information of the next iteration;
[0033] In a case where the number indicated by the candidate global batch information is greater than or equal to the number indicated by the basic global batch information, the candidate global batch information is used as the global batch information of the next iteration.
[0034] In an exemplary embodiment, the distributed model training includes multiple training cycles, each of the training cycles includes the multiple iterations; and obtaining the reference model parameters and the reference model loss of each iteration includes:
[0035] Obtaining reference model parameters and reference model losses for each iteration obtained by sequentially performing corresponding iterations based on the reference training samples in a training cycle previous to the current training cycle;
[0036] Wherein, if the previous training cycle is the first training cycle among the multiple training cycles, the reference training sample order in the first training cycle is obtained by packaging samples based on a greedy bucketing method under the KL divergence constraint;
[0037] If the previous training cycle is any training cycle after the first training cycle, the reference training sample sequence in any training cycle is the target training sample sequence corresponding to the any training cycle.
[0038] According to a second aspect of an embodiment of the present disclosure, there is provided a data processing apparatus, including:
[0039] A sample sequence acquisition unit is configured to execute acquisition of a training sample sequence for distributed model training, wherein the training sample sequence includes a plurality of training sample sets corresponding one-to-one to a plurality of iterations, and the distributed model training includes training of the plurality of iterations;
[0040] a cluster performance loss determining unit configured to determine a cluster performance loss for executing the distributed model training based on computing power consumption of training samples corresponding to each iteration in the training sample sequence;
[0041] A model effect loss prediction unit is configured to predict the model effect loss of a model obtained by sequentially executing the distributed model training based on the training samples;
[0042] The target sample sequence determination unit is configured to determine the comprehensive loss based on the cluster performance loss and the model effect loss, with the goal of minimizing the comprehensive loss. The training sample sequence is updated by reorganizing samples between training sample sets of different iterations until a preset end condition is met, thereby obtaining a target training sample sequence; the target training sample sequence is used to execute the distributed model training.
[0043] In an exemplary embodiment, the cluster performance loss determination unit is specifically configured to perform: for each of the iterations, based on the computing power consumption of the training samples corresponding to the iteration in the training sample sequence, determine the load ratio of the iteration; determine the relative deviation of the load ratio of each iteration relative to the target load ratio; accumulate the relative deviations corresponding to each of the iterations to obtain the cluster performance loss of executing the distributed model training.
[0044] In an exemplary embodiment, the model effect loss prediction unit includes:
[0045] a reference data acquisition unit configured to acquire reference model parameters and reference model loss for each iteration; the reference model parameters being obtained by successfully executing the corresponding iteration of training based on the reference training samples, and the reference model loss being used to characterize the difference between the predicted value of the validation sample in the validation sample set based on the reference model parameters and the corresponding true label value;
[0046] a parameter trajectory prediction unit configured to perform parameter trajectory prediction on the training sample sequence based on the reference model parameters of each iteration, and obtain prediction model parameters corresponding to the training sample sequence in each iteration;
[0047] A prediction model loss determination unit is configured to execute prediction model parameters based on each of the iterations to determine the prediction model loss of each of the iterations; the prediction model loss is used to represent the difference between the predicted value of the validation sample in the validation sample set obtained based on the prediction model parameters and the corresponding true label value;
[0048] The model effect loss determining unit is configured to determine the model effect loss based on the difference between the prediction model loss and the reference model loss of each iteration.
[0049] In an exemplary embodiment, the model effect loss determination unit includes:
[0050] a confidence determination unit configured to determine a confidence corresponding to each iteration based on a deviation between the prediction model parameters of each iteration and the reference model parameters of the iteration, wherein the confidence represents a degree of credibility of the prediction model parameters of the corresponding iteration;
[0051] The model effect loss determination subunit is configured to perform an accumulation of the difference between the prediction model loss and the reference model loss of each iteration based on the confidence of each iteration to obtain the model effect loss.
[0052] In an exemplary embodiment, the target sample sequence determination unit, when determining the comprehensive loss based on the cluster performance loss and the model effect loss, is specifically configured to determine a fusion coefficient according to the application scenario of the model; and linearly fuse the cluster performance loss and the model effect loss based on the fusion coefficient to obtain the comprehensive loss.
[0053] In an exemplary embodiment, the target sample sequence determining unit is specifically configured to perform, when updating the training sample sequence: traversing the multiple iterations, and determining, for a current iteration traversed, a load of the current iteration based on computing power consumption of training samples corresponding to the current iteration in the training sample sequence;
[0054] Determining a relative load deviation based on a difference between the load of the current iteration and the target load; determining a global batch adjustment coefficient based on a product of the relative load deviation and an adjustment rate coefficient;
[0055] Adjusting the number of training samples corresponding to the current iteration in the training sample sequence based on the global batch adjustment coefficient to obtain global batch information for the next iteration, where the global batch information represents the number of training samples corresponding to the next iteration;
[0056] Based on the global batch information of the next iteration, sample reorganization is performed between training sample sets of different iterations until the next iteration is the last iteration of the multiple iterations.
[0057] In an exemplary embodiment, the target sample sequence determining unit, when adjusting the number of training samples corresponding to the current iteration in the training sample sequence based on the global batch adjustment coefficient to obtain the global batch information of the next iteration, is specifically configured to perform: determining the product of the global batch adjustment coefficient and the number of training samples corresponding to the current iteration in the training sample sequence to obtain the candidate global batch information of the next iteration;
[0058] In a case where the number indicated by the candidate global batch information is less than the number indicated by the basic global batch information, using the basic global batch information as the global batch information of the next iteration;
[0059] In a case where the number indicated by the candidate global batch information is greater than or equal to the number indicated by the basic global batch information, the candidate global batch information is used as the global batch information of the next iteration.
[0060] In an exemplary embodiment, the distributed model training includes multiple training cycles, each of which includes the multiple iterations; the reference data acquisition unit is specifically configured to perform: obtaining reference model parameters and reference model losses of each iteration obtained by sequentially executing corresponding iterations based on the reference training samples in a training cycle before the current training cycle;
[0061] Among them, if the previous training cycle is the first training cycle among the multiple training cycles, the reference training sample order in the first training cycle is obtained by sample packaging based on the greedy bucketing method while satisfying the KL divergence constraint; if the previous training cycle is any training cycle after the first training cycle, the reference training sample order in any training cycle is the target training sample order corresponding to any training cycle.
[0062] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:
[0063] processor;
[0064] a memory for storing instructions executable by the processor;
[0065] The processor is configured to execute the instructions to implement the data processing method in the first aspect above.
[0066] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the data processing method in the first aspect above.
[0067] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the data processing method in the first aspect is implemented.
[0068] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:
[0069] By obtaining a training sample sequence for distributed model training, the training sample sequence includes multiple training sample sets corresponding to multiple iterations one by one, and then based on the computing power consumption of the training samples corresponding to each iteration in the training sample sequence, the cluster performance loss of executing the distributed model training is determined, and the model effect loss of the model obtained by executing the distributed model training based on the training sample sequence is predicted. The comprehensive loss is determined based on the cluster performance loss and the model effect loss, and with the goal of minimizing the comprehensive loss, the training sample sequence is updated by reorganizing samples between training sample sets of different iterations until the preset end condition is met, and the target training sample sequence actually used to execute the distributed model training is obtained, thereby ensuring the model training effect while minimizing the computing / communication bubble problem caused by the packaging of training samples of different lengths, improving hardware utilization and overall training efficiency, and being particularly suitable for hybrid parallel distributed model training.
[0070] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0072] Figure 1 is a schematic diagram of an implementation environment of a data processing method according to an exemplary embodiment;
[0073] Figure 2 is a flowchart illustrating a data processing method according to an exemplary embodiment;
[0074] Figure 3 is a flowchart illustrating another data processing method according to an exemplary embodiment;
[0075] Figure 4 is a flowchart illustrating another data processing method according to an exemplary embodiment;
[0076] Figure 5 is a structural block diagram of a data processing device according to an exemplary embodiment;
[0077] Figure 6 is a block diagram showing an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0078] In order to enable ordinary people in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0079] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0080] See also Figure 1 , which is a schematic diagram of an implementation environment of a data processing method according to an exemplary embodiment. The implementation environment may include a computing device cluster 110 and a data processing server 120.
[0081] Computing device cluster 110 includes multiple computing devices used to perform distributed model training of large deep learning models. When training large deep learning models, these multiple computing devices form a large computing network. Training data can be exchanged between multiple devices in this computing network. After the training data is exchanged to the computing devices, the computing devices will perform the model training calculations, and the intermediate results will be exchanged as training data to other computing devices. Distributed model training can include data parallelism (DP), pipeline parallelism (PP), and hybrid parallelism (DP+PP). Large deep learning models can include large language models (LLM).
[0082] Among them, computing devices have the characteristic of large computing power. Computing devices may include but are not limited to graphics processing units (GPUs), neural-network process units (NPUs), tensor processing units (TPUs), FPGA (Field Programmable Gate Array) hardware, etc.
[0083] The data processing server 120 is used to obtain a target training sample sequence for performing distributed model training on a large deep learning model based on the data processing method of an embodiment of the present disclosure, and use the target training sample sequence to perform distributed model training on the model in the computing device cluster 110, thereby minimizing the computing / communication cavitation problem caused by the packaging of training samples of different lengths while ensuring the model training effect, thereby improving hardware utilization and overall training efficiency, and is particularly suitable for hybrid parallel distributed model training.
[0084] It should be noted that the server involved in the embodiments of the present disclosure can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0085] Figure 2 is a flow chart showing a data processing method according to an exemplary embodiment. Figure 2 As shown, the method may include the following steps S201 to S207:
[0086] In step 201, a training sample sequence for distributed model training is obtained, where the training sample sequence includes a plurality of training sample sets corresponding one-to-one to a plurality of iterations, and the distributed model training includes training of the plurality of iterations.
[0087] In a specific implementation, the initial training sample order can be obtained by randomly packaging the samples in the training dataset based on the parallel configuration parameters. Taking the distributed model training as hybrid parallel as an example, the parallel configuration parameters can include the number of data parallel groups dp_size (i.e., the number of model copies trained simultaneously) and the number of micro-batches num_micro_bs (i.e., the minimum computing unit split in pipeline parallelism). Based on the product of dp_size and num_micro_bs, the total number of micro-batches corresponding to each iteration can be determined. Based on the total number of micro-batches corresponding to each iteration, the samples in the training dataset are packaged to obtain the training sample set corresponding to each iteration. For example, suppose there is a training dataset containing T iterations, denoted as , where B t represents the training sample set of the tth iteration, then the initial training sample order is {B0, B2,…, B T-1}.
[0088] The training samples may be multimodal data, which may be one or a combination of documents, pictures, and videos.
[0089] In step 203, based on the computing power consumption of the training samples corresponding to each iteration in the training sample sequence, the cluster performance loss of executing the distributed model training is determined.
[0090] The computing power consumption of the training samples in the training sample set corresponding to each iteration can represent the number of machine resources consumed when calculating the training samples. It can be quantified as the number of cores (central processing units, CPUs) consumed, the number of floating-point operations per second (FLOPS) consumed, etc. Taking FLOPS as an example to represent the computing power consumption of training samples, the FLOPS of a training sample can be calculated by multiplying the square of the sequence length by the number of model parameters.
[0091] The cluster performance loss represents the severity of the load imbalance caused by the order of training samples during multiple iterations (such as T iterations). Generally, the greater the cluster performance loss, the more severe the load imbalance caused by the order of training samples.
[0092] In some exemplary embodiments, the cluster performance loss can be quantified by the cumulative cost of the load ratio of the training sample sequence deviating from the target load ratio during multiple iterations. Figure 3 As shown, the above step 203, when determining the cluster performance loss of executing the distributed model training based on the computing power consumption of the training samples corresponding to each iteration in the training sample sequence, may include:
[0093] In step S301, for each iteration, a load ratio of the iteration is determined based on the computing power consumption of the training samples corresponding to the iteration in the training sample sequence.
[0094] Among them, the load ratio of the t-th iteration represents the degree of load imbalance in the training sample set corresponding to the t-th iteration. When the load ratio is close to 1, it indicates that the load in the training sample set is highly balanced. When the load ratio is greater than 1, it indicates that the load in the training sample set is unbalanced. The larger the load ratio, the higher the degree of imbalance.
[0095] For example, under the training sample order, the load ratio of the t-th iteration can be obtained by the following formula (1):
[0096] (1)
[0097] in, Indicates the order of training samples Next, the load ratio of the tth iteration; represents the training sample set corresponding to the t-th iteration, that is, the training sample set processed in the t-th iteration; Indicates the size of the set, that is, the number of training samples in the set; Represents training samples The computing power consumption; express The maximum computational power consumption of all training samples in reflects the heaviest load in the tth iteration; express The arithmetic mean of the computing power consumption of all training samples in reflects the average load level of the t-th iteration.
[0098] In step S303 , a relative deviation of the load ratio of each iteration relative to the target load ratio is determined.
[0099] The target load ratio represents the desired load balancing state of the cluster and can usually be set to a value equal to 1, such as 0.9.
[0100] In step S305, the relative deviations corresponding to the iterations are accumulated to obtain the cluster performance loss of executing the distributed model training.
[0101] Specifically, only items with relative deviations greater than 0 are considered when accumulating relative deviations. At the same time, the impact of large relative deviations (such as a load ratio far exceeding the target load ratio) on cluster performance loss can be amplified.
[0102] In practical applications, the importance of different iterations can also be considered, and a performance weight coefficient can be assigned to each iteration. For example, in critical iterations (such as in the late stages of training), load balancing may be more desirable. Therefore, a higher performance weight coefficient can be assigned to the critical iteration to enhance its contribution to cluster performance loss.
[0103] For example, the clustering performance loss under the training sample order can be obtained by the following formula (2):
[0104] (2)
[0105] in, Indicates the order of training samples Cluster performance loss under Indicates the order of training samples Next, the load ratio of the tth iteration; represents the target load ratio; represents the relative deviation of the load ratio of the tth iteration relative to the target load ratio; Indicates that only the cases exceeding the target load ratio are considered; represents the performance weight coefficient of the t-th iteration; T represents T iterations.
[0106] In the above embodiment, the training sample sequence is Accurately quantify the order of training samples by the cumulative cost of the load ratio deviating from the target load ratio during multiple iterations The cluster performance loss under this condition provides an accurate basis for the subsequent comprehensive consideration of system overhead.
[0107] In step 205, the model effect loss of the model obtained by sequentially executing the distributed model training based on the training samples is predicted.
[0108] Among them, the model effect loss is used to predict the effect of different training sample orders on the model effect without actually training the model. The model effect is used to measure the performance of the model, such as the accuracy of the model on a specific task.
[0109] In some exemplary embodiments, Figure 4 As shown, the above step S205 may include:
[0110] In step S401, the reference model parameters and reference model loss of each iteration are obtained; the reference model parameters are obtained by successfully executing the training of the corresponding iteration based on the reference training samples, and the reference model loss is used to characterize the difference between the predicted value obtained by the verification sample in the verification sample set based on the reference model parameters and the corresponding true label value.
[0111] The reference model parameters for each iteration refer to the model parameters at the end of that iteration under the reference training sample sequence, which can be obtained by performing actual distributed training on the model using the reference training sample sequence. Exemplarily, the reference training sample sequence can be a random training sample sequence.
[0112] The reference model loss of each iteration is used to characterize the difference between the predicted value of the validation sample in the validation sample set based on the reference model parameters of that iteration and the corresponding true label value.
[0113] by represents the validation sample set, Indicates the order of reference training samples The model parameters at the end of the tth iteration are the reference model parameters of the tth iteration. Indicates that the model has parameters The loss value of the validation sample (x, y) when , represents the predicted value, y represents the corresponding true label value, then the reference model loss of the tth iteration can be expressed as , the corresponding loss function can be cross entropy, mean square error, etc.
[0114] In step S403, parameter trajectory prediction is performed on the training sample sequence based on the reference model parameters of each iteration to obtain prediction model parameters corresponding to the training sample sequence of each iteration.
[0115] The prediction model parameters corresponding to the training sample sequence of each iteration refer to the model parameters at the end of the iteration under the training sample sequence.
[0116] Specifically, the parameter trajectory prediction based on Talor expansion can be used, as shown in the following formula (3):
[0117] (3)
[0118] in, ; Indicates the order of reference training samples Next, the reference model parameters of the t+1th iteration; Indicates the order of training samples Next, the prediction model parameters of the t+1th iteration; Indicates the order of training samples The predicted model parameters at the tth iteration are: K represents the order of the Talor expansion. Considering the computational complexity, K can be set to 2. The above parameter trajectory prediction based on the Talor expansion can predict the impact of different training sample orders on model parameters without actually training the model.
[0119] In step S405, based on the prediction model parameters of each iteration, the prediction model loss of each iteration is determined; the prediction model loss is used to characterize the difference between the predicted value of the verification sample in the verification sample set based on the prediction model parameters and the corresponding true label value.
[0120] In step S407, the model effect loss is determined based on the difference between the prediction model loss and the reference model loss of each iteration.
[0121] Specifically, for the tth iteration, the difference between the prediction model loss and the reference model loss of this iteration can be expressed as: ,in, Indicates the order of training samples Next, the prediction model loss corresponding to the validation sample (x, y) in the tth iteration; Indicates the single sample loss difference on the validation sample set in the tth iteration. The model effect loss can be determined based on the single sample loss difference on the validation sample set in the t-th iteration, and the average loss difference on the validation sample set in the t-th iteration can be accumulated.
[0122] In practical applications, the importance of different iterations can also be considered, and an effect weight coefficient can be assigned to each iteration. For example, in key iterations (such as in the later stages of training), a higher effect weight coefficient can be assigned to the key iteration to enhance the contribution of the key iteration to the model effect loss.
[0123] For example, in the training sample order The model performance loss under can be obtained by the following formula (4):
[0124] (4)
[0125] in, Indicates the order of training samples Model effect loss under ; represents the effect weight coefficient of the t-th iteration; T represents T iterations. E is the expectation, which means averaging the loss difference of all validation samples on the validation sample set to avoid the influence of noise of individual validation samples.
[0126] The above implementation method predicts the model parameters under the new training sample sequence based on the reference model parameters of each iteration, and determines the model effect loss under the new training sample sequence in combination with the reference model loss based on each iteration, thereby achieving accurate estimation of the model effect loss without actual training.
[0127] In some exemplary embodiments, in order to improve the accuracy of parameter trajectory prediction and to improve the accuracy of the target training sample sequence finally optimized, the above-mentioned step S407 may include, when implemented: determining the confidence corresponding to each iteration based on the deviation between the prediction model parameters of each iteration and the reference model parameters of the iteration, the confidence representing the credibility of the prediction model parameters of the corresponding iteration; based on the confidence of each iteration, accumulating the difference between the prediction model loss of each iteration and the reference model loss to obtain the model effect loss.
[0128] In the specific implementation, under the training sample order, the confidence corresponding to each iteration can be obtained using the following formula (5):
[0129] (5)
[0130] in, Indicates the order of training samples Next, the confidence level of the t-th iteration; Represents a hyperparameter used to control sensitivity.
[0131] Then in the training sample order Under this condition, the model effect loss after confidence adjustment can be obtained by the following formula (6):
[0132] (6)
[0133] The above implementation introduces the confidence of each iteration as the coefficient of the loss term corresponding to the corresponding iteration, so that when the prediction model parameters are unreliable (the confidence is approximately zero), the loss difference of this iteration can be ignored. When the prediction model parameters are reliable (the confidence is approximately 1), the loss difference of this iteration is retained. This can improve the stability and accuracy of the sequential evaluation of training samples.
[0134] In step 207, a comprehensive loss is determined based on the cluster performance loss and the model effect loss. With the goal of minimizing the comprehensive loss, the training sample sequence is updated by reorganizing samples between training sample sets of different iterations until a preset end condition is met, thereby obtaining a target training sample sequence.
[0135] The target training sample sequence is used to perform the distributed model training on the model in the computing device cluster.
[0136] Specifically, a weight coefficient can be assigned to the cluster performance loss and the model effect loss, and the weighted sum of the two can be calculated based on the assigned weight coefficient to obtain the comprehensive loss. For example, the comprehensive loss can be expressed as:
[0137] (7)
[0138] in, Represents the weight coefficient, which can be set based on actual needs; Indicates the order of training samples The comprehensive loss under .
[0139] In some exemplary embodiments, to improve the adaptability of the target training sample sequence to the application scenario, determining the comprehensive loss based on the cluster performance loss and the model effect loss may include: determining a fusion coefficient based on the application scenario of the model; and linearly fusing the cluster performance loss and the model effect loss based on the fusion coefficient to obtain the comprehensive loss. Specifically, the comprehensive loss may be obtained by combining the fusion coefficient using the following formula (8):
[0140] (8)
[0141] in, Represents the fusion coefficient, which can be determined based on the application scenario of the model. For example, for high-throughput scenarios (such as recommendation services), cluster performance is usually prioritized, so the fusion coefficient It can be set to greater than 0.5, for example, 0.8. For high-quality scenarios (such as model training scenarios), the model effect is usually prioritized, so the fusion coefficient It can be set to less than 0.5, such as 0.3; for balanced scenes, it can be set The value is set to 0.5 to achieve a trade-off between the two.
[0142] In the disclosed embodiments, sample reorganization between training sample sets from different iterations can achieve global sample reorganization across iterations. Compared to sample reorganization within a single iteration, this can achieve global load balancing, minimize resource consumption caused by packaging samples of different lengths, and improve hardware utilization. In specific implementations, sample sequences can be exchanged or reallocated between adjacent iterations or neighboring iterations. The preset termination condition can be that the comprehensive loss reaches a preset minimum threshold.
[0143] In an exemplary embodiment, in order to further improve the global load optimization effect across iterations while taking into account both cluster performance and model effect, thereby improving hardware utilization and overall training efficiency, the above step S207 may include the following steps when updating the training sample order:
[0144] Traversing the plurality of iterations, and determining, for a current iteration traversed, a load of the current iteration based on computing power consumption of training samples corresponding to the current iteration in the training sample sequence;
[0145] Determining a relative load deviation based on a difference between the load of the current iteration and the target load; determining a global batch adjustment coefficient based on a product of the relative load deviation and an adjustment rate coefficient;
[0146] Adjusting the number of training samples corresponding to the current iteration in the training sample sequence based on the global batch adjustment coefficient to obtain global batch information for the next iteration, where the global batch information represents the number of training samples corresponding to the next iteration;
[0147] Based on the global batch information of the next iteration, sample reorganization is performed between training sample sets of different iterations until the next iteration is the last iteration of the multiple iterations.
[0148] In the specific implementation, the following formula (9) can be used to dynamically adjust the global batch size (that is, the number of training samples in the training sample set corresponding to each iteration):
[0149] (9)
[0150] in, Indicates the global batch size used in the current iteration t, that is, the number of training samples in the training sample set corresponding to the current iteration t; Indicates the load of the current iteration t, which can be the maximum computing power consumption of the training sample for the current iteration t; Indicates the target load, that is, the load expected to be achieved in a single iteration, which can be set based on the actual cluster situation; is the relative load deviation, which indicates the ratio of the difference between the load of the current iteration and the target load relative to the target load; Represents the adjustment rate coefficient, which is a hyperparameter greater than 0 and is used to control the magnitude and aggressiveness of global batch size adjustment. The larger the value, the more aggressive the adjustment (i.e., faster response, but may be unstable), and the smaller the value, the more gradual the adjustment (i.e., more stable, but may respond slower). Represents the global batch adjustment coefficient. If it is greater than 1, it means that the global batch needs to be increased. If it is less than 1, it means that the global batch needs to be reduced. If it is equal to 1, it means that the global batch remains unchanged. It represents the global batch size of the next iteration t+1, that is, the number of training samples in the training sample set corresponding to the t+1th iteration.
[0151] The above implementation dynamically adjusts the global batch size in distributed training based on the load differences between iterations, thereby stabilizing the load of a single iteration near the target load, maximizing the utilization of cluster computing resources (avoiding idleness or overload), and thus accelerating the entire training process.
[0152] In some exemplary embodiments, adjusting the number of training samples corresponding to the current iteration in the training sample sequence based on the global batch adjustment coefficient to obtain the global batch information for the next iteration may include: determining the product of the global batch adjustment coefficient and the number of training samples corresponding to the current iteration in the training sample sequence to obtain candidate global batch information for the next iteration; if the number indicated by the candidate global batch information is less than the number indicated by the basic global batch information, using the basic global batch information as the global batch information for the next iteration; if the number indicated by the candidate global batch information is greater than or equal to the number indicated by the basic global batch information, using the candidate global batch information as the global batch information for the next iteration.
[0153] Specifically, the number of basic global batch information indications can be understood as the set minimum global batch size, and the global batch information of the next iteration can be obtained by the following formula (10):
[0154] B t+1=max( B min , )(10)
[0155] Wherein, Bmin represents the number of basic global batch information indications; Indicates the number of candidate global batch information indicators.
[0156] The above implementation method coordinates the basic share with the dynamic share to achieve dynamic adjustment of the global batch size while ensuring that the global batch size of each iteration will not fall below the basic level that guarantees training stability and minimum data throughput.
[0157] In some exemplary embodiments, the distributed model training may include multiple training cycles (epochs), each training cycle including the aforementioned multiple iterations (iters), and the aforementioned step S401, when obtaining the reference model parameters and reference model loss of each iteration, may be: obtaining the reference model parameters and reference model loss of each iteration obtained by sequentially executing corresponding iterations based on the reference training samples in a training cycle before the current training cycle;
[0158] Wherein, if the previous training cycle is the first training cycle among the multiple training cycles, the reference training sample order in the first training cycle is obtained by packaging samples based on a greedy bucketing method under the KL divergence constraint;
[0159] If the previous training cycle is any training cycle after the first training cycle, the reference training sample sequence in any training cycle is the target training sample sequence corresponding to the any training cycle.
[0160] In the specific implementation, for the first training cycle (Epoch0), the input is the original training dataset and the parallel configuration (such as DP_size, PP_size, micro_bs). Then, in Epoch0, greedy bucketing is performed under the KL divergence constraint to obtain the corresponding training sample order. The greedy bucketing process is as follows: for samples of multiple consecutive iterations in the original training dataset, sort them in descending order according to the computing power consumption of the samples; initialize num_buckets = dp_size × micro_bs buckets, select the bucket with the lowest current cumulative computing power consumption each time, put the next sample in, and use the above formula (1) to perform load balancing constraints so that the load ratio of each iteration is less than the target load ratio (such as 0.9). Considering that changes in sample distribution will affect the model training effect, Epoch0 uses the KL divergence constraint, specifically: divide the original training dataset into K groups according to the length / modality / category of the samples to obtain the original data distribution p0; for each iteration obtained based on greedy bucketing, determine the data distribution p of the iteration, if KL(p||p0)> (It can be set based on actual needs, for example, set to 0.05), then the samples are exchanged with the subsequent iterations of this iteration until KL(p||p0) is less than or equal to . Perform complete Epoch0 training based on the training samples corresponding to Epoch0 above, and record the reference model parameters of each iteration _t_ref (t=1,2,...,T) and reference model loss L _t_ref (such as cross entropy loss) to get the reference data for the second training cycle.
[0161] For the second training cycle (Epoch 1), the input is the reference data obtained in Epoch 0, that is, the reference model parameters recorded for each iteration in Epoch 0 training. _t_ref (t=1,2,...,T) and reference model loss L _t_ref (such as cross entropy loss), and the model parameters at the end of Epoch0, and then use the data processing method of the embodiment of the present disclosure to obtain the target training sample sequence corresponding to Epoch1. Perform the complete Epoch1 training based on the target training sample sequence corresponding to Epoch1, and record the reference model parameters of each iteration _t_ref (t=1,2,...,T) and reference model loss L _t_ref (such as cross entropy loss) to get the reference data for the third training cycle.
[0162] For subsequent training cycles (Epoch K ≥ 2), the input is the reference data obtained in the previous training cycle, that is, the reference model parameters recorded for each iteration in Epoch K-1 training. _t_ref (t=1,2,...,T) and reference model loss L _t_ref (such as cross entropy loss), and the model parameters at the end of EpochK-1, and then use the data processing method of the embodiment of the present disclosure to obtain the target training sample sequence corresponding to EpochK. Perform the complete EpochK training based on the target training sample sequence corresponding to EpochK, and record the reference model parameters of each iteration _t_ref (t=1,2,...,T) and reference model loss L _t_ref (such as cross entropy loss) to obtain the reference data for the K+1th training cycle.
[0163] The above implementation ensures the stability of distributed model training by implementing a greedy bucketing order based on the KL divergence constraint in the first epoch. Subsequent epochs are gradually optimized based on the previous epoch. This multi-epoch progressive optimization mechanism not only avoids training crashes caused by aggressive sample reshuffling, but also improves the accuracy of parameter trajectory prediction as reference data accumulates, thereby significantly improving cluster performance while ensuring model training effectiveness.
[0164] Figure 5 FIG. 1 is a block diagram of a data processing device according to an exemplary embodiment. Figure 5 , the data processing device 500 includes:
[0165] A sample sequence acquisition unit 510 is configured to execute acquisition of a training sample sequence for distributed model training, wherein the training sample sequence includes a plurality of training sample sets corresponding one-to-one to a plurality of iterations, and the distributed model training includes training of the plurality of iterations;
[0166] A cluster performance loss determining unit 520 is configured to determine a cluster performance loss for executing the distributed model training based on the computing power consumption of the training samples corresponding to each iteration in the training sample sequence;
[0167] A model effect loss prediction unit 530 is configured to predict the model effect loss of a model obtained by sequentially executing the distributed model training based on the training samples;
[0168] The target sample sequence determination unit 540 is configured to determine the comprehensive loss based on the cluster performance loss and the model effect loss, with the goal of minimizing the comprehensive loss. The training sample sequence is updated by reorganizing samples between training sample sets of different iterations until a preset end condition is met, thereby obtaining a target training sample sequence; the target training sample sequence is used to execute the distributed model training.
[0169] In an exemplary embodiment, the cluster performance loss determination unit 520 is specifically configured to perform: for each of the iterations, based on the computing power consumption of the training samples corresponding to the iteration in the training sample sequence, determine the load ratio of the iteration; determine the relative deviation of the load ratio of each iteration relative to the target load ratio; accumulate the relative deviations corresponding to each of the iterations to obtain the cluster performance loss of executing the distributed model training.
[0170] In an exemplary embodiment, the model effect loss prediction unit 530 includes:
[0171] a reference data acquisition unit configured to acquire reference model parameters and reference model loss for each iteration; the reference model parameters being obtained by successfully executing the corresponding iteration of training based on the reference training samples, and the reference model loss being used to characterize the difference between the predicted value of the validation sample in the validation sample set based on the reference model parameters and the corresponding true label value;
[0172] a parameter trajectory prediction unit configured to perform parameter trajectory prediction on the training sample sequence based on the reference model parameters of each iteration, and obtain prediction model parameters corresponding to the training sample sequence in each iteration;
[0173] A prediction model loss determination unit is configured to execute prediction model parameters based on each of the iterations to determine the prediction model loss of each of the iterations; the prediction model loss is used to represent the difference between the predicted value of the validation sample in the validation sample set obtained based on the prediction model parameters and the corresponding true label value;
[0174] The model effect loss determining unit is configured to determine the model effect loss based on the difference between the prediction model loss and the reference model loss of each iteration.
[0175] In an exemplary embodiment, the model effect loss determination unit includes:
[0176] a confidence determination unit configured to determine a confidence corresponding to each iteration based on a deviation between the prediction model parameters of each iteration and the reference model parameters of the iteration, wherein the confidence represents a degree of credibility of the prediction model parameters of the corresponding iteration;
[0177] The model effect loss determination subunit is configured to perform an accumulation of the difference between the prediction model loss and the reference model loss of each iteration based on the confidence of each iteration to obtain the model effect loss.
[0178] In an exemplary embodiment, the target sample sequence determination unit 540, when determining the comprehensive loss based on the cluster performance loss and the model effect loss, is specifically configured to determine a fusion coefficient according to the application scenario of the model; and linearly fuse the cluster performance loss and the model effect loss based on the fusion coefficient to obtain the comprehensive loss.
[0179] In an exemplary embodiment, the target sample sequence determining unit 540 is specifically configured to perform, when updating the training sample sequence: traversing the plurality of iterations, and determining, for a current iteration traversed, a load of the current iteration based on computing power consumption of training samples corresponding to the current iteration in the training sample sequence;
[0180] Determining a relative load deviation based on a difference between the load of the current iteration and the target load; determining a global batch adjustment coefficient based on a product of the relative load deviation and an adjustment rate coefficient;
[0181] Adjusting the number of training samples corresponding to the current iteration in the training sample sequence based on the global batch adjustment coefficient to obtain global batch information for the next iteration, where the global batch information represents the number of training samples corresponding to the next iteration;
[0182] Based on the global batch information of the next iteration, sample reorganization is performed between training sample sets of different iterations until the next iteration is the last iteration of the multiple iterations.
[0183] In an exemplary embodiment, the target sample sequence determination unit 540, when adjusting the number of training samples corresponding to the current iteration in the training sample sequence based on the global batch adjustment coefficient to obtain the global batch information of the next iteration, is specifically configured to perform: determining the product of the global batch adjustment coefficient and the number of training samples corresponding to the current iteration in the training sample sequence to obtain candidate global batch information for the next iteration; if the number indicated by the candidate global batch information is less than the number indicated by the basic global batch information, using the basic global batch information as the global batch information for the next iteration; if the number indicated by the candidate global batch information is greater than or equal to the number indicated by the basic global batch information, using the candidate global batch information as the global batch information for the next iteration.
[0184] In an exemplary embodiment, the distributed model training includes multiple training cycles, each of which includes the multiple iterations; the reference data acquisition unit is specifically configured to perform: obtaining reference model parameters and reference model losses of each iteration obtained by sequentially executing corresponding iterations based on the reference training samples in a training cycle before the current training cycle;
[0185] Among them, if the previous training cycle is the first training cycle among the multiple training cycles, the reference training sample order in the first training cycle is obtained by sample packaging based on the greedy bucketing method while satisfying the KL divergence constraint; if the previous training cycle is any training cycle after the first training cycle, the reference training sample order in any training cycle is the target training sample order corresponding to any training cycle.
[0186] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0187] In an exemplary embodiment, an electronic device is also provided, including a processor; a memory for storing processor-executable instructions; wherein, when the processor is configured to execute the instructions stored in the memory, it implements the steps of any data processing method provided in the above embodiments.
[0188] The electronic device may be a terminal, a server or a similar computing device. For example, the computer may be running on a server. Figure 6 This is a hardware structure diagram of a server running a data processing method provided by an embodiment of the present invention, such as Figure 6As shown, the server 600 may vary significantly depending on its configuration or performance. It may include one or more central processing units (CPUs) 610 (processor 610 may include, but is not limited to, a microprocessor (MCU) or a processing device such as a programmable logic device (FPGA), a memory 630 for storing data, and one or more storage media 620 (e.g., one or more mass storage devices) for storing application programs 623 or data 622. The memory 630 and storage media 620 may be either transient or persistent storage. The program stored in the storage medium 620 may include one or more modules, each of which may include a series of instruction operations on the server. Furthermore, the CPU 610 may be configured to communicate with the storage medium 620 to execute the series of instruction operations in the storage medium 620 on the server 600. The server 600 may also include one or more power supplies 660, one or more wired or wireless network interfaces 650, one or more input and output interfaces 640, and / or one or more operating systems 621, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0189] The input / output interface 640 can be used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the server 600. In one embodiment, the input / output interface 640 may include a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the input / output interface 640 may be a radio frequency (RF) module for wireless communication with the Internet.
[0190] It can be understood by those skilled in the art that Figure 6 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 6 More or fewer components than shown, or with Figure 6 Different configurations shown.
[0191] In an exemplary embodiment, a computer-readable storage medium including instructions is further provided, such as a memory 630 including instructions. The instructions may be executed by the processor 610 of the apparatus 600 to implement the data processing method of the embodiment of the present disclosure. Alternatively, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, or the like.
[0192] In an exemplary embodiment, a computer program product is further provided, including a computer program, wherein when the computer program is executed by a processor, the data processing method provided by the embodiment of the present disclosure is implemented.
[0193] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0194] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A data processing method, characterized in that: include: Obtaining a training sample sequence for distributed model training, the training sample sequence including a plurality of training sample sets corresponding one-to-one to a plurality of iterations, the distributed model training including training of the plurality of iterations; Determining a cluster performance loss for executing the distributed model training based on computing power consumption of the training samples corresponding to each iteration in the training sample sequence, wherein the cluster performance loss represents a severity of load imbalance caused by the training sample sequence in the multiple iterations; Obtaining reference model parameters and reference model loss for each of the iterations; The reference model parameters are obtained by successfully performing corresponding iterative training based on the reference training samples, and the reference model loss is used to characterize the difference between the predicted value of the validation sample in the validation sample set based on the reference model parameters and the corresponding true label value; Based on the reference model parameters of each iteration, parameter trajectory prediction is performed on the training sample sequence to obtain prediction model parameters corresponding to the training sample sequence in each iteration; Determining a prediction model loss for each iteration based on the prediction model parameters of each iteration; the prediction model loss is used to characterize the difference between a predicted value of a validation sample in the validation sample set based on the prediction model parameters and a corresponding true label value; Determining a model performance loss based on a difference between the prediction model loss and the reference model loss for each of the iterations; A comprehensive loss is determined based on the cluster performance loss and the model effect loss. With the goal of minimizing the comprehensive loss, the training sample sequence is updated by reorganizing samples between training sample sets of different iterations until a preset end condition is met, thereby obtaining a target training sample sequence; the target training sample sequence is used to perform the distributed model training on the model in the computing device cluster.
2. The data processing method according to claim 1, wherein: The determining, based on the computing power consumption of the training samples corresponding to each iteration in the training sample sequence, the cluster performance loss of executing the distributed model training comprises: For each of the iterations, determining a load ratio of the iteration based on the computing power consumption of the training samples corresponding to the iteration in the training sample sequence; determining a relative deviation of the load ratio from a target load ratio for each of said iterations; The relative deviations corresponding to the iterations are accumulated to obtain a cluster performance loss for executing the distributed model training.
3. The data processing method according to claim 1, wherein: The determining of the model effect loss based on the difference between the prediction model loss and the reference model loss of each iteration includes: Determining a confidence level corresponding to each iteration based on a deviation between the prediction model parameters of each iteration and the reference model parameters of the iteration, wherein the confidence level represents a degree of credibility of the prediction model parameters of the corresponding iteration; Based on the confidence of each iteration, the difference between the prediction model loss and the reference model loss of each iteration is accumulated to obtain the model effect loss.
4. The data processing method according to claim 1, wherein: The determining of the comprehensive loss based on the cluster performance loss and the model effect loss includes: Determine the fusion coefficient according to the application scenario of the model; The cluster performance loss and the model effect loss are linearly fused based on the fusion coefficient to obtain the comprehensive loss.
5. The data processing method according to claim 1, wherein: The updating of the training sample sequence includes: Traversing the plurality of iterations, and determining, for a current iteration traversed, a load of the current iteration based on computing power consumption of training samples corresponding to the current iteration in the training sample sequence; Determining a relative load deviation based on a difference between the load of the current iteration and the target load; determining a global batch adjustment coefficient based on a product of the relative load deviation and an adjustment rate coefficient; Adjusting the number of training samples corresponding to the current iteration in the training sample sequence based on the global batch adjustment coefficient to obtain global batch information for the next iteration, where the global batch information represents the number of training samples corresponding to the next iteration; Based on the global batch information of the next iteration, sample reorganization is performed between training sample sets of different iterations until the next iteration is the last iteration of the multiple iterations.
6. The data processing method according to claim 5, characterized in that: The step of adjusting the number of training samples corresponding to the current iteration in the training sample sequence based on the global batch adjustment coefficient to obtain global batch information for the next iteration includes: Determine the product of the global batch adjustment coefficient and the number of training samples corresponding to the current iteration in the training sample sequence to obtain candidate global batch information for the next iteration; In a case where the number indicated by the candidate global batch information is less than the number indicated by the basic global batch information, using the basic global batch information as the global batch information of the next iteration; In a case where the number indicated by the candidate global batch information is greater than or equal to the number indicated by the basic global batch information, the candidate global batch information is used as the global batch information of the next iteration.
7. The data processing method according to claim 1, wherein: The distributed model training includes multiple training cycles, each of which includes the multiple iterations; and obtaining the reference model parameters and the reference model loss of each iteration includes: Obtaining reference model parameters and reference model losses for each iteration obtained by sequentially performing corresponding iterations based on the reference training samples in a training cycle previous to the current training cycle; Wherein, if the previous training cycle is the first training cycle among the multiple training cycles, the reference training sample order in the first training cycle is obtained by packaging samples based on a greedy bucketing method under the KL divergence constraint; If the previous training cycle is any training cycle after the first training cycle, the reference training sample sequence in any training cycle is the target training sample sequence corresponding to the any training cycle.
8. A data processing device, characterized in that: include: A sample sequence acquisition unit is configured to execute acquisition of a training sample sequence for distributed model training, wherein the training sample sequence includes a plurality of training sample sets corresponding one-to-one to a plurality of iterations, and the distributed model training includes training of the plurality of iterations; a cluster performance loss determining unit configured to determine a cluster performance loss for executing the distributed model training based on computing power consumption of the training samples corresponding to each iteration in the training sample sequence, wherein the cluster performance loss represents a severity of load imbalance caused by the training sample sequence in the plurality of iterations; A model effect loss prediction unit, the model effect loss prediction unit includes: a reference data acquisition unit, configured to execute acquisition of reference model parameters and reference model loss for each iteration; the reference model parameters are obtained by successfully executing the training of the corresponding iteration based on the reference training sample, and the reference model loss is used to characterize the difference between the predicted value obtained by the verification sample in the verification sample set based on the reference model parameters and the corresponding true label value; a parameter trajectory prediction unit, configured to execute parameter trajectory prediction for the training sample sequence based on the reference model parameters of each iteration, and obtain the prediction model parameters corresponding to the training sample sequence of each iteration; a prediction model loss determination unit, configured to execute prediction model parameters based on each iteration, and determine the prediction model loss for each iteration; the prediction model loss is used to characterize the difference between the predicted value obtained by the verification sample in the verification sample set based on the prediction model parameters and the corresponding true label value; a model effect loss determination unit, configured to execute the difference between the prediction model loss of each iteration and the reference model loss, and determine the model effect loss; The target sample sequence determination unit is configured to determine the comprehensive loss based on the cluster performance loss and the model effect loss, with the goal of minimizing the comprehensive loss. The training sample sequence is updated by reorganizing samples between training sample sets of different iterations until a preset end condition is met, thereby obtaining a target training sample sequence; the target training sample sequence is used to execute the distributed model training.
9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the data processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the data processing method according to any one of claims 1 to 7.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Training data processing method and device and storage medium
CN114332984A
Model training method, model training device and electronic equipment
CN114764593A