Breakpoint continuous training method and device, equipment, storage medium and program product
By calculating the recommended breakpoint saving step size and determining the breakpoint saving timing during the model training process, the inefficiency problem caused by the dependence on artificial settings of the breakpoint saving timing in the prior art is solved, and the effectiveness of breakpoints is ensured by saving verification information, and efficient breakpoint training is achieved.
Patent Information
- Application Number
- CN202411848420.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-06
AI Technical Summary
In the existing breakpoint training mechanism, the timing of breakpoint storage relies solely on artificial settings, resulting in low efficiency of breakpoint training, and the model data saved by breakpoints is not included in the verification mechanism, which may lead to incorrect breakpoint training.
By receiving the model training parameters and resource parameters sent by the user, the recommended breakpoint saving step size is calculated, and the breakpoint saving time is determined based on this step size during the model training process. In addition, when saving breakpoints, in addition to saving model data, the verification information of the model data is also saved to verify the validity of the breakpoint during training.
Effectively rationalize the frequency of breakpoint storage, improve the efficiency of breakpoint training, ensure the effectiveness of breakpoints, and avoid failure of training due to wrong breakpoints.
Smart Images

Figure CN119938260A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a breakpoint-resume training method and apparatus, equipment, storage medium, and program product. Background Art
[0002] Breakpoint saving refers to saving model data once during model training. During the model training process, model training may be interrupted due to various reasons. In this case, the model data saved at the most recent breakpoint will be selected as the breakpoint reference data, and the model training before the interruption will continue using the breakpoint reference data; this model training process can be called the breakpoint continuation training process.
[0003] In the current breakpoint resume training mechanism, the timing of breakpoint saving (referred to as breakpoint saving timing) is simply set according to human subjectivity. The breakpoint saving timing can only be adjusted after each training is completed or interrupted, and the efficiency of breakpoint resume training is low. Summary of the invention
[0004] Embodiments of the present application provide a breakpoint resume training method and apparatus, a processing device, a computer-readable storage medium, and a computer program product.
[0005] The breakpoint resume training method provided in the embodiment of the present application includes:
[0006] Receiving the model training parameters sent by the user terminal, and calculating the first step length of breakpoint saving based on the model training parameters and resource parameters;
[0007] The first step length is sent to the user terminal, and a second step length sent by the client is received, where the second step length is determined based on the first step length.
[0008] The breakpoint resume training method provided in the embodiment of the present application includes:
[0009] Receive a model training instruction sent by the task management module, wherein the model training instruction carries a second step length for breakpoint saving;
[0010] In response to the model training instruction, a resource pull-up instruction is sent to the resource management module, and the resource pull-up instruction is used to trigger the resource management module to pull up resources to perform the model training task; and, based on the second step size, a breakpoint save timing is determined, and breakpoint save is performed during the model training process based on the breakpoint save timing.
[0011] The breakpoint-resuming training device provided in the embodiment of the present application includes:
[0012] A first communication unit, used to receive model training parameters sent by a user terminal;
[0013] A processing unit, configured to calculate a first step length of breakpoint saving based on the model training parameters and resource parameters;
[0014] The first communication unit is further used to send the first step length to the user terminal and receive a second step length sent by the client, where the second step length is determined based on the first step length.
[0015] The breakpoint-resuming training device provided in the embodiment of the present application includes:
[0016] A second communication unit is used to receive a model training instruction sent by the task management module, wherein the model training instruction carries a second step length for saving breakpoints; in response to the model training instruction, a resource pull-up instruction is sent to the resource management module, wherein the resource pull-up instruction is used to trigger the resource management module to pull up resources to execute the model training task;
[0017] A saving unit is used to determine a breakpoint saving timing based on the second step length, and perform breakpoint saving during the model training process based on the breakpoint saving timing.
[0018] The processing device provided in the embodiment of the present application includes: a processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute any one of the above-mentioned breakpoint resume training methods.
[0019] The computer-readable storage medium provided in the embodiment of the present application is used to store a computer program, and the computer program enables a computer to execute any one of the above-mentioned breakpoint resume training methods.
[0020] The computer program product provided in the embodiment of the present application includes computer program instructions, which enable a computer to execute any one of the above-mentioned breakpoint resume training methods.
[0021] The technical solution of the embodiment of the present application supports the ability to recommend the step length for breakpoint resumption training. Specifically, the recommended step length is calculated based on the model training parameters and resource parameters. This can effectively rationalize the frequency of breakpoint saving and improve the efficiency of breakpoint resumption training. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flowchart of the breakpoint continuous training process;
[0023] Figure 2 This is a diagram of the implementation architecture of breakpoint-resume training provided in an embodiment of the present application;
[0024] Figure 3 This is a schematic diagram of the process of the breakpoint continuous training method provided in the embodiment of the present application. Figure 1 ;
[0025] Figure 4This is a schematic diagram of the process of the breakpoint continuous training method provided in the embodiment of the present application. Figure 2 ;
[0026] Figure 5 This is a schematic diagram of the process of the breakpoint continuous training method provided in the embodiment of the present application. Figure 3 ;
[0027] Figure 6 This is a schematic diagram of the structure of the breakpoint continuous training device provided in the embodiment of the present application. Figure 1 ;
[0028] Figure 7 This is a schematic diagram of the structure of the breakpoint continuous training device provided in the embodiment of the present application. Figure 2 ;
[0029] Figure 8 is a schematic structural diagram of a communication device provided in an embodiment of the present application;
[0030] Fig. 9 It is a schematic structural diagram of the chip of an embodiment of the present application. DETAILED DESCRIPTION
[0031] It should be noted that the term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0032] To facilitate understanding of the technical solutions of the embodiments of the present application, the relevant technologies / terms involved in the embodiments of the present application are explained below.
[0033] 1. Distributed parallel training
[0034] Parallel training and distributed training are methods to improve training efficiency and accelerate the training process in deep learning. They use multiple computing resources (such as multiple graphics processing units (GPUs) or multiple servers) to perform training tasks simultaneously.
[0035] Parallel training refers to the process of using multiple processors (such as multiple GPUs) on a single machine to perform training simultaneously. For example, parallel training has the following forms:
[0036] 1) Data Parallelism
[0037] In data parallelism, the training data is split into multiple mini-batches, and each processor processes a mini-batch of data at the same time and calculates the corresponding gradients. These gradients are then aggregated and used to update the model parameters.
[0038] 2) Model Parallelism
[0039] In model parallelism, different parts of the model are distributed across different processors, with each processor responsible for computing the forward and backward propagation of its corresponding part.
[0040] In addition to the above parallel training forms, there may also be tensor parallelism, pipeline parallelism and other forms. The technical solution of the embodiment of the present application does not restrict the specific form of parallel training.
[0041] Distributed training refers to the process of training on multiple machines (or nodes), each of which may have one or more processors. When distributed training adopts data parallelism and model parallelism strategies, this training method is called distributed parallel training.
[0042] It should be noted that the technical solution of the embodiment of the present application can be applied to the distributed parallel training process, that is, the "breakpoint resume training method / device" described in the embodiment of the present application can also be called "breakpoint resume training method / device based on distributed parallel training". Not limited to this, the technical solution of the embodiment of the present application can also be applied to other types of training processes.
[0043] 2. Breakpoint continuation training
[0044] Figure 1 The process of breakpoint continuous training is shown in Figure 1. Figure 1 As shown, the process includes the following steps:
[0045] Step 101: The user inputs model training parameters.
[0046] Step 102: The training task lifecycle management module initiates training based on the model training parameters.
[0047] Step 103: The task scheduling module sends a resource pull-up instruction to the resource management module.
[0048] Step 104: The resource management module pulls up related resources based on the resource pull-up instruction.
[0049] Step 105: The pulled-up related resources execute the training task.
[0050] Step 106: The resource reports a successful pull-up message to the resource management module and reports that the training task is being executed.
[0051] Step 107: The resource management module replies to the task scheduling module that the launch is successful and replies that the training task is being executed.
[0052] Step 108: The task scheduling module saves breakpoints according to the breakpoint requirements set by the training task lifecycle management module.
[0053] Here, breakpoint saving refers to saving the model data once during the model training process. The task scheduling module saves the model data to the specified location according to the breakpoint requirements set by the training task lifecycle management module.
[0054] Here, the breakpoint requirements set by the training task lifecycle management module include the step length, which refers to the number of training rounds (abbreviated as Epoch step length). A single round of training is referred to as an Epoch (a single round may be trained multiple times). For example, if the Epoch step length is EP, it means that a breakpoint save is executed after every EP Epoch training during the training process.
[0055] Here, model data includes but is not limited to model weights, optimizer status, scheduler status, etc.
[0056] Step 109: The resource management module reports the resource failure event to the task scheduling module, and the task scheduling module reports the resource failure event to the training task lifecycle management module.
[0057] Step 110: The training task lifecycle management module reallocates nodes or restarts nodes.
[0058] For example, if the resource failure event is a software failure, the training task lifecycle management will restart the node. For example, if the resource failure event is a hardware failure, the training task lifecycle management will reallocate the node.
[0059] Step 111: The resource management module starts / restarts related resources.
[0060] Step 112: The resource starts executing the training task from the breakpoint.
[0061] Specifically, the model data saved at the most recent breakpoint is selected as the breakpoint reference data, and the resource uses the breakpoint reference data to continue training the model before the interruption.
[0062] Step 113: The resource reports to the resource management module that the startup / restart is successful, and reports that the training task is being executed.
[0063] Step 114: The resource management module replies to the task scheduling module that the startup / restart is successful, and replies that the training task is being executed.
[0064] Step 115: The task scheduling module saves breakpoints according to the breakpoint requirements set by the training task lifecycle management module.
[0065] Step 116: When the training task is completed, the resource reports that the training is completed.
[0066] Above Figure 1In the breakpoint-resume training process shown, the timing of breakpoint saving (referred to as breakpoint saving timing) is simply set according to human subjectivity. The breakpoint saving timing needs to be adjusted each time training is completed or interrupted, and the efficiency of breakpoint-resume training is low. For example: if the breakpoint saving timing is set more frequently (i.e., the step size is smaller), then the overall training duration of the task will be longer; if the breakpoint saving timing is set more sparsely (i.e., the step size is larger), then although the overall training duration of the task is less affected by the breakpoint saving step, the validity of the model data saved at the breakpoint will be reduced. In addition, the model data saved at the breakpoint is not included in the verification mechanism, which may cause the model data saved at the most recent breakpoint to be unreliable. If training is resumed according to the model data saved at the most recent breakpoint, there is a problem of continued training failure due to continued training at the wrong breakpoint. To this end, the following technical solution of the embodiment of the present application is proposed.
[0067] To facilitate understanding of the technical solutions of the embodiments of the present application, the technical solutions of the present application are described in detail below through specific embodiments. The above related technologies can be combined arbitrarily with the technical solutions of the embodiments of the present application as optional solutions, and they all belong to the protection scope of the embodiments of the present application.
[0068] It should be noted that the “training” described in the embodiments of the present application may be but is not limited to “distributed parallel training”.
[0069] It should be noted that the model training process described in the embodiments of the present application includes E Epochs, where E can be understood as the total number of training rounds of the model, and E is a positive integer. An Epoch refers to the process of training a round of models using all samples in the training set. In other words, a single round of training is referred to as an Epoch (a single round may be trained multiple times).
[0070] It should be noted that the step size described in the embodiment of the present application refers to the number of training rounds (referred to as Epoch step size for short), and the step size is used to determine the timing of breakpoint saving. For example, if the Epoch step size is EP, it means that a breakpoint save is performed after every EP rounds of Epoch training during the training process.
[0071] It should be noted that the breakpoint saving described in the embodiments of the present application refers to saving the model data once during the model training process.
[0072] Figure 2 This is a diagram of the implementation architecture of breakpoint continuous training provided in the embodiment of the present application, such as Figure 2 As shown, the architecture includes: a training task lifecycle management module, a task scheduling module, a resource management module and resources.
[0073] The training task lifecycle management module is used to manage the start, pause, stop, restart, breakpoint saving, and breakpoint resumption of training tasks.
[0074] The task scheduling module is used to start, pause, stop, restart, save at breakpoints, and resume training at breakpoints for the related functions of the training task lifecycle management module.
[0075] The resource management module is used to manage the resources that carry the training and is scheduled by the task scheduling module. The managed resources include but are not limited to containers, virtual machines, bare metal, storage, networks, clusters, etc.
[0076] Resources refer to intelligent computing resources, which are managed by the resource management module and can provide the upper layer with resource monitoring information such as resource status information and resource networking information.
[0077] Figure 3 This is a schematic diagram of the process of the breakpoint continuous training method provided in the embodiment of the present application. Figure 1 ,The breakpoint continuation training method is applied to the training task lifecycle management module (referred to as the task management module), such as Figure 3 As shown, the breakpoint continuous training method includes the following steps:
[0078] Step 301: Receive the model training parameters sent by the user terminal, and calculate the first step length of breakpoint saving based on the model training parameters and resource parameters; wherein breakpoint saving refers to saving the model data once during the model training process.
[0079] In an embodiment of the present application, the model training parameters include model parameters and / or training parameters.
[0080] In some embodiments, the model training parameters include at least one of the following:
[0081] The first parameter indicates the size of the model data;
[0082] The second parameter indicates the number of training rounds included in the model training process;
[0083] The third parameter represents the training duration of a single round, that is, the average training duration of a single round in Epoch.
[0084] In some implementations, the resource parameter includes at least one of the following:
[0085] The fourth parameter indicates the maximum storage space for saving breakpoints;
[0086] The fifth parameter indicates the stability working time of the device, that is, the average single stability working time of the device.
[0087] For ease of description, the first parameter may be expressed as M, the second parameter may be expressed as E, the third parameter may be expressed as T, the fourth parameter may be expressed as Dm, and the fifth parameter may be expressed as Ts.
[0088] In the embodiment of the present application, the first step of calculating the breakpoint saving based on the model training parameters and resource parameters includes:
[0089] Calculate the value range of the number of breakpoint save times based on the first parameter and the fourth parameter;
[0090] Calculate the value range of the first step length based on the third parameter and the fifth parameter;
[0091] Based on the value range of the second parameter and the number of breakpoint save times, a value that meets the requirements is screened out from the value range of the first step length, and the value that meets the requirements is used as the value of the first step length.
[0092] For ease of description, the number of breakpoint saves can be expressed as N, and the first step length can be expressed as EP, which is also the recommended step length. In addition, assume that the space occupied by saving model data is Dn, and the time interval between two model data saves is Tn.
[0093] The relationship between the above parameters can be expressed by the following formula:
[0094]
[0095] Dn = M * N; (2-1)
[0096] Dn = M * K; (2-2)
[0097] Tn = T * EP (3)
[0098] Among them, if the model data saved at N breakpoints are stored, Dn satisfies formula (2-1); if the model data saved at the most recent K breakpoints are stored, Dn satisfies formula (2-2).
[0099] It should be noted that among the above parameters, M, T, Ts, E, and Dm are known quantities, and N, EP, and K are unknown quantities. The following example illustrates how EP is calculated through model training parameters.
[0100] Example 1
[0101] If all model data saved at N breakpoints are saved, the following conditions are met:
[0102] Dn ≤ Dm, Tn ≤ Ts (4-1)
[0103] Substituting the above formula (2-1) and formula (3) into formula (4-1), we can get the following formula to represent the conditions:
[0104] M * N ≤ Dm, T * EP ≤ Ts (4-2)
[0105] In addition, EP also needs to satisfy the constraints expressed in the following formula:
[0106] E / N ≤ EP ≤ E / (N-1) (4-3)
[0107] For example, M = 0.2G, E = 10000 times, T = 30min, Ts = 300min, Dm = 300G. According to the above formula (4-2), the value ranges of EP and N can be calculated, where the value range of EP is [1, 10], and the value range of N is (0, 1500). In the value range of EP, the values that meet the requirements (i.e., formula (4-3)) are selected in the following manner:
[0108] When EP is 1, E / N≤1≤E / (N-1), then the value range of N is [10000,10001], which obviously does not meet the requirements (that is, it is inconsistent with the value range of N determined above as (0,1500]).
[0109] When EP is 2, E / N≤2≤E / (N-1), then the value range of N is [5000,5001], which obviously does not meet the requirements (that is, it is inconsistent with the value range of N determined above as (0,1500]).
[0110] When EP is 3, E / N≤3≤E / (N-1), then the value range of N is [3333.4,3334.3], which obviously does not meet the requirements (that is, it is inconsistent with the value range of N determined above as (0,1500]).
[0111] When EP is 4, E / N≤4≤E / (N-1), then the value range of N is [2500,2501], which obviously does not meet the requirements (that is, it is inconsistent with the value range of N determined above as (0,1500]).
[0112] When EP is 5, E / N≤5≤E / (N-1), then the value range of N is [2500,2501], which obviously does not meet the requirements (that is, it is inconsistent with the value range of N determined above as (0,1500]).
[0113] When EP is 6, E / N≤6≤E / (N-1), then the value range of N is [1666.7,1667.6], which obviously does not meet the requirements (that is, it is inconsistent with the value range of N determined above as (0,1500]).
[0114] When EP is 7, E / N≤7≤E / (N-1), then the value range of N is [1428.6,1429.5], and N is 1429, which meets the requirement of N≤1500. EP can be 7 and N can be 1429.
[0115] When EP is 8, E / N≤8≤E / (N-1), then the value range of N is [1250,1251], N can be 1250 or 1251, and EP can be 8.
[0116] When EP is 9, E / N≤9≤E / (N-1), then the value range of N is [1111.2,1112.2], N can be 1112, and EP can be 9.
[0117] When EP is 10, E / N≤10≤E / (N-1), then the value range of N is [1000,1001]. If N can be 1000 or 1001, EP can be 10.
[0118] As shown above, the values of EP that meet the requirements are {7, 8, 9, 10}, and the corresponding values of N are {1429, 1250 / 1251, 1112, 1000 / 1001}. Select one of the values of EP that meet the requirements (for example, 8) as the first step length.
[0119] Example 2
[0120] If the model data saved at the most recent K breakpoints is saved, the following conditions are met:
[0121] Dn ≤ Dm, Tn ≤ Ts, K ≤ N (5-1)
[0122] Substituting the above formula (2-1) and formula (3) into formula (5-1), we can get the following formula to represent the conditions:
[0123] M * K ≤ Dm, M * N ≥ Dm, T * EP ≤ Ts (5-2)
[0124] In addition, EP also needs to satisfy the constraints expressed in the following formula:
[0125] E / N ≤ EP ≤ E / (N-1) (5-3)
[0126] For example, M = 0.2G, E = 10000 times, T = 30min, Ts = 300min, Dm = 300G. According to the above formula (5-2), EP≤10 can be calculated. EP is the recommended step size and cannot be zero. The value range of EP is [1,10]. In addition, K≤1500 can be calculated. The value range of K is (0,1500]. In addition, N≥1500 can be calculated. The value range of N is [1500,∞). In the value range of EP, the values that meet the requirements (i.e., formula (5-3)) are selected as follows:
[0127] When EP is 1, E / N≤1≤E / (N-1), then the value range of N is [10000,10001]. If N is 10000 or 10001, K can be any value in the range of (0,1500).
[0128] When EP is 2, E / N≤2≤E / (N-1), then the value range of N is [5000,5001]. If N is 5000 or 5001, K can be any value in the range of (0,1500].
[0129] When EP is 3, E / N≤3≤E / (N-1), then the value range of N is [3333.4,3334.3]. If N is 3334, K can be any value in the range of (0,1500).
[0130] When EP is 4, E / N≤4≤E / (N-1), then the value range of N is [2500,2501]. If N is 2500, then K can be any value in the range of (0,1500).
[0131] When EP is 5, E / N≤5≤E / (N-1), then the value range of N is [2500,2501]. If N is 2500 or 2501, K can be any value in the range of (0,1500).
[0132] When EP is 6, E / N≤6≤E / (N-1), then the value range of N is [1666.7,1667.6]. When N is 1667, K can be any value in the range of (0,1500].
[0133] When EP is 7, E / N≤7≤E / (N-1), then the value range of N is [1428.6,1429.5]. When N is 1429, K can be any value in the range of (0,1429].
[0134] When EP is 8, E / N≤8≤E / (N-1), then the value range of N is [1250,1251]. If N is 1250 or 1251, K can be any value in the range of (0,1250 / 1251].
[0135] When EP is 9, E / N≤9≤E / (N-1), then the value range of N is [1111.2,1112.2]. When N is 1112, K can take any value in the range of (0,1112].
[0136] When EP is 10, E / N≤10≤E / (N-1), then the value range of N is [1000,1001]. If N is 1000 or 1001, K can be any value in the range of (0,1000 / 1001].
[0137] As shown above, the EP value that meets the requirements is [1,10], and the corresponding N and K values are as above. You can choose a suitable step size based on the customer's requirements for the K value. For example, from the perspective of high efficiency and storage space utilization efficiency, you can choose a value of 5 as the first step size, and accordingly, K can be any positive integer value below 1500, such as 1500.
[0138] Step 302: Send the first step length to the user terminal, and receive the second step length sent by the client, where the second step length is determined based on the first step length.
[0139] Here, the client determines the second step length according to the first step length. In some embodiments, the second step length is consistent with the first step length. In other embodiments, the second step length is inconsistent with the first step length, for example, the second step length is equal to the first step length plus an offset.
[0140] It should be noted that the second step length is the step length actually used in breakpoint training.
[0141] In some embodiments, the method further includes the following steps: sending a model training instruction to a task scheduling module, the model training instruction carries a second step length, the second step length is used by the task scheduling module to determine the breakpoint save timing, and save the breakpoint based on the breakpoint save timing.
[0142] In an embodiment of the present application, a training task lifecycle management module (referred to as the task management module for short) initiates model training, and carries a second step length in the model training instruction sent to the task scheduling module. The task scheduling module determines the breakpoint save timing based on the second step length, and saves the breakpoint based on the breakpoint save timing.
[0143] In some implementations, in addition to saving the model data, the breakpoint save also saves verification information of the model data.
[0144] In some embodiments, if the training task lifecycle management module (referred to as the task management module for short) receives a resource fault instruction sent by the task scheduling module, the verification information of the saved model data is verified, and the latest model data that has passed the verification is selected as the breakpoint reference data; a resource fault recovery instruction is sent to the task scheduling module, and the resource fault recovery instruction carries the breakpoint reference data, and the breakpoint reference data is used for the continued execution of the model training process.
[0145] In some embodiments, the verification information includes at least one of the following:
[0146] The first random value is a random value generated by saving the current breakpoint;
[0147] The first context value is the context value of this breakpoint save, and the first context value is calculated based on the second context value and the first random value, and the second context value is the context value of the breakpoint save before this breakpoint save; if this breakpoint save is the first breakpoint save, then the corresponding first context value is equal to the first random value generated by this breakpoint save;
[0148] First training-related data, the first training-related data is the training-related data saved at this breakpoint;
[0149] The first time data is the time data saved at this breakpoint.
[0150] In one example, for each breakpoint save, in addition to saving the model data, the verification information corresponding to the current model data is also saved. The content of the verification information can be expressed as: [random value, context value, training-related typical algorithm data, time data]. For ease of description, the random value can be expressed as A, the context value can be expressed as B, the training-related typical algorithm data can be expressed as C, and the time data can be expressed as D. For the nth breakpoint save, the content of the verification information can be calculated by the following formula:
[0151] A[n] = random value (6-1)
[0152] B[n]= B[n-1]+A[n] (6-2)
[0153] C[n] = typical algorithm (n), for example, C[n] = log n (6-3)
[0154] D[n] = breakpoint save time (6-4)
[0155] To verify the verification information saved at the nth breakpoint, you can use the following methods:
[0156] 1) Calculate B[n] using the check information A[0] to A[n] saved at the breakpoints, and compare the calculated B[n] with the check information B[n] saved at the nth breakpoint to verify whether they are consistent.
[0157] 2) Calculate C[n] according to formula (6-3), and compare the calculated C[n] with the verification information C[n] saved at the nth breakpoint to verify whether they are consistent.
[0158] 3) Compare the file save time of the model data with the verification information D[n] saved at the nth breakpoint to verify whether they are consistent. Here, "consistent" means "whether the time difference exceeds a fixed time limit". If it exceeds the fixed time limit, it indicates inconsistency. If it does not exceed the fixed time limit, it indicates consistency.
[0159] If the above three items are verified to be consistent, it is determined that the model data saved at the nth breakpoint has been verified.
[0160] In some embodiments, if the most recent K breakpoint saves are stored, then when the next breakpoint save (recorded as the i-th breakpoint save) is performed, a certain breakpoint save among the K breakpoint saves (recorded as the j-th breakpoint save) needs to be verified, and the verification method can refer to the above description of "verifying the verification information of the n-th breakpoint save"; if the j-th breakpoint save passes the verification, the content of the breakpoint save before the j-th breakpoint save is deleted, and the content of the i-th breakpoint save is stored, so as to maintain that the most recent K breakpoint saves are always stored. Among them, the content stored in each breakpoint save includes model data and corresponding verification information.
[0161] The technical solution of the embodiment of the present application, on the one hand, supports the ability to recommend the step length for breakpoint continuation training. Specifically, the recommended step length is calculated based on the model training parameters (such as the size of the model data, the maximum storage space, the number of training rounds included in the model training process, the average single-round training time of the Epoch, the average single-time stability working time of the resource, etc.), which can effectively reduce the frequency of breakpoint saving. On the other hand, when saving the breakpoint, in addition to saving the model data, the verification information corresponding to the model data is also saved. The validity of the breakpoint (i.e., the model data saved at the breakpoint) is verified during the continuation of training through the verification information, so as to achieve effective continuation of breakpoint training and optimize the final training time. The following problems caused by some hardware and software failures or link virtual connection can be avoided: using invalid model data for continuation training to obtain meaningless results, wasting resources and time.
[0162] Figure 4 This is a schematic diagram of the process of the breakpoint continuous training method provided in the embodiment of the present application. Figure 2 ,This breakpoint continuous training method is applied to the task scheduling module, such as Figure 4As shown, the breakpoint continuous training method includes the following steps:
[0163] Step 401: Receive a model training instruction sent by a task management module, where the model training instruction carries a second step length for breakpoint saving; wherein breakpoint saving refers to saving model data once during the model training process.
[0164] In the embodiment of the present application, the calculation method of the second step length can refer to the aforementioned Figure 3 Related description.
[0165] Step 402: In response to the model training instruction, a resource pull-up instruction is sent to the resource management module, where the resource pull-up instruction is used to trigger the resource management module to pull up resources to perform the model training task; and, based on the second step length, a breakpoint save timing is determined, and breakpoint save is performed during the model training process based on the breakpoint save timing.
[0166] For example, assuming that the second step length = 8, it means that a breakpoint save is performed every time 8 epochs are completed during the model training process.
[0167] In some implementations, in addition to saving the model data, the breakpoint save also saves verification information of the model data.
[0168] In some embodiments, if the resource management module receives a resource failure instruction, it sends the resource failure instruction to the task management module / task scheduling module; the resource management module receives the resource failure recovery instruction sent by the task management module / task scheduling module, and the resource failure recovery instruction carries breakpoint reference data, which is the model data verified by the task management module; the breakpoint reference data is used for the continued execution of the model training process.
[0169] In some embodiments, the verification information includes at least one of the following:
[0170] The first random value is a random value generated by saving the current breakpoint;
[0171] The first context value is the context value of this breakpoint save, and the first context value is calculated based on the second context value and the first random value, and the second context value is the context value of the breakpoint save before this breakpoint save; if this breakpoint save is the first breakpoint save, then the corresponding first context value is equal to the first random value generated by this breakpoint save;
[0172] First training-related data, the first training-related data is the training-related data saved at this breakpoint;
[0173] The first time data is the time data saved at this breakpoint.
[0174] In one example, for each breakpoint save, in addition to saving the model data, the verification information corresponding to the current model data is also saved. The content of the verification information can be expressed as: [random value, context value, training-related typical algorithm data, time data]. For ease of description, the random value can be expressed as A, the context value can be expressed as B, the training-related typical algorithm data can be expressed as C, and the time data can be expressed as D. For the nth breakpoint save, the content of the verification information can be calculated by the following formula:
[0175] A[n] = random value (6-1)
[0176] B[n]= B[n-1]+A[n] (6-2)
[0177] C[n] = typical algorithm (n), for example, C[n] = log n (6-3)
[0178] D[n] = breakpoint save time (6-4)
[0179] To verify the verification information saved at the nth breakpoint, you can use the following methods:
[0180] 1) Calculate B[n] using the check information A[0] to A[n] saved at the breakpoints, and compare the calculated B[n] with the check information B[n] saved at the nth breakpoint to verify whether they are consistent.
[0181] 2) Calculate C[n] according to formula (6-3), and compare the calculated C[n] with the verification information C[n] saved at the nth breakpoint to verify whether they are consistent.
[0182] 3) Compare the file save time of the model data with the verification information D[n] saved at the nth breakpoint to verify whether they are consistent. Here, "consistent" means "whether the time difference exceeds a fixed time limit". If it exceeds the fixed time limit, it indicates inconsistency. If it does not exceed the fixed time limit, it indicates consistency.
[0183] If the above three items are verified to be consistent, it is determined that the model data saved at the nth breakpoint has been verified.
[0184] In some embodiments, if the most recent K breakpoint saves are stored, then when the next breakpoint save (recorded as the i-th breakpoint save) is performed, a certain breakpoint save among the K breakpoint saves (recorded as the j-th breakpoint save) needs to be verified, and the verification method can refer to the above description of "verifying the verification information of the n-th breakpoint save"; if the j-th breakpoint save passes the verification, the content of the breakpoint save before the j-th breakpoint save is deleted, and the content of the i-th breakpoint save is stored, so as to maintain that the most recent K breakpoint saves are always stored. Among them, the content stored in each breakpoint save includes model data and corresponding verification information.
[0185] The technical solution of the embodiment of the present application, on the one hand, supports the ability to recommend the step length of breakpoint continuation training. Specifically, the recommended step length is calculated based on the model training parameters and resource parameters (such as the size of the model data, the maximum storage space, the number of training rounds included in the model training process, the average single-round training time of the Epoch, the average single-time stability working time of the resource, etc.), which can effectively reduce the frequency of breakpoint saving. On the other hand, when saving the breakpoint, in addition to saving the model data, the verification information corresponding to the model data is also saved. The validity of the breakpoint (i.e., the model data saved at the breakpoint) is verified during the continuation of training through the verification information, so as to achieve effective continuation of breakpoint training and optimize the final training time. The following problems caused by some hardware and software failures or link virtual connection can be avoided: using invalid model data for continuation training to obtain meaningless results, wasting resources and time.
[0186] Figure 5 This is a schematic diagram of the process of the breakpoint continuous training method provided in the embodiment of the present application. Figure 3 ,like Figure 5 As shown, the breakpoint continuous training method includes the following steps:
[0187] Step 501: The user inputs model training parameters.
[0188] Step 502: The training task lifecycle management module calculates the recommended step length (ie, the first step length) according to the model training parameters and resource parameters.
[0189] Here, the calculation method of the recommended step length (i.e. the first step length) can refer to the above Figure 3 Description of the relevant scheme.
[0190] Step 503: The training task lifecycle management module returns the recommended step size to the user end.
[0191] Step 504: The client selects the step size actually used for breakpoint training (ie, the second step size) according to the recommended step size, and returns the step size actually used to the training task lifecycle management module.
[0192] Step 505: The training task lifecycle management module initiates training.
[0193] Step 506: The task scheduling module sends a resource pull-up instruction to the resource management module.
[0194] Step 507: The resource management module pulls up related resources based on the resource pull-up instruction.
[0195] Step 508: The pulled-up related resources execute the training task.
[0196] Step 509: The resource checks with the resource management module and reports that the pull-up is successful, and reports that the training task is being executed.
[0197] Step 510: The resource management module replies to the task scheduling module that the launch is successful and replies that the training task is being executed.
[0198] Step 511: The task scheduling module saves breakpoints according to the step size (ie, the second step size) set by the training task lifecycle management module.
[0199] Here, if the task scheduling module only saves the model data saved at the most recent K breakpoints, then before each model data deletion operation is performed, the verification information of the deleted model data is verified to ensure that the earliest model data saved after deletion is accurate.
[0200] Step 512: The resource management module reports the resource failure event to the task scheduling module, and the task scheduling module reports the resource failure event to the training task lifecycle management module.
[0201] Step 513: The training task lifecycle management module performs breakpoint verification and selects the model data saved at the most recent breakpoint that has passed the verification as the breakpoint reference data for subsequent training.
[0202] Here, the breakpoint verification method can refer to the above Figure 3 Description of the relevant scheme.
[0203] Step 514: The training task lifecycle management module reallocates nodes or restarts nodes.
[0204] For example, if the resource failure event is a software failure, the training task lifecycle management will restart the node. For example, if the resource failure event is a hardware failure, the training task lifecycle management will reallocate the node.
[0205] Step 515: The resource management module starts / restarts related resources.
[0206] Step 516: The resource starts executing the training task from the breakpoint.
[0207] Specifically, the resource uses the breakpoint reference data to continue training the model before the interruption.
[0208] Step 517: The resource checks with the resource management module and reports that the startup / restart is successful, and reports that the training task is being executed.
[0209] Step 518: The resource management module replies to the task scheduling module that the startup / restart is successful, and replies that the training task is being executed.
[0210] Step 519: The task scheduling module saves breakpoints according to the step size (ie, the second step size) set by the training task lifecycle management module.
[0211] Step 520: When the training task is completed, the resource reports that the training is completed.
[0212] Figure 6 This is a schematic diagram of the structure of the breakpoint continuous training device provided in the embodiment of the present application. Figure 1 ,like Figure 6 As shown, the breakpoint continuous training device comprises:
[0213] The first communication unit 601 is used to receive the model training parameters sent by the user end;
[0214] The processing unit 602 is used to calculate the first step length of breakpoint saving based on the model training parameters and resource parameters; wherein the breakpoint saving refers to saving the model data once during the model training process;
[0215] The first communication unit 601 is further configured to send the first step length to the user terminal and receive a second step length sent by the client, wherein the second step length is determined based on the first step length.
[0216] In some embodiments, the first communication unit 601 is further used to send a model training instruction to the task scheduling module, and the model training instruction carries the second step length. The second step length is used by the task scheduling module to determine the breakpoint save timing and perform breakpoint save based on the breakpoint save timing.
[0217] In some implementations, the breakpoint save saves verification information of the model data in addition to the model data;
[0218] The processing unit 602 is used to receive the resource failure instruction sent by the task scheduling module, verify the verification information of the saved model data, and select the model data that passes the verification as the breakpoint reference data;
[0219] The first communication unit 601 is used to send a resource fault recovery instruction to the task scheduling module, and the resource fault recovery instruction carries the breakpoint reference data, and the breakpoint reference data is used for the continued execution of the model training process.
[0220] In some implementations, the verification information includes at least one of the following:
[0221] A first random value, where the first random value is a random value generated by saving the current breakpoint;
[0222] a first context value, the first context value being the context value of the current breakpoint save, the first context value being calculated based on the second context value and the first random value, the second context value being the context value of the last breakpoint save before the current breakpoint save; if the current breakpoint save is the first breakpoint save, then the corresponding first context value is equal to the first random value generated by the current breakpoint save;
[0223] First training-related data, where the first training-related data is the training-related data saved at this breakpoint;
[0224] The first time data is the time data saved at this breakpoint.
[0225] In some embodiments, the model training parameters include at least one of the following:
[0226] A first parameter, wherein the first parameter indicates the size of the model data;
[0227] A second parameter, wherein the second parameter represents the number of training rounds included in the model training process;
[0228] The third parameter represents the training duration of a single round, that is, the average training duration of a single round in Epoch.
[0229] In some implementations, the resource parameter includes at least one of the following:
[0230] The fourth parameter indicates the maximum storage space for saving breakpoints;
[0231] The fifth parameter represents the stability working time of the device, that is, the average single stability working time of the device.
[0232] In some embodiments, the processing unit 602 is used to calculate the value range of the breakpoint save times based on the first parameter and the fourth parameter; calculate the value range of the first step length based on the third parameter and the fifth parameter; based on the second parameter and the value range of the breakpoint save times, filter out a value that meets the requirements from the value range of the first step length, and use the value that meets the requirements as the value of the first step length.
[0233] Those skilled in the art should understand that Figure 6 The functions implemented by each unit in the breakpoint-resuming training device shown can be understood by referring to the relevant description of the aforementioned method. Figure 6 The functions of each unit in the breakpoint continuous training device shown can be implemented by a program running on a processor, or by a specific logic circuit.
[0234] Figure 7 This is a schematic diagram of the structure of the breakpoint continuous training device provided in the embodiment of the present application. Figure 2 ,like Figure 7 As shown, the breakpoint continuous training device comprises:
[0235] The second communication unit 701 is used to receive a model training instruction sent by the task management module, wherein the model training instruction carries a second step length for breakpoint saving; wherein the breakpoint saving refers to saving the model data once during the model training process; in response to the model training instruction, a resource pull-up instruction is sent to the resource management module, wherein the resource pull-up instruction is used to trigger the resource management module to pull up resources to execute the model training task;
[0236] The saving unit 702 is used to determine the breakpoint saving timing based on the second step length, and perform breakpoint saving during the model training process based on the breakpoint saving timing.
[0237] In some implementations, the breakpoint save saves not only the model data but also verification information of the model data.
[0238] In some embodiments, the second communication unit 701 is used to receive a resource fault instruction sent by the resource management module, and forward the resource fault instruction to the task management module; receive a resource fault recovery instruction sent by the task management module, and the resource fault recovery instruction carries breakpoint reference data, and the breakpoint reference data is model data verified by the task management module; the breakpoint reference data is used for the continued execution of the model training process.
[0239] In some implementations, the verification information includes at least one of the following:
[0240] A first random value, where the first random value is a random value generated by saving the current breakpoint;
[0241] a first context value, the first context value being the context value of the current breakpoint save, the first context value being calculated based on the second context value and the first random value, the second context value being the context value of the last breakpoint save before the current breakpoint save; if the current breakpoint save is the first breakpoint save, then the corresponding first context value is equal to the first random value generated by the current breakpoint save;
[0242] First training-related data, where the first training-related data is the training-related data saved at this breakpoint;
[0243] The first time data is the time data saved at this breakpoint.
[0244] Those skilled in the art should understand that Figure 7 The functions implemented by each unit in the breakpoint-resuming training device shown can be understood by referring to the relevant description of the aforementioned method. Figure 7 The functions of each unit in the breakpoint continuous training device shown can be implemented by a program running on a processor, or by a specific logic circuit.
[0245] Figure 8 It is a schematic structural diagram of a processing device 800 provided in an embodiment of the present application. Figure 8 The processing device 800 shown includes a processor 810, which can call and run a computer program from a memory to implement the method in the embodiment of the present application.
[0246] Alternatively, if Figure 8 As shown, the processing device 800 may further include a memory 820. The processor 810 may call and run a computer program from the memory 820 to implement the method in the embodiment of the present application.
[0247] The memory 820 may be a separate device independent of the processor 810 , or may be integrated into the processor 810 .
[0248] Alternatively, if Figure 8 As shown, the processing device 800 may further include a transceiver 830, and the processor 810 may control the transceiver 830 to communicate with other devices, specifically, may send information or data to other devices, or receive information or data sent by other devices.
[0249] The processing device 800 can specifically be an intelligent computing platform, and the processing device 800 can implement the corresponding processes of each method implemented in the embodiments of the present application. For the sake of brevity, they will not be repeated here.
[0250] Fig. 9It is a schematic structural diagram of the chip of an embodiment of the present application. Fig. 9 The chip 900 shown includes a processor 910, which can call and run a computer program from a memory to implement the method in the embodiment of the present application.
[0251] Alternatively, if Fig. 9 As shown, the chip 900 may further include a memory 920. The processor 910 may call and run a computer program from the memory 920 to implement the method in the embodiment of the present application.
[0252] The memory 920 may be a separate device independent of the processor 910 , or may be integrated into the processor 910 .
[0253] Optionally, the chip 900 may further include an input interface 930. The processor 910 may control the input interface 930 to communicate with other devices or chips, and specifically, may obtain information or data sent by other devices or chips.
[0254] Optionally, the chip 900 may further include an output interface 940. The processor 910 may control the output interface 940 to communicate with other devices or chips, and specifically, may output information or data to other devices or chips.
[0255] The chip can be applied to an intelligent computing platform, and the chip can implement the corresponding processes of each method implemented in the embodiments of the present application. For the sake of brevity, they will not be repeated here.
[0256] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0257] It should be understood that the processor of the embodiment of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiment can be completed by the hardware integrated logic circuit or software instructions in the processor. The above processor can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or the hardware and software modules in the decoding processor are combined and performed. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, and other mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0258] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0259] It should be understood that the above-mentioned memory is exemplary but not restrictive. For example, the memory in the embodiments of the present application may also be static random access memory (static RAM, SRAM), dynamic random access memory (dynamic RAM, DRAM), synchronous dynamic random access memory (synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (double data rate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (enhanced SDRAM, ESDRAM), synchronous link dynamic random access memory (synch link DRAM, SLDRAM) and direct memory bus random access memory (Direct Rambus RAM, DR RAM), etc. That is to say, the memory in the embodiments of the present application is intended to include but not limited to these and any other suitable types of memory.
[0260] The embodiment of the present application also provides a computer-readable storage medium for storing a computer program. The computer-readable storage medium can be applied to the intelligent computing platform, and the computer program enables the computer to execute the corresponding processes implemented by each method of the embodiment of the present application, which will not be described here for the sake of brevity.
[0261] The embodiment of the present application also provides a computer program product, including computer program instructions. The computer program product can be applied to an intelligent computing platform, and the computer program instructions enable a computer to execute the corresponding processes implemented by each method of the embodiment of the present application, which will not be described here for brevity.
[0262] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0263] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0264] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0265] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0266] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0267] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0268] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A breakpoint resume training method, characterized in that: The method comprises: Receiving the model training parameters sent by the user terminal, and calculating the first step length of breakpoint saving based on the model training parameters and resource parameters; The first step length is sent to the user terminal, and a second step length sent by the client is received, where the second step length is determined based on the first step length.
2. The method according to claim 1, characterized in that The breakpoint saving saves not only the model data but also verification information of the model data.
3. The method according to claim 1, characterized in that: The method further comprises: A model training instruction is sent to the task scheduling module, wherein the model training instruction carries the second step length, and the second step length is used by the task scheduling module to determine a breakpoint saving timing, and to save the breakpoint based on the breakpoint saving timing.
4. The method according to claim 2, characterized in that: The method further comprises: If a resource failure instruction is received from the task scheduling module, the verification information of the saved model data is verified, and the model data that has passed the verification is selected as the breakpoint reference data; A resource fault recovery instruction is sent to the task scheduling module, wherein the resource fault recovery instruction carries the breakpoint reference data, and the breakpoint reference data is used for the continued execution of the model training process.
5. The method according to claim 2 or 4, characterized in that: The verification information includes at least one of the following: A first random value, where the first random value is a random value generated by saving the current breakpoint; a first context value, the first context value being a context value saved at the current breakpoint, the first context value being calculated based on the second context value and the first random value, the second context value being a context value saved at a breakpoint before the current breakpoint is saved; First training-related data, where the first training-related data is the training-related data saved at this breakpoint; The first time data is the time data saved at this breakpoint.
6. The method according to any one of claims 1 to 4, characterized in that The model training parameters include at least one of the following: A first parameter, wherein the first parameter indicates the size of the model data; A second parameter, wherein the second parameter represents the number of training rounds included in the model training process; A third parameter, wherein the third parameter represents the training duration of a single round; The resource parameters include at least one of the following: A fourth parameter, the fourth parameter indicates a maximum storage space for saving breakpoints; A fifth parameter indicates the stable working time of the device.
7. The method according to claim 6, characterized in that The first step of calculating the breakpoint saving based on the model training parameters and resource parameters includes: Calculate a value range of the number of breakpoint save times based on the first parameter and the fourth parameter; Calculate a value range of the first step length based on the third parameter and the fifth parameter; Based on the value range of the second parameter and the number of breakpoint save times, a value that meets the requirements is screened out from the value range of the first step length, and the value that meets the requirements is used as the value of the first step length.
8. A breakpoint resume training method, characterized in that: The method comprises: Receive a model training instruction sent by the task management module, wherein the model training instruction carries a second step length for breakpoint saving; In response to the model training instruction, a resource pull-up instruction is sent to the resource management module, and the resource pull-up instruction is used to trigger the resource management module to pull up resources to perform the model training task; and, based on the second step size, a breakpoint save timing is determined, and breakpoint save is performed during the model training process based on the breakpoint save timing.
9. The method according to claim 8, characterized in that The breakpoint saving saves not only the model data but also verification information of the model data.
10. The method according to claim 9, characterized in that The method further comprises: If a resource failure instruction is received, forwarding the resource failure instruction to the task management module; Receive a resource fault recovery instruction sent by the task management module, the resource fault recovery instruction carries breakpoint reference data, the breakpoint reference data is the model data verified by the task management module; the breakpoint reference data is used for the continued execution of the model training process.
11. The method according to claim 9 or 10, characterized in that: The verification information includes at least one of the following: A first random value, where the first random value is a random value generated by saving the current breakpoint; a first context value, the first context value being a context value saved at the current breakpoint, the first context value being calculated based on the second context value and the first random value, the second context value being a context value saved at a breakpoint before the current breakpoint is saved; First training-related data, where the first training-related data is the training-related data saved at this breakpoint; The first time data is the time data saved at this breakpoint.
12. A breakpoint resume training device, characterized in that: The device comprises: A first communication unit, used to receive model training parameters sent by a user terminal; A processing unit, configured to calculate a first step length of breakpoint saving based on the model training parameters and resource parameters; The first communication unit is further used to send the first step length to the user terminal and receive a second step length sent by the client, where the second step length is determined based on the first step length.
13. The device according to claim 12, characterized in that The first communication unit is also used to send a model training instruction to the task scheduling module, and the model training instruction carries the second step length. The second step length is used by the task scheduling module to determine the breakpoint saving timing and perform breakpoint saving based on the breakpoint saving timing.
14. A breakpoint resume training device, characterized in that: The device comprises: A second communication unit is used to receive a model training instruction sent by the task management module, wherein the model training instruction carries a second step length for saving breakpoints; in response to the model training instruction, a resource pull-up instruction is sent to the resource management module, wherein the resource pull-up instruction is used to trigger the resource management module to pull up resources to execute the model training task; A saving unit is used to determine a breakpoint saving timing based on the second step length, and perform breakpoint saving during the model training process based on the breakpoint saving timing.
15. A processing device, characterized in that: include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 11.
16. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein the computer program causes a computer to execute the method according to any one of claims 1 to 11.
17. A computer program product, characterized in that The method comprises computer program instructions which cause a computer to execute the method as claimed in any one of claims 1 to 11.
Citation Information
Cited By
Large language model breakpoint continuous training method, system and device and storage medium
CN121808393A