Cluster-based training method, device, electronic device and storage medium
By introducing multiple spare nodes into the cluster and quickly initializing the spare node to take over the training task when a hardware failure occurs in the training node, the problem of cluster training stagnation due to a single node failure is solved, and more efficient model training recovery is achieved.
Patent Information
- Application Number
- CN202311767093.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-12-20
AI Technical Summary
When using a cluster to train a large model, a hardware failure of a single node will cause the cluster to stop training jobs, affecting the training efficiency of the model.
A cluster-based training method is provided. By introducing multiple spare nodes into the cluster, responding to the hardware failure of the training node, selecting the target spare node from the spare node, and obtaining the target training snapshot of the model training task of the training node. Based on the model training program container image and training status data of the training node, the target spare node is initialized to replace the training node and continue to perform the model training task.
When a hardware failure occurs in a training node, the backup node is initialized at a fast speed to quickly restore the model training task, shorten the fault recovery time, and improve the model training efficiency.
Smart Images

Figure CN117742959B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, specifically to the fields of artificial intelligence and deep learning, and in particular to a cluster-based training method, device, electronic device and storage medium. Background Art
[0002] With the continuous development of artificial intelligence (AI), the demand for model training is also increasing.
[0003] For training scenarios with high computing power requirements, such as large model training scenarios, in order to improve the training efficiency of large models, relevant researchers have proposed using clusters to train large models. However, when using clusters to train large models, although clusters provide higher computing power and improve the model training efficiency to a certain extent, the failure of a single node in the cluster will cause the cluster to stop training operations, thereby affecting the model training efficiency. Therefore, when using clusters to train models, it is necessary to further improve the model training efficiency. Summary of the invention
[0004] The present disclosure provides a cluster-based training method, apparatus, device, and storage medium.
[0005] According to one aspect of the present disclosure, a cluster-based training method is provided, wherein the cluster includes a training node and a plurality of standby nodes for performing a model training task, including:
[0006] In response to a hardware failure in a training node, a target standby node is selected from a plurality of standby nodes, and a target training snapshot of a model training task in the training node is obtained; the target training snapshot includes training status data of the model training task;
[0007] Based on the container image of the model training program in the training node and the training status data, the target standby node is initialized so that the target standby node takes over the training node to continue to execute the model training task.
[0008] According to a second aspect of the present disclosure, a cluster-based training device is provided, wherein the cluster includes a training node and a plurality of standby nodes for performing a model training task, including:
[0009] A first response module is used to select a target standby node from multiple standby nodes in response to a hardware failure of a training node, and obtain a target training snapshot of a model training task in the training node; the target training snapshot includes training status data of the model training task;
[0010] An initialization module is used to initialize the target standby node based on the container image of the model training program in the training node and the training status data, so that the target standby node takes over the training node to continue to perform the model training task.
[0011] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0012] at least one processor; and
[0013] a memory communicatively connected to the at least one processor; wherein,
[0014] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the cluster-based training method as described in the first aspect.
[0015] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the cluster-based training method as described in the first aspect.
[0016] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement the steps of the cluster-based training method as described in the first aspect.
[0017] The present disclosure provides a cluster-based training method, device and electronic device, which have the following beneficial effects: in response to a hardware failure in a training node, the present disclosure selects a target standby node from multiple standby nodes, and obtains a target training snapshot of a model training task in the training node; the target training snapshot includes the training status data of the model training task; based on the container image and training status data of the model training program in the training node, the target standby node is initialized so that the target standby node takes over the training node to continue to perform the model training task. Based on this, the present disclosure can quickly initialize the target standby node based on the container image and the target training snapshot when a hardware failure occurs in the training node, so that the target standby node quickly takes over the training node to continue to perform the model training task, thereby shortening the fault recovery time and improving the model training efficiency.
[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0020] Figure 1 is a flow chart of a cluster-based training method provided according to an embodiment of the present disclosure;
[0021] Figure 2 is a flow chart of a cluster-based training method provided according to another embodiment of the present disclosure;
[0022] Figure 3 is a flow chart of a cluster-based training method provided according to another embodiment of the present disclosure;
[0023] Figure 4 A schematic diagram of the structure of a cluster-based training device provided according to an embodiment of the present disclosure;
[0024] Figure 5 It is a block diagram of an electronic device for implementing the cluster-based training method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] The cluster-based training method, device, and electronic device according to the embodiments of the present disclosure are described below with reference to the accompanying drawings.
[0027] It should be noted that the executor of the cluster-based training method of this embodiment is a cluster-based training device, and the cluster-based training device can be implemented by software and / or hardware and can be configured in an electronic device.
[0028] Figure 1 It is a flowchart of a cluster-based training method provided according to an embodiment of the present disclosure.
[0029] The cluster includes a training node and multiple standby nodes for executing model training tasks. The number of training nodes included in the cluster can be one or more. When the number of training nodes in the cluster is multiple, a model can be trained through multiple training nodes, and each training node undertakes a part of the training task of the model.
[0030] The training nodes and standby nodes may refer to nodes that have been stress tested, for example, nodes that have been stress tested by the communication unit, the computing unit, and the model end-to-end. The period of stress testing on the nodes may be determined by the user according to actual needs.
[0031] like Figure 1 As shown, the cluster-based training method includes:
[0032] Step 101, in response to a hardware failure in a training node, select a target backup node from multiple backup nodes, and obtain a target training snapshot of a model training task in the training node; the target training snapshot includes training status data of the model training task.
[0033] Optionally, the hardware failure may refer to a failure of hardware in the training node. As an example and not limitation, the hardware may refer to a graphics processing unit (GPU), a central processing unit (CPU), a network card, etc.
[0034] Optionally, selecting a target standby node from a plurality of standby nodes includes: randomly selecting a standby node from the plurality of standby nodes as the target standby node.
[0035] In order to improve node utilization and make full use of node resources, standby nodes are generally not idle, that is, standby nodes generally perform tasks other than model training tasks.
[0036] Optionally, selecting a target standby node from a plurality of standby nodes includes: obtaining a priority of the standby node, where the priority indicates the importance of other tasks; and selecting a target standby node from the plurality of standby nodes based on the priority.
[0037] Selecting the target standby node based on priority can, on the one hand, prevent other important tasks being executed in the standby node from being terminated, and on the other hand, ensure the high availability of the cluster and improve node utilization.
[0038] Assuming that the higher the priority, the more important the other tasks performed by the standby node are, when selecting the target standby node based on the priority, the target standby node can be selected in order from low to high priority. When there is no low-priority standby node or the low-priority standby node is unavailable, the high-priority standby node is selected as the target standby node.
[0039] As an example but not a limitation, assuming that the priorities corresponding to multiple backup nodes are A and B, respectively, and A is higher than B, when selecting a target backup node, a backup node with a priority of B may be preferentially selected as the target backup node, and when there is no backup node with a priority of B or the backup node with a priority of B is unavailable, a backup node with a priority of A may be selected as the target backup node. The number of backup nodes with a priority of A may be greater than the number of backup nodes with a priority of B.
[0040] For GPU hardware, if a GPU fails, it is likely to fail again during subsequent use. In order to reduce the overall failure frequency of the cluster and improve training efficiency, the present disclosure proposes that GPUs that have previously failed can be centrally deployed in one or more nodes, and such nodes can be used as low-priority backup nodes for priority scheduling.
[0041] Optionally, the present disclosure further includes: obtaining an estimated availability of the cluster and an estimated number of failed nodes per unit time; and determining the number of spare nodes according to the estimated availability and the estimated number.
[0042] In an optional implementation, Poisson fitting may be performed based on the estimated availability and the estimated quantity to obtain the number of standby nodes.
[0043] As an example and not a limitation, the number of spare nodes may be determined according to the following formula:
[0044]
[0045] Among them, α is the estimated availability, μ is the estimated number, and x is the number of backup nodes to be solved. Solving the above formula can obtain the number of backup nodes x.
[0046] The estimated number of faulty nodes per unit time can be obtained based on historical data statistics, or the estimated number of faulty nodes per unit time can be obtained by converting the estimated failure frequency of the cluster.
[0047] Determining the number of standby nodes based on the estimated availability and estimated quantity can enable the standby nodes to cover failed nodes and achieve high availability of the cluster.
[0048] Optionally, the target training snapshot may refer to a training snapshot of a model training task when a hardware failure occurs in a training node, or may refer to a training snapshot of a model training task at a certain moment before a hardware failure occurs in a training node.
[0049] The training status data of the model training task included in the target training snapshot may refer to data such as training data status, model parameters, and optimizer status. The training data status is used to indicate whether the training data has been used. Based on the training data status, the target standby node can directly train the model based on the unused training data to avoid repeated training of the model based on the same training data.
[0050] Step 102, based on the container image and training status data of the model training program in the training node, initialize the target standby node so that the target standby node takes over the training node to continue to perform the model training task.
[0051] The container image of the model training program can be prefabricated in advance. After the target standby node is determined, the container image can be deployed in the target standby node to improve the deployment speed of the model training program.
[0052] After the container image is deployed, data configuration can be performed on the target standby node based on the training status data.
[0053] The process of initializing the target standby node can be understood as seizing the resources of the target standby node to perform the model training task.
[0054] Initializing the target standby node based on the container image and training status data can automatically perform hardware failure recovery, shorten hardware failure recovery time, and improve hardware failure recovery efficiency.
[0055] In an embodiment of the present disclosure, in response to a hardware failure in a training node, a target standby node is selected from multiple standby nodes, and a target training snapshot of a model training task in the training node is obtained; the target training snapshot includes training status data of the model training task; based on the container image and training status data of the model training program in the training node, the target standby node is initialized so that the target standby node takes over the training node to continue to perform the model training task. Based on this, the present disclosure can perform extremely fast initialization of the target standby node based on the container image and the target training snapshot when a hardware failure occurs in the training node, so that the target standby node quickly takes over the training node to continue to perform the model training task, thereby shortening the fault recovery time and improving the model training efficiency.
[0056] Figure 2 It is a flowchart of a cluster-based training method provided according to another embodiment of the present disclosure.
[0057] like Figure 2 As shown, the cluster-based training method includes:
[0058] Step 201, in response to a hardware failure in a training node, select a target backup node from multiple backup nodes, and obtain at least one training snapshot of the model training task in the training node; the training snapshot is obtained from the training node according to a target period, and a target training snapshot is selected from at least one training snapshot.
[0059] After the training node starts executing the model training task, the training snapshot of the model training task is obtained regularly according to the target cycle, which can avoid data loss and invalid training caused by hardware failure of the training node.
[0060] The cluster also includes a storage node. Optionally, obtaining at least one training snapshot of the model training task in the training node includes: controlling the training node to save the training snapshot to the memory of the training node based on a first process according to a target cycle; controlling the training node to read the training snapshot from the memory based on a second process, and saving the training snapshot to the solid-state hard disk of the training node; controlling the training node to read the training snapshot from the solid-state hard disk based on a third process, and sending the training snapshot to the storage node.
[0061] If the training node saves the training snapshot to the memory, SSD and storage node in sequence through only one process, the process can only return to execute other instructions after storing the training snapshot to the storage node, and the overall execution efficiency of the program is low.
[0062] By asynchronously acquiring or storing training snapshots through multiple processes, blocking of program execution can be avoided, thereby improving the concurrent performance of the program.
[0063] Alternatively, the target period may be determined empirically by the user.
[0064] Optionally, the target period can also be determined by:
[0065] Obtain a single estimated save cost of a training snapshot of a model training task; construct a total save cost function with a target period as a variable based on the set duration and the single estimated save cost; construct an invalid training cost function with a target period as a variable; determine a target cost function based on the total save cost function and the invalid training cost function; solve the target cost function to determine the target period.
[0066] As an example but not limitation, a single estimated saving overhead may refer to the saving overhead corresponding to controlling the training node to save the training snapshot to the storage node through three processes, or may refer to the overhead of saving the training snapshot through other saving methods.
[0067] As an example but not limitation, the set duration may refer to the average duration from the start of the model training task by the training node to the interruption of the model training task, or it may be the duration determined by the user based on experience.
[0068] As an example and not a limitation, the target period may be determined by the following formula:
[0069]
[0070] Where H(t) is the target cost function, H 1 (t) is the total storage cost function, T is the set duration, c is the single estimated storage cost, t is a variable, that is, the target period, t is greater than 0, H 2(t) is the invalid training cost function. Invalid training may refer to the training between the time when the last training snapshot is obtained before the hardware failure occurs and the time when the hardware failure occurs. It should be noted that H 2 The coefficient of the variable t in (t) can be determined according to actual conditions.
[0071] By solving the minimum value of H(t), the target cycle can be obtained. The above total storage overhead and invalid training overhead may refer to time overhead.
[0072] The target period obtained by solving the above formula is an optimal period, which can balance the total storage overhead of training snapshots and the invalid training overhead, and minimize the overhead during model training.
[0073] As an example but not a limitation, the target period may also be determined by the following formula:
[0074]
[0075] Among them, f failure The estimated failure frequency of a training node during model training tasks.
[0076] Optionally, selecting a target training snapshot from at least one training snapshot includes: selecting the target training snapshot from at least one training snapshot according to a corresponding training effect, wherein the training effect may be determined based on a loss value corresponding to the training snapshot.
[0077] Selecting the training snapshot with the best corresponding training effect as the target training snapshot can enable the target standby node to continue to perform model training tasks based on the best model training results before the hardware failure occurs, thereby improving model training efficiency and reducing model training costs.
[0078] The training snapshot corresponds to an acquisition time. Optionally, a target training snapshot is selected from at least one training snapshot, including: obtaining a fault time corresponding to a hardware fault of a training node; and selecting a training snapshot with the shortest time interval between the corresponding acquisition time and the fault as the target training snapshot.
[0079] Selecting the last training snapshot obtained before the hardware failure occurs as the target training snapshot enables the target standby node to continue to perform the model training task based on the latest model training results, making the model training continuous and reducing invalid training overhead.
[0080] Step 202: Perform communication bandwidth verification and / or silent data error verification on the target standby node.
[0081] Optionally, performing communication bandwidth verification on the target standby node refers to verifying whether the communication bandwidth corresponding to the target standby node is greater than or equal to a bandwidth threshold, if the communication bandwidth is greater than or equal to the bandwidth threshold, the communication bandwidth verification passes, if the communication bandwidth is less than the bandwidth threshold, the communication bandwidth verification fails. The communication bandwidth may refer to a collective communication bandwidth.
[0082] Optionally, performing silent data error checking on the target standby node refers to checking whether the target standby node has a silent data error. If no silent data error exists, the silent data error check passes; if a silent data error exists, the silent data error check fails.
[0083] As an example but not limitation, the training node can be made to perform multiple specific operations and compare multiple operation results. If the multiple operation results are the same, the training node does not have a silent data error. If there are at least two different operation results among the multiple operation results, the training node has a silent data error.
[0084] By performing communication bandwidth verification and / or silent data error checking on the target standby node, risk monitoring of the performance of the target standby node can be performed to avoid introducing nodes with poor performance or failures, and reduce the probability of subsequent failures of the target standby node.
[0085] Step 203, in response to the communication bandwidth verification passing and / or the silent data error check passing, the target standby node is initialized based on the container image and training status data of the model training program in the training node.
[0086] For the relevant contents in step 203, reference may be made to the relevant description in step 102, which will not be repeated here.
[0087] In an embodiment of the present disclosure, in response to a hardware failure in a training node, a target standby node is selected from a plurality of standby nodes, and at least one training snapshot of a model training task in the training node is obtained; the training snapshot is obtained from the training node according to a target period, and a target training snapshot is selected from at least one training snapshot; communication bandwidth verification and / or silent data error checking are performed on the target standby node; in response to the communication bandwidth verification and / or silent data error checking passing, the target standby node is initialized based on the container image and training status data of the model training program in the training node. Based on this, the present disclosure obtains at least one training snapshot to avoid data loss and invalid training problems caused by hardware failure of the training node. In addition, communication bandwidth verification and / or silent data error checking are performed on the target standby node, and risk monitoring of the performance of the target standby node can be performed to avoid the introduction of nodes with poor performance or failures.
[0088] Figure 3It is a flowchart of a cluster-based training method provided according to another embodiment of the present disclosure.
[0089] like Figure 3 As shown, the cluster-based training method includes:
[0090] Step 301: Perform fault monitoring on the training node.
[0091] Optionally, the model training program in the training node may be monitored, that is, whether the model training program exits due to a failure may be monitored.
[0092] Step 302 , in response to acquiring the fault code of the training node, based on the mapping relationship between the fault code and the fault type, determine the fault type of the training node, where the fault type includes a hardware fault or a software fault.
[0093] When a fault is detected in the model training program, the corresponding fault code is obtained, and the fault type corresponding to the fault code is searched based on the mapping relationship to determine whether the training node is a hardware fault or a software fault.
[0094] The present disclosure can realize self-diagnosis of hardware and software faults. No matter it is a hardware fault or a software fault, its manifestation in the model training program is program fault exit. Therefore, rapid perception of faults can be achieved based on fault codes.
[0095] Step 303, in response to a hardware failure in the training node, select a target backup node from multiple backup nodes, and obtain a target training snapshot of the model training task in the training node.
[0096] Compared with the fault monitoring method of scanning the hardware error log at a certain period, hardware fault monitoring based on fault codes is more timely.
[0097] In addition, there is a mapping relationship between fault codes and hardware faults. Therefore, faulty hardware can be located based on the fault codes.
[0098] Step 304, in response to a software failure in the training node, determining an abnormal process in the training node, and restarting the abnormal process.
[0099] There is also a mapping relationship between fault codes and software faults. Therefore, the abnormal process can be determined based on the fault code, and the abnormal process can be restarted.
[0100] The present disclosure can respond to different fault types and select corresponding fault recovery methods, thereby shortening the fault recovery time and ensuring the training stability of the cluster.
[0101] Optionally, the fault type may be mapped to a fault level, and each fault level corresponds to a fault recovery method. As an example and not a limitation, the fault level corresponding to a software fault may be M, and the fault level corresponding to a hardware fault may be N. The fault level of a software fault is lower than the fault level of a hardware fault, and therefore, the fault level M may be lower than the fault level N. Among them, the fault recovery method corresponding to the fault level M is to restart the abnormal process, and the fault recovery method corresponding to the fault level N is to select a target standby node and initialize the target standby node.
[0102] Specifically, the fault level may be determined based on the fault code, and then a corresponding fault recovery method may be selected based on the fault level.
[0103] Optionally, when the target standby node fails to initialize or the abnormal process fails to restart, the cluster can be scheduled, that is, another cluster takes over the current cluster to continue training the model.
[0104] Based on the cluster-based training method disclosed in the present invention, the effective training rate of the cluster for the model is effectively improved. The effective training rate can be used to indicate the proportion of time actually contributed to the model parameter update within the expected model training time period. Optionally, the effective training rate of the cluster can be estimated by the following formula:
[0105] ρ=α-F failure *C recovery -F checkpoint *c
[0106] Where ρ is the effective training rate, α is the estimated availability of the cluster, and Ff ailure is the estimated failure frequency of the cluster, C recovery is the estimated single fault recovery time ratio, C recovery It can be understood as the ratio of a single fault recovery time to the estimated training time of the model training task. checkpoint is the estimated saving frequency of training snapshots, and c is the estimated single saving cost.
[0107] It should be noted that after the training node completes the model training task, the actual value of each parameter in the above formula can be obtained to determine the actual effective training rate of the cluster.
[0108] The effective training time can be determined based on the effective training rate and training time. The total throughput data volume of the cluster when executing the model training task can be determined based on the effective training time and the throughput data volume per unit time.
[0109] In addition to training nodes and backup nodes, the cluster also includes training service nodes, such as scheduling nodes for selecting target backup nodes and monitoring nodes for fault monitoring of training nodes. For training service nodes, fault monitoring of training service nodes can be performed through the communication status between training service nodes.
[0110] In the embodiment of the present disclosure, the training node is monitored for faults, and in response to obtaining the fault code of the training node, the fault type of the training node is determined based on the mapping relationship between the fault code and the fault type; in response to a hardware fault in the training node, a target spare node is selected from multiple spare nodes, and a target training snapshot of the model training task in the training node is obtained; in response to a software fault in the training node, an abnormal process in the training node is determined, and the abnormal process is restarted. Based on this, the present disclosure can select corresponding fault recovery methods in response to different fault types to shorten the fault recovery time.
[0111] Figure 4 1 is a schematic diagram of a cluster-based training device according to an embodiment of the present disclosure. The cluster includes a training node for executing a model training task and multiple standby nodes.
[0112] like Figure 4 As shown, the cluster-based training device 400 includes:
[0113] The first response module 401 is used to select a target standby node from multiple standby nodes in response to a hardware failure of the training node, and obtain a target training snapshot of the model training task in the training node; the target training snapshot includes training status data of the model training task;
[0114] The initialization module 402 is used to initialize the target standby node based on the container image and training status data of the model training program in the training node, so that the target standby node takes over the training node to continue to perform the model training task.
[0115] Optionally, the first response module 401 is used to:
[0116] Obtain at least one training snapshot of the model training task in the training node; the training snapshot is obtained from the training node according to the target period;
[0117] From at least one training snapshot, select a target training snapshot.
[0118] Optionally, the device further comprises:
[0119] The cost acquisition module is used to obtain the single estimated saving cost of the training snapshot of the model training task;
[0120] The first construction module is used to construct a total storage cost function with a target period as a variable according to a set duration and a single estimated storage cost;
[0121] The second building module is used to build an invalid training cost function with the target period as a variable;
[0122] A function determination module, used for determining a target cost function according to a total preservation cost function and an invalid training cost function;
[0123] The solution module is used to solve the target cost function and determine the target period.
[0124] Optionally, the cluster includes a storage node, and the first response module 401 is used to:
[0125] According to the target period, controlling the training node to save the training snapshot to the memory of the training node based on the first process;
[0126] Controlling the training node to read the training snapshot from the memory based on the second process, and saving the training snapshot to the solid state drive of the training node;
[0127] The control training node reads the training snapshot from the solid state drive based on the third process, and sends the training snapshot to the storage node.
[0128] Optionally, the training snapshot corresponds to an acquisition time, and the first response module 401 is used to:
[0129] Get the failure time corresponding to the hardware failure of the training node;
[0130] The training snapshot with the shortest acquisition time and fault time interval is selected as the target training snapshot.
[0131] Optionally, the standby node is used to perform other tasks besides the model training task, and the first response module 401 is used to:
[0132] Get the priority of the standby node, which indicates the importance of other tasks;
[0133] Based on the priority, a target standby node is selected from multiple standby nodes.
[0134] Optionally, the initialization module 402 is used to:
[0135] Perform communication bandwidth verification and / or silent data error checking on the target standby node;
[0136] In response to the communication bandwidth verification passing and / or the silent data error check passing, the target standby node is initialized.
[0137] Optionally, the device further comprises:
[0138] The second response module is used to determine an abnormal process in the training node in response to a software failure in the training node, and restart the abnormal process.
[0139] Optionally, the device further comprises:
[0140] A fault monitoring module is used to monitor the faults of training nodes;
[0141] The type determination module is used to determine the fault type of the training node in response to obtaining the fault code of the training node based on the mapping relationship between the fault code and the fault type, where the fault type includes a hardware fault or a software fault.
[0142] Optionally, the device further comprises:
[0143] The data acquisition module is used to obtain the estimated availability of the cluster and the estimated number of failed nodes per unit time;
[0144] The quantity determination module is used to determine the quantity of spare nodes according to the estimated availability and the estimated quantity.
[0145] It should be noted that the aforementioned explanation of the cluster-based training method is also applicable to the cluster-based training device of this embodiment, and will not be repeated here.
[0146] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0147] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device 500 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0148] like Figure 5As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 to a RAM (Random Access Memory) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.
[0149] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0150] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as a cluster-based training method. For example, in some embodiments, the cluster-based training method may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the cluster-based training method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform a cluster-based training method in any other appropriate manner (eg, by means of firmware).
[0151] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor that may be a dedicated or general-purpose programmable processor that may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0152] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0153] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include electrical connections based on one or more lines, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0154] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0155] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0156] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0157] It should be noted that artificial intelligence is a discipline that studies how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.
[0158] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0159] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A cluster-based training method, wherein the cluster includes a training node for performing a model training task and a plurality of standby nodes, include: In response to a hardware failure in a training node, selecting a target standby node from a plurality of standby nodes, and obtaining a target training snapshot of a model training task in the training node; The target training snapshot includes the training status data of the model training task; Initializing the target standby node based on the container image of the model training program in the training node and the training status data, so that the target standby node takes over the training node to continue to perform the model training task; The target training snapshot is obtained from the training node according to the target period, and the method further includes: Obtain a single estimated saving cost of a training snapshot of the model training task; According to the set duration and the single estimated storage cost, construct a total storage cost function with the target period as a variable; Constructing an invalid training cost function with the target period as a variable; Determining a target cost function according to the total preservation cost function and the invalid training cost function; The target cost function is solved to determine the target period.
2. The method according to claim 1, wherein obtaining a target training snapshot of the model training task in the training node, include: Obtain at least one training snapshot of the model training task in the training node; The training snapshot is obtained from the training node according to the target period; The target training snapshot is selected from at least one training snapshot.
3. The method according to claim 2, wherein the cluster includes a storage node, and the obtaining of at least one training snapshot of the model training task in the training node, include: According to the target period, controlling the training node to save the training snapshot to the memory of the training node based on the first process; Controlling the training node to read the training snapshot from the memory based on a second process, and saving the training snapshot to a solid state drive of the training node; The training node is controlled to read the training snapshot from the solid state drive based on a third process, and the training snapshot is sent to the storage node.
4. The method according to claim 2, wherein the training snapshot corresponds to an acquisition time, and the target training snapshot is selected from at least one training snapshot, include: Obtaining a fault time corresponding to the hardware fault of the training node; A training snapshot whose corresponding acquisition time is the shortest time interval with the fault time is selected as the target training snapshot.
5. According to the method of claim 1, the standby node is used to perform tasks other than the model training task, and the target standby node is selected from the plurality of standby nodes. include: Obtaining a priority of the standby node, where the priority indicates the importance of the other tasks; Based on the priority, the target standby node is selected from the multiple standby nodes.
6. The method according to claim 1, wherein the target standby node is initialized. include: Performing communication bandwidth verification and / or silent data error checking on the target standby node; In response to the communication bandwidth verification passing and / or the silent data error check passing, the target standby node is initialized.
7. The method according to claim 1, further comprising: include: In response to a software failure in the training node, an abnormal process in the training node is determined, and the abnormal process is restarted.
8. The method according to claim 1 or 7, further comprising: include: Performing fault monitoring on the training node; In response to acquiring the fault code of the training node, the fault type of the training node is determined based on a mapping relationship between the fault code and the fault type, where the fault type includes a hardware fault or a software fault.
9. The method according to any one of claims 1 to 7, further comprising: include: Obtaining the estimated availability of the cluster and the estimated number of failed nodes per unit time; The number of the standby nodes is determined according to the estimated availability and the estimated number.
10. A cluster-based training device, wherein the cluster includes a training node for executing a model training task and a plurality of standby nodes, include: A first response module, configured to select a target standby node from a plurality of standby nodes in response to a hardware failure of a training node, and obtain a target training snapshot of a model training task in the training node; The target training snapshot includes the training status data of the model training task; An initialization module, used to initialize the target standby node based on the container image of the model training program in the training node and the training status data, so that the target standby node takes over the training node to continue to perform the model training task; Wherein, the target training snapshot is obtained from the training node according to the target period, and the device further includes: A cost acquisition module, used to obtain a single estimated saving cost of a training snapshot of the model training task; A first construction module is used to construct a total storage cost function with the target period as a variable according to a set duration and the single estimated storage cost; A second construction module is used to construct an invalid training cost function with the target period as a variable; A function determination module, used to determine a target cost function according to the total preservation cost function and the invalid training cost function; A solution module is used to solve the target cost function and determine the target period.
11. The device according to claim 10, in, The first response module is used to: Obtain at least one training snapshot of the model training task in the training node; the training snapshot is obtained from the training node according to the target period; The target training snapshot is selected from at least one training snapshot.
12. The device according to claim 11, in, The cluster includes a storage node, and the first response module is used to: According to the target period, controlling the training node to save the training snapshot to the memory of the training node based on the first process; Controlling the training node to read the training snapshot from the memory based on a second process, and saving the training snapshot to a solid state drive of the training node; The training node is controlled to read the training snapshot from the solid state drive based on a third process, and the training snapshot is sent to the storage node.
13. The device according to claim 11, in, The training snapshot corresponds to an acquisition time, and the first response module is used to: Obtaining a fault time corresponding to the hardware fault of the training node; A training snapshot whose corresponding acquisition time is the shortest time interval with the fault time is selected as the target training snapshot.
14. The device according to claim 10, in, The standby node is used to perform tasks other than the model training task, and the first response module is used to: Obtaining a priority of the standby node, where the priority indicates the importance of the other tasks; Based on the priority, the target standby node is selected from the multiple standby nodes.
15. The device according to claim 10, in, The initialization module is used to: Performing communication bandwidth verification and / or silent data error checking on the target standby node; In response to the communication bandwidth verification passing and / or the silent data error check passing, the target standby node is initialized.
16. The device according to claim 10, in, The device also includes: The second response module is used to determine an abnormal process in the training node in response to a software failure in the training node, and restart the abnormal process.
17. The device according to claim 10 or 16, in, The device also includes: A fault monitoring module, used for performing fault monitoring on the training node; The type determination module is used to determine the fault type of the training node in response to obtaining the fault code of the training node based on the mapping relationship between the fault code and the fault type, wherein the fault type includes a hardware fault or a software fault.
18. The device according to any one of claims 10 to 16, in, The device also includes: A data acquisition module, used to obtain the estimated availability of the cluster and the estimated number of failed nodes per unit time; The quantity determination module is used to determine the quantity of the standby nodes according to the estimated availability and the estimated quantity.
19. An electronic device, include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
20. A non-transitory computer-readable storage medium storing computer instructions, in, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.
21. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Auto snapshot method and device
CN101359301A
Online training method, device and system and computer readable storage medium
CN109635948A