Model training recovery method and electronic equipment
By updating the communication topology relationship, the long-term interruption problem caused by GPU failure in large model training is solved, efficient training recovery is achieved, and communication reconstruction time is reduced.
Patent Information
- Application Number
- CN202510560473.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-18
AI Technical Summary
During the training process of large-scale model, when some GPU devices fail, the existing technology needs to restart the entire training task, resulting in long-term interruptions, and the communication group reconstruction process takes a long time, affecting the training efficiency.
By obtaining the task process and device information of the processing device, updating the communication topology relationship, re-establishing the communication connection between virtual devices, avoiding restarting the entire training task and reducing interrupt time.
It significantly reduces the interrupt time of training tasks, improves the efficiency of communication reconstruction during training, and ensures that communication between unfailed devices is not affected.
Smart Images

Figure CN120336068A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of model training, and more specifically, to a method for restoring model training and an electronic device. Background Art
[0002] When training a large model, the training task can be executed by GPU devices. When some of the GPU devices fail, the communication connections between the devices will be interrupted, and the entire task needs to be restarted to recover from the failure, resulting in a long interruption of the training task. Summary of the Invention
[0003] In view of this, the present disclosure provides a method for restoring model training and an electronic device.
[0004] The first aspect of the present disclosure provides a method for restoring model training, including: In response to a first processing device among a plurality of processing devices failing, obtaining the task processes of the plurality of processing devices executing the large model training task, and the device information of the plurality of processing devices; Based on the task processes and device information, updating a first communication topology relationship of the plurality of processing devices to obtain a second communication topology relationship; the communication topology relationship represents the communication connection manner between the virtual devices corresponding to the respective processing devices; Sending the second communication topology relationship to the processing devices that have not failed.
[0005] According to an embodiment of the present disclosure, before responding to a first processing device among a plurality of processing devices failing, the method further includes: Based on the heartbeat messages sent by the plurality of processing devices, determining whether a first processing device among the plurality of processing devices has failed.
[0006] According to an embodiment of the present disclosure, based on the heartbeat messages sent by the plurality of processing devices, determining whether a first processing device among the plurality of processing devices has failed includes: Receiving the heartbeat messages sent by the plurality of processing devices at a preset time interval; When the heartbeat message sent by the first processing device meets a preset message type, confirming that the first processing device has failed; Or, when the first processing device does not send a heartbeat message at the preset time interval, confirming that the first processing device has failed.
[0007] According to an embodiment of the present disclosure, based on the task processes and device information, updating a first communication topology relationship of the plurality of processing devices to obtain a second communication topology relationship includes: According to the task process and device information of the first processing device, adjusting the device attribute of the first processing device to a closed attribute; Adjust the device attribute of the second processing device to the enabled attribute according to the task processes and device information of multiple processing devices; the second processing device is a processing device without a fault. Update the first communication topology relationship of multiple processing devices based on the device attributes of the multiple processing devices to obtain a second communication topology relationship.
[0008] According to an embodiment of the present disclosure, the device information includes the device addresses, network names, and communication connection information of multiple processing devices.
[0009] A second aspect of the present disclosure provides another method for resuming model training, including: Obtain a second communication topology relationship; the second communication topology relationship is obtained by updating the first communication topology relationship based on the task processes of multiple processing devices performing large model training tasks and the device information of the multiple processing devices; the communication topology relationship characterizes the communication connection methods between virtual devices corresponding to each processing device. Update the communication connection with other processing devices based on the second communication topology relationship, and obtain task data for executing the task process from other processing devices. Execute the task process based on the task data.
[0010] According to an embodiment of the present disclosure, updating the communication connection with other processing devices based on the second communication topology relationship includes: Confirm a first processing device with a closed device attribute and a second processing device with an enabled device attribute according to the second communication topology relationship. Establish a communication connection with the second processing device and delete the communication connection with the first processing device.
[0011] According to an embodiment of the present disclosure, executing the task process based on the task data includes: Obtain the data volume of the task data. Compare the data volume of the task data with the storage capacity of the local memory. Write the task data into the local memory based on the comparison result.
[0012] According to an embodiment of the present disclosure, writing the task data into the local memory based on the comparison result includes: When the data volume is less than the storage capacity, save the task data to the local memory. Write the task data from the local memory.
[0013] According to an embodiment of the present disclosure, writing the task data into the local memory based on the comparison result further includes: When the data volume is greater than or equal to the storage capacity, save the first part of the task data to the local memory; save the second part of the task data to a pre-configured database. Write the first part of the task data from local memory; at the same time, write the second part of the task data to local memory.
[0014] A third aspect of the present disclosure provides an electronic device, including a transceiver and a processor; and The processor is coupled to the transceiver; the processor is configured to: in response to a failure of a first processing device among a plurality of processing devices, obtain the task processes of the plurality of processing devices for performing the large model training task, and the device information of the plurality of processing devices; Based on the task processes and the device information, update the first communication topology relationship of the plurality of processing devices to obtain a second communication topology relationship; the communication topology relationship represents the communication connection manner between the virtual devices corresponding to each processing device; Send the second communication topology relationship to the processing devices that have not failed.
[0015] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings: Figure 1 Schematically shows one of the scenario diagrams of a method for recovering large model training according to an embodiment of the present disclosure; Figure 2 Schematically shows the flowchart of a method for recovering large model training according to an embodiment of the present disclosure; Figure 3 Schematically shows the communication topology diagram between virtual devices according to an embodiment of the present disclosure; Figure 4 Schematically shows the flowchart of a method for updating the communication topology relationship according to an embodiment of the present disclosure; Figure 5 Schematically shows another scenario diagram of a method for recovering large model training according to an embodiment of the present disclosure; Figure 6 Schematically shows another flowchart of a method for recovering large model training according to an embodiment of the present disclosure; Figure 7 Schematically shows the principle diagram of establishing communication connections between multiple devices according to an embodiment of the present disclosure; Figure 8 Schematically shows the block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0018] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0019] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0020] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0021] The present disclosure provides a method for restoring model training. Before introducing the specific solutions of the embodiments of the present disclosure, the related technologies involved in the present disclosure will be described first.
[0022] When training a large model, usually a cluster system composed of multiple Graphics Processing Units (GPUs) is used to execute the training task. When a software or hardware failure occurs in the cluster system, the entire training task needs to be restarted to resume the large model training. Since the entire recovery process takes a long time, it will cause a long interruption in the large model training task, affecting model training.
[0023] In the related art, when a GPU device fails during the execution of a certain task in the process of training a large model, other task processes are first "frozen" and will not exit but continue. At the same time, since the task process restarted due to the failure will restart on healthy machines in the cluster again, the affected communication group will be rebuilt after the restart. In addition, a new application programming interface (API) for adjusting the collective communication group is added, and then the TrainMover controller will coordinate all the reconstruction processes of the collective communication groups uniformly. Finally, the state of the job (model parameters, optimizer, etc.) is restored based on the neighbor node backup or the Checkpoint data in the remote storage system. The following problems exist in the above solution: The three major links of task restart, communication group reconstruction, and state recovery based on Checkpoint data are executed serially, and the end-to-end time consumption is relatively long; the communication group communication construction process depends on the modification of the NVIDIA Collective Communications Library, and an outsourcing mechanism is required to coordinate the entire reconstruction process. The actual overhead is close to the overhead of cancellation and reconstruction, and the portability is relatively weak.
[0024] The following introduces the relevant terms in the embodiments of the present disclosure.
[0025] The NVIDIA Collective Communications Library (NCCL) is a high-performance collective communication library developed by NVIDIA, optimized for distributed computing in multi-GPU and multi-node environments, aiming to improve the communication efficiency between GPU clusters.
[0026] A GPU cluster is a computer cluster in which each node is equipped with a graphics processing unit GPU. By using the general computing on the graphics processing unit, the computing power of modern GPUs can be utilized, and very fast computing can be performed using the GPU cluster.
[0027] Figure 1 Figure 1 schematically shows one of the scenario diagrams of a recovery method for large model training according to an embodiment of the present disclosure.
[0028] As Figure 1 shown, each processing device communicates with other processing devices in the GPU cluster 102 through a pre-built communication connection. The processing devices include, but are not limited to, electronic devices such as GPUs and CPUs that can be used to execute model training tasks. In the present disclosure, a GPU device is used as an exemplary illustration.
[0029] In the embodiments of the present disclosure, each processing device is a physical GPU device, and there is a corresponding virtual device. The virtual device establishes a mapping relationship with the processing device through the processing device ID corresponding to the virtual device ID. In the embodiments of the present disclosure, the virtual device ID does not change before and after a failure occurs.
[0030] In the embodiments of the present disclosure, as Figure 1 shown, each slave node includes at least one virtual device and a corresponding processing device. The slave node manages the processing device that executes a certain task process, and the slave node sends the task process and device information of each processing device to the master node, so that the master node can update the communication topology relationship between multiple virtual devices.
[0031] In the embodiments of the present disclosure, a method for restoring large model training is provided and applied to the master node.
[0032] Figure 2 Schematically shows a flowchart of a method for restoring large model training according to an embodiment of the present disclosure.
[0033] Specifically, as Figure 2 shown, the method includes operations S201 to S203.
[0034] Operation S201, in response to a first processing device among multiple processing devices having a failure, obtains the task processes of the multiple processing devices executing the large model training task, and the device information of the multiple processing devices.
[0035] Operation S202, based on the task processes and device information, updates the first communication topology relationship of the multiple processing devices to obtain a second communication topology relationship; the communication topology relationship represents the communication connection manner between the virtual devices corresponding to the respective processing devices.
[0036] Operation S203, sends the second communication topology relationship to the processing devices that have not failed.
[0037] In the embodiments of the present disclosure, the first processing device is at least one processing device confirmed by the master node to have failed.
[0038] In operation S201, when performing large model training, multiple processing devices are usually used to execute the training task in parallel. During the execution process, if a certain processing device fails, the task process executed by the processing device will be interrupted. At this time, the communication connection between the multiple processing devices will change. At this time, the master node obtains the task processes of the multiple processing devices executing the large model training task, and the device information of the multiple processing devices through the slave node.
[0039] In operation S202, since the master node records the first communication topology relationship of multiple processing devices before a failure occurs, after the failure occurs, by adjusting the connection relationships of the virtual devices in the first communication topology relationship, a second communication topology relationship is obtained.
[0040] In operation S203, the second communication topology relationship is sent from the slave node to the processing devices that have not failed. Through the second communication topology relationship, the communication topology relationship between each virtual device can be re-established, so that according to the mapping relationship between the virtual device and the processing device, the communication connections of multiple processing devices can be restored.
[0041] Through the above method, only the communication topology relationship of the faulty device is updated, avoiding the problem of restarting the entire job in the traditional method, and significantly reducing the interruption time of the training task. At the same time, by reconstructing the communication topology relationship, it is ensured that the communication between the non-faulty processing devices is not affected, improving the efficiency of reconstructing the communication process during training.
[0042] According to an embodiment of the present disclosure, the device information includes the device addresses, network names, and communication connection information of multiple processing devices.
[0043] Exemplarily, the device address is the device ID of each processing device, the network name is the address name of the communication group where each processing device is located, and the communication connection information is the information of other processing devices connected to each processing device.
[0044] Before operation S201, the method further includes operation S204.
[0045] Operation S204 constructs the first communication topology relationship of multiple processing devices.
[0046] Specifically, after the training process is successfully deployed to the processing devices in each slave node, the information of each training process can be determined. The information of the training process includes, but is not limited to, the local identifier, global identifier, processing device used, and the mapping relationship between the processing device and the virtual device in the entire training task, and the network information of the slave node where it is located. Each slave node reports the information of the corresponding training process to the master node, and the master node constructs the first communication topology relationship after receiving the information sent by all slave nodes.
[0047] Exemplarily, the mapping relationship between the processing device and the virtual device is determined by a device mapping table. Among them, the device ID of the processing device is the IP address and device number of the local host where the GPU device is located; the device ID of the virtual device is the collective communication number where the GPU device is located. The collective communication number is the device number of each virtual device in different communication connection methods. For example, the device ID of the virtual device is dp0-tp0-pp1, indicating that its number in the data parallel (DP) communication method is 0, its number in the tensor parallel (TP) communication method is 0, and its number in the pipeline parallel (PP) communication method is 1. For example, the device ID of the processing device is node1-gpu0, indicating that it is the 0th device on the slave node.
[0048] Based on the above, a device mapping table as shown in Table 1 is constructed to further determine the mapping relationship between the virtual device and the processing device.
[0049] Table 1 Device Mapping Table Device ID of the virtual device Device ID of the processing device dp0-tp0-pp0 node1-gpu0 dp0-tp1-pp0 node1-gpu1 dp1-tp0-pp0 node1-gpu2 dp1-tp1-pp0 node1-gpu3 dp0-tp0-pp1 node2-gpu0 dp0-tp1-pp1 node2-gpu1 dp1-tp0-pp1 node2-gpu2 dp1-tp1-pp1 node2-gpu3 dp0-tp0-pp2 node3-gpu0 dp0-tp1-pp2 node3-gpu1 dp1-tp0-pp2 node3-gpu2 dp1-tp1-pp2 node3-gpu3 dp0-tp0-pp3 node4-gpu0 dp0-tp1-pp3 node4-gpu1 dp1-tp0-pp3 node4-gpu2 dp1-tp1-pp3 node4-gpu3 Figure 3 Schematically shows a communication topology diagram between virtual devices according to an embodiment of the present disclosure.
[0050] As Figure 3 shown, based on the device information of the processing device collected and the pre-defined communication group, an undirected graph is constructed to represent the communication relationship between virtual devices. Nodes in the graph represent virtual devices, and edges represent communication connections between virtual devices.
[0051] Exemplarily, it includes four slave nodes, namely node1, node2, node3, and node4. Each slave node contains the identification information of the virtual device (such as device ID, network name, IP address, etc.) and its role in the communication group (such as DP0, TP0, PP0, etc., respectively representing the first virtual device in the data parallel group, model parallel group, and pipeline parallel group).
[0052] As Figure 3 shown, each edge represents a communication connection between two virtual devices. To more intuitively represent the relationship between different communication groups, different colors can be used to mark different communication groups in the communication topology diagram.
[0053] According to an embodiment of the present disclosure, before a first processing device fails among multiple processing devices, the method further includes operation S205.
[0054] Operation S205: Based on the heartbeat messages sent by multiple processing devices, determine whether a first processing device fails among the multiple processing devices.
[0055] In operation S205, the master node periodically obtains heartbeat messages through the slave node where the processing device is located. When the obtained heartbeat message does not match the preset message type, it can be determined that a first processing device among the multiple processing devices has failed.
[0056] According to an embodiment of the present disclosure, determining whether a first processing device among multiple processing devices has failed based on heartbeat messages sent by the multiple processing devices includes: Receiving heartbeat messages sent by the multiple processing devices according to a preset time interval; confirming that the first processing device has failed when the heartbeat message sent by the first processing device meets the preset message type; or, confirming that the first processing device has failed when the first processing device does not send a heartbeat message according to the preset time interval.
[0057] Exemplarily, in distributed model training, the master node monitors the running status of processing devices in multiple slave nodes through a heartbeat mechanism. When a failure occurs in slave node 1 (such as process crash, network interruption, hardware damage), the master node will not be able to receive heartbeat messages regularly. Therefore, it is confirmed that all processing devices in the entire slave node 1 have failed.
[0058] Exemplarily, when a certain processing device in slave node 1 fails, the device information of the processing device is sent to the master node through the message field carried in the heartbeat message. After the master node receives the device information of the processing device, it is confirmed that the processing device has failed.
[0059] Judging whether a failure occurs during the execution of the task process through the type of the heartbeat message and the time interval for obtaining the heartbeat message can more accurately identify the failure of the processing device and reduce misjudgment and missed judgment.
[0060] Next, a method for updating the communication topology relationship after a processing device fails in the embodiments of the present disclosure will be described.
[0061] Figure 4 A flowchart of a method for updating a communication topology relationship according to an embodiment of the present disclosure is schematically shown.
[0062] As Figure 4 shown, based on the task process and device information, updating the first communication topology relationship of multiple processing devices to obtain a second communication topology relationship includes operations S401 to S403.
[0063] Operation S401, adjusting the device attribute of the first processing device to a closed attribute according to the task process and device information of the first processing device; Operation S402: Adjust the device attribute of the second processing device to the enabled attribute according to the task processes and device information of multiple processing devices; the second processing device is a processing device that has not failed. Operation S403: Update the first communication topology relationship of multiple processing devices based on the device attributes of the multiple processing devices to obtain a second communication topology relationship.
[0064] Specifically, based on the task process and device information of the failed processing device, in the first communication topology relationship, the topological point where the virtual device corresponding to the failed processing device is located can be determined, and the value of the device attribute is adjusted to the disabled attribute, indicating that the virtual device corresponding to this topological point is not used for the execution of the task process. At the same time, the processing devices that have not failed are determined from the cluster of multiple processing devices, and the value of the device attribute is adjusted to the enabled attribute, indicating that the virtual device corresponding to this processing device will be used for the execution of the task process.
[0065] By dynamically adjusting the device attributes of each processing device in the communication topology relationship according to the task process and device information of the failed processing device, so as to shut down the faulty device and reallocate resources to healthy devices, it is ensured that the communication connections of multiple processing devices represented by the communication topology relationship can be correctly applied to execute the task process.
[0066] The process of updating the first communication topology relationship of multiple processing devices to obtain a second communication topology relationship will be described in detail below.
[0067] Figure 5 Schematically shows a second scenario diagram of a large model training recovery method according to an embodiment of the present disclosure.
[0068] Exemplarily, as Figure 5 shown, after the processing devices 1021 to 1024 in the slave nodes fail and go down, in the first communication topology relationship, the topological point where the virtual device corresponding to the failed processing device is located is determined, and at the same time, the device attribute values of the processing devices 1021 to 1024 are adjusted to the disabled attribute "FLASE", and the processing devices that have not failed are determined from the cluster of multiple processing devices, such as the processing device 1025, and the device attribute value is adjusted to the enabled attribute "TURE", and the device information of the processing device 1025 and the corresponding virtual device is updated to facilitate the subsequent recovery of the execution of the task process.
[0069] Exemplarily, after the processing devices 1021 to 1024 fail and go down, in the device mapping table, the processing devices 1021 to 1024 are deleted, and new processing devices are configured to build the mapping relationship with the virtual devices. For details, please refer to Table II.
[0070] Table II Updated device mapping table Device ID of the virtual device Device ID of the processing device dp0-tp0-pp0 node1-gpu0 dp0-tp1-pp0 node1-gpu1 dp1-tp0-pp0 node1-gpu2 dp1-tp1-pp0 node1-gpu3 dp0-tp0-pp1 node2-gpu0 dp0-tp1-pp1 node2-gpu1 dp1-tp0-pp1 node2-gpu2 dp1-tp1-pp1 node2-gpu3 dp0-tp0-pp2 node3-gpu0 dp0-tp1-pp2 node3-gpu1 dp1-tp0-pp2 node3-gpu2 dp1-tp1-pp2 node3-gpu3 dp0-tp0-pp3 node5-gpu0 dp0-tp1-pp3 node5-gpu1 dp1-tp0-pp3 node5-gpu2 dp1-tp1-pp3 node5-gpu3 The second aspect of the present disclosure provides a method for resuming large model training, which is applied to a processing device.
[0071] Figure 6 Schematically shows a second flowchart of a method for resuming large model training according to an embodiment of the present disclosure.
[0072] Specifically, as Figure 6 shown, the method includes operations S601 to S603.
[0073] Operation 601, obtain a second communication topology relationship; the second communication topology relationship is obtained by updating a first communication topology relationship based on the task processes of multiple processing devices executing a large model training task and the device information of the multiple processing devices; the communication topology relationship represents the communication connection method between virtual devices corresponding to each processing device; Operation 602, update the communication connection with other processing devices based on the second communication topology relationship, and obtain task data for executing the task process from other processing devices; Operation 603, execute the task process based on the task data.
[0074] Specifically, after obtaining the updated second communication topology relationship, a processing device that has not failed (hereinafter referred to as the second processing device) can establish a communication connection with other processing devices. After the communication connection is established, obtain the task data of the task process that stopped due to a failure to execute the task process corresponding to the task data.
[0075] Exemplarily, as Figure 5 shown, after receiving the second communication topology relationship in the slave node, the virtual devices that need to be established can be determined according to the second communication topology relationship. After the communication connection of the virtual devices is established. Processing devices 1025 to 1028 query the corresponding virtual devices based on the updated device mapping table and establish a mapping relationship.
[0076] According to an embodiment of the present disclosure, updating the communication connection with other processing devices based on the second communication topology relationship includes: According to the second communication topology relationship, confirm a first processing device with a closed device attribute and a second processing device with an open device attribute; establish a communication connection with the second processing device and delete the communication connection with the first processing device.
[0077] Exemplarily, as Figure 5As shown, the processing device 1025 queries the second processing devices with the device attribute of the on attribute according to the second communication topology relationship, including the processing device 1026, the processing device 1027, and the processing device 1028, and establishes a communication connection. At the same time, in the communication connection of the processing devices, after receiving the second communication topology connection in the slave node, the communication connections between the processing devices 1021 to 1024 and other processing devices are deleted.
[0078] Figure 7 Schematically shows a schematic diagram of the principle of constructing communication connections between multiple devices according to an embodiment of the present disclosure.
[0079] As Figure 7 shown, where the green GPUs represent virtual devices and the white GPUs represent processing devices. In one slave node, four virtual devices form a communication group. When a virtual device in the communication group fails, after receiving the second communication topology connection in the slave node, delete Figure 7 the virtual device marked in red. At this time, the communication connection in the communication group is interrupted. From the second communication topology connection, select the virtual device corresponding to the processing device with the device attribute of the on attribute, and construct a communication connection with the other three virtual devices in the communication group. After the communication connection is established. Through the device mapping table, the four virtual devices in the communication group respectively establish a mapping relationship with the corresponding processing devices.
[0080] In the embodiments of the present disclosure, the communication connections between the processing devices are the same as the connection relationships of the respective virtual devices in the communication topology diagram.
[0081] According to an embodiment of the present disclosure, based on the task data, execute the task process, including: Obtain the data volume of the task data; compare the data volume of the task data with the storage volume of the local memory; based on the comparison result, write the task data from the local memory.
[0082] In operation S603, for the processing devices 1025 to 1028 that execute the task process stopped due to a failure, when the data volume of the task data is large, it may cause insufficient storage space in the local memory of the slave node, and then cause the task data to not be fully written into the slave node. By comparing the data volume of the task data with the storage volume of the local memory; based on the comparison result, write all or part of the task data from the local memory into the processing devices 1025 to 1028.
[0083] According to an embodiment of the present disclosure, based on the comparison result, write the task data from the local memory, including: when the data volume is less than the storage volume, save the task data to the local memory; write the task data from the local memory.
[0084] Exemplarily, if the local memory capacity is sufficient, the data is directly read into the memory file system path based on tmpfs. After the training process completes the necessary initialization operations, it reads the data from this path into the GPU video memory.
[0085] Among them, tmpfs is a memory-based temporary file system.
[0086] According to an embodiment of the present disclosure, writing task data from local memory based on the comparison result further includes: In the case where the data volume is greater than or equal to the storage volume, save the first part of the task data to local memory; save the second part of the task data to a pre-configured database; write the first part of the task data from local memory; at the same time, write the second part of the task data to local memory.
[0087] Exemplarily, in distributed large model training, a slave node pulls the task data (such as Checkpoint files or gradient data) left by a faulty device from other slave nodes, but the local memory capacity is not sufficient to store the complete data at one time and needs to be written in stages. The slave node obtains the total size of the task data (such as 12GB) through a data interface. Calls the API of the slave node to obtain the remaining space in local memory (such as 8GB).
[0088] Since the data volume is greater than or equal to the storage volume, the first 8GB of the task data (matching the remaining memory capacity) is written into the buffer of local memory through TCP streaming. After the writing is completed, this part of the data is immediately loaded into the GPU video memory of the GPU device to perform partial gradient aggregation or parameter update. Write the remaining 4GB of data to the temporary directory of the local file system (i.e., the pre-configured database). When this part of the data is needed, the data in the local file system is loaded into local memory through memory mapping.
[0089] Exemplarily, when the local memory capacity is insufficient, when the processing device executes the data pulling process, it will apply for a memory area of a fixed size in local memory, create a thread P1 to save the Checkpoint data therein first. When the memory area is full, write the remaining data to the local file system (non-memory). At the same time, start a new thread P2 to continuously and asynchronously try to write the data in the file system to the memory area. If the training process has not started reading the Checkpoint data in the memory, the memory area is still full at this time, and P2 will continue to try. If the training process has started reading the data in the memory, it will continuously release the storage space occupied by the read data. At this time, thread P2 will sequentially write the data in the local file system to the free memory area in local memory. In the case where the data volume is greater than the storage volume, by writing task data in stages, it is possible to avoid data overflow in the Bendu memory and ensure data integrity. While writing the data in the Bendu memory to execute the task process, the local memory is released to reduce the waiting time when the task data is written and improve the model training efficiency.
[0090] A third aspect of the present disclosure provides an electronic device, including a transceiver and a processor; The processor is coupled to the transceiver; the processor is configured to: in response to a first processing device among a plurality of processing devices having a failure, obtain the task processes of the plurality of processing devices for executing the large model training task, and the device information of the plurality of processing devices; Based on the task processes and the device information, update the first communication topology relationship of the plurality of processing devices to obtain a second communication topology relationship; the communication topology relationship represents the communication connection manner between the virtual devices corresponding to each processing device; Send the second communication topology relationship to the processing devices that have not failed.
[0091] In an embodiment of the present disclosure, the processor is further configured to: before a first processing device among a plurality of processing devices has a failure, based on the heartbeat messages sent by the plurality of processing devices, determine whether a first processing device among the plurality of processing devices has a failure.
[0092] In an embodiment of the present disclosure, the processor is further configured to: receive the heartbeat messages sent by the plurality of processing devices according to a preset time interval; in the case where the heartbeat message sent by the first processing device meets the preset message type, confirm that the first processing device has a failure; or, in the case where the first processing device does not send a heartbeat message according to the preset time interval, confirm that the first processing device has a failure.
[0093] In an embodiment of the present disclosure, the processor is further configured to: according to the task process and the device information of the first processing device, adjust the device attribute of the first processing device to a closed attribute; according to the task processes and the device information of the plurality of processing devices, adjust the device attribute of a second processing device to an open attribute; the second processing device is a processing device that has not failed; based on the device attributes of the plurality of processing devices, update the first communication topology relationship of the plurality of processing devices to obtain a second communication topology relationship.
[0094] It should be noted that the electronic device part in the embodiment of the present disclosure corresponds to the model training recovery method part in the embodiment of the present disclosure, and their specific implementation details are also the same, which will not be elaborated here.
[0095] Figure 8 A block diagram of another electronic device suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. Figure 8The illustrated electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0096] As Figure 8 shown, the electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general microprocessor (e.g., CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application specific integrated circuit (ASIC)), and so on. The processor 801 may also include on-board memory configured for caching purposes. The processor 801 may include a single processing unit or multiple processing units configured to perform different actions of the method flow according to an embodiment of the present disclosure.
[0097] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to an embodiment of the present disclosure by executing the program in the ROM 802 and / or the RAM 803. It should be noted that the program may also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 may also perform various operations of the method flow according to an embodiment of the present disclosure by executing the program stored in the one or more memories.
[0098] According to an embodiment of the present disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read therefrom can be installed into the storage portion 808 as needed.
[0099] According to an embodiment of the present disclosure, the method flow according to the embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium. The computer program includes program code configured to execute the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system according to the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, device, apparatus, module, unit, etc. can be implemented by computer program modules.
[0100] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.
[0101] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0102] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803.
[0103] An embodiment of the present disclosure also includes a computer program product, which includes a computer program. The computer program includes program code configured to execute the method provided by the embodiment of the present disclosure. When the computer program product runs on an electronic device, the program code is configured to enable the electronic device to implement the remote sensing image detection method based on a deep neural network provided by the embodiment of the present disclosure.
[0104] When the computer program is executed by the processor 801, the above functions defined in the system / apparatus of the embodiments of the present disclosure are executed. According to the embodiments of the present disclosure, the systems, apparatuses, modules, units, etc. described above can be implemented by computer program modules.
[0105] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program can also be transmitted and distributed in the form of signals on a network medium, and be downloaded and installed through the communication part 809, and / or be installed from the removable medium 811. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0106] According to the embodiments of the present disclosure, the program code configured to execute the computer program provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include but are not limited to, such as Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions configured to implement the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art will appreciate that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0108] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A method for restoring model training, comprising: In response to a failure of a first processing device among multiple processing devices, obtaining the task progress of the multiple processing devices performing the large model training task, and the device information of the multiple processing devices; Based on the task progress and the device information, updating the first communication topology relationship of the multiple processing devices to obtain a second communication topology relationship; the communication topology relationship characterizes the communication connection mode between the virtual devices corresponding to each processing device; Sending the second communication topology relationship to the processing devices that have not failed.
2. The method according to claim 1, before responding to a failure of a first processing device among multiple processing devices, the method further comprises: Based on the heartbeat messages sent by the multiple processing devices, determining whether a first processing device among the multiple processing devices has failed.
3. The method according to claim 2, wherein the determining whether a first processing device among the multiple processing devices has failed based on the heartbeat messages sent by the multiple processing devices comprises: Receiving the heartbeat messages sent by the multiple processing devices according to a preset time interval; When the heartbeat message sent by the first processing device meets the preset message type, confirming that the first processing device has failed; Or, when the first processing device does not send a heartbeat message according to the preset time interval, confirming that the first processing device has failed.
4. The method according to claim 1, updating the first communication topology relationship of the multiple processing devices based on the task progress and the device information to obtain a second communication topology relationship, comprising: According to the task progress and device information of the first processing device, adjusting the device attribute of the first processing device to a closed attribute; According to the task progress and device information of the multiple processing devices, adjusting the device attribute of a second processing device to an open attribute; the second processing device is a processing device that has not failed; Based on the device attributes of the multiple processing devices, updating the first communication topology relationship of the multiple processing devices to obtain a second communication topology relationship.
5. The method according to claim 1, wherein the device information includes the device addresses, network names, and communication connection information of the multiple processing devices.
6. A method for restoring model training, comprising: Obtaining a second communication topology relationship; The second communication topology relationship is obtained by updating the first communication topology relationship based on the task progress of the multiple processing devices performing the large model training task and the device information of the multiple processing devices; the communication topology relationship characterizes the communication connection mode between the virtual devices corresponding to each processing device; Based on the second communication topology relationship, updating the communication connection with other processing devices, and obtaining task data for executing the task progress from the other processing devices; Based on the task data, executing the task progress.
7. The method according to claim 6, updating the communication connection with other processing devices based on the second communication topology relationship, comprising: According to the second communication topology relationship, confirming a first processing device with a closed device attribute and a second processing device with an open device attribute; Establish a communication connection with the second processing device and delete the communication connection with the first processing device.
8. The method according to claim 6, performing the task process based on the task data, including: Obtain the data volume of the task data; Compare the data volume of the task data with the storage capacity of the local memory; Based on the comparison result, write the task data from the local memory.
9. The method according to claim 8, the writing the task data from the local memory based on the comparison result includes: In the case where the data volume is less than the storage capacity, save the task data to the local memory; Write the task data from the local memory; Alternatively, in the case where the data volume is greater than or equal to the storage capacity, save the first part of the task data to the local memory; Save the second part of the task data to a pre-configured database; Write the first part of the task data from the local memory; at the same time, write the second part of the task data to the local memory.
10. An electronic device, including a transceiver and a processor; and The processor is coupled to the transceiver; the processor is configured to: In response to a failure of the first processing device among multiple processing devices, obtain the task process of the multiple processing devices performing the large model training task, and the device information of the multiple processing devices; Based on the task process and the device information, update the first communication topology relationship of the multiple processing devices to obtain a second communication topology relationship; the communication topology relationship represents the communication connection method between the virtual devices corresponding to each processing device; Send the second communication topology relationship to the processing devices that have not failed.
Citation Information
Cited By
Communication method and device
CN120582963A
Communication method and apparatus
CN120582963B