NCCL task scheduling method and device, related equipment, storage medium and computer program product

By monitoring the fault conditions and performance of the channel in the NCCL system, dynamically adjusting the scheduling priority and rebuilding the communication topology, the problem of insufficient fault detection and diagnosis capabilities of the NCCL system is solved, and the communication efficiency and reliability of the system are improved.

CN120066772APending Publication Date: 2025-05-30CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510125473.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Various failures may occur when NCCL systems perform parallel computing tasks, affecting communication efficiency and reliability. The existing technology lacks internal fault detection and diagnosis capabilities, resulting in low fault tolerance.

Method used

By monitoring the failure conditions, performance, and load of each channel in the main process of the NCCL process group, dynamically adjust the scheduling priorities of the channel, and rebuild the topology and channels of the communication network for fault detection, diagnosis, and fault tolerance.

Benefits of technology

It improves the communication efficiency and reliability of the NCCL system, realizes real-time monitoring of faults and dynamic load adjustment, and reduces task recovery time and operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066772A_ABST
    Figure CN120066772A_ABST
Patent Text Reader

Abstract

The invention discloses an NVIDIA collective communication library (NCCL) task scheduling method and device, first equipment, second equipment, a storage medium and a computer program product. The method comprises the following steps: a first process in an NCCL process group monitors one or more of second information, third information and fourth information of each channel in a first NCCL topology based on first information periodically reported by each second process in the NCCL process group, the first information comprises related information of the operation condition and / or performance of first equipment, the second information represents whether a corresponding channel breaks down or not, the third information represents the performance of the corresponding channel, and the fourth information represents the load condition of the corresponding channel; determining or adjusting the scheduling priority of the corresponding channel based on one or more of the second information, the third information and the fourth information of each channel; and scheduling the NCCL tasks in the task queue based on the scheduling priority of each channel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of infrastructure and information technology (IT) support, and particularly to a method, apparatus, related device, storage medium, and computer program product for NVIDIA Collective Communications Library (NCCL) task scheduling. Background Art

[0002] NCCL is an open-source parallel computing library, specifically a communication library for accelerating collective communication computing of a distributed graphics processing unit (GPU) set; NCCL is designed specifically for NVIDIA GPUs and is used to accelerate applications such as deep learning training and high-performance computing (HPC). NCCL provides a set of efficient and scalable primitives and provides developers with a simple interface to perform collective communication operations between multiple GPUs on a GPU cluster, realizing high-performance communication between multiple devices (i.e., multiple GPUs). Since applications such as deep learning training and HPC may need to synchronously transmit data between multiple GPUs when performing parallel computing tasks, the above-mentioned collective communication operations between multiple GPUs are crucial for applications such as deep learning training and HPC. Generally speaking, NCCL has characteristics such as high performance, easy integration, strong scalability, and high flexibility, and is a powerful parallel computing library suitable for performing tasks such as deep learning training and HPC on NVIDIA GPUs; moreover, by providing efficient collective communication primitives, NCCL can significantly improve the data transmission speed and synchronization performance between multiple GPUs, thereby accelerating the calculation process of corresponding tasks.

[0003] However, various failures may occur in the NCCL system, thus affecting the communication efficiency and reliability of the system. Summary of the Invention

[0004] To solve the related technical problems, embodiments of this application provide a method, apparatus, related device, storage medium, and computer program product for NCCL task scheduling.

[0005] The technical solution of the embodiments of this application is implemented as follows:

[0006] Embodiments of this application provide a method for NCCL task scheduling, which is applied to a first process in an NCCL process group and includes:

[0007] Monitor one or more of the second information, third information, and fourth information of each channel in the first NCCL topology based on the first information periodically reported by each second process in the NCCL process group. The first information includes information related to the operating status and / or performance of the first device, and at least the corresponding second process is running on the first device. The second information characterizes whether the corresponding channel has a fault, the third information characterizes the performance of the corresponding channel, and the fourth information characterizes the load condition of the corresponding channel. The first NCCL topology is the topology of the communication network corresponding to the NCCL process group, the first process is the main process of the NCCL process group, and the second processes include other processes in the NCCL process group except the first process.

[0008] Determine or adjust the scheduling priority of the corresponding channel based on one or more of the second information, third information, and fourth information of each channel.

[0009] Schedule the NCCL tasks in the task queue based on the scheduling priority of each channel.

[0010] In the above solution, the determining or adjusting the scheduling priority of the corresponding channel includes:

[0011] Determine or adjust the scheduling priority of the corresponding channel by using one or more of the following rules:

[0012] The scheduling priority of the first channel is higher than that of the second channel. The first channel has no fault, and the second channel has a fault.

[0013] The scheduling priority of the third channel is higher than that of the fourth channel. The performance of the third channel is better than that of the fourth channel.

[0014] The scheduling priority of the fifth channel is higher than that of the sixth channel. The load of the fifth channel is less than that of the sixth channel.

[0015] In the above solution, the monitoring one or more of the second information, third information, and fourth information of each channel in the first NCCL topology based on the first information periodically reported by each second process in the NCCL process group includes:

[0016] Based on the first information reported by each second process, perform a fault diagnosis on each first device in the first NCCL topology to obtain a fault diagnosis result.

[0017] For each first device, based on the fault diagnosis result, determine second information of the channel associated with the corresponding first device; wherein, when the fault diagnosis result indicates that the corresponding first device has a fault, the second information of the channel associated with the corresponding first device indicates that the corresponding channel has a fault.

[0018] In the above solution, the method further includes:

[0019] When the fault diagnosis result indicates that the corresponding first device has a fault, obtain preset fifth information, where the fifth information includes alarm-related information;

[0020] Based on the fifth information, send sixth information, where the sixth information is used to perform fault alarm for the corresponding first device and / or the corresponding channel.

[0021] In the above solution, the method further includes:

[0022] For each first device, when the fault diagnosis result indicates that the corresponding first device has a fault, determine the state of the corresponding first device as an unavailable state, and determine the state of the channel associated with the corresponding first device as an unavailable state. The first device and the channel in the unavailable state cannot be scheduled;

[0023] When the first parameter is greater than or equal to the first threshold, reconstruct the topology and channels of the communication network to obtain a second NCCL topology; the first parameter represents the ratio of the number of channels in the unavailable state in the first NCCL topology to the total number of channels, and the second NCCL topology does not include the first device and channels in the unavailable state in the first NCCL topology;

[0024] Based on the second NCCL topology, schedule the NCCL tasks in the task queue;

[0025] Send seventh information to each second process, where the seventh information is used to indicate the second NCCL topology.

[0026] In the above solution, when the fault diagnosis result indicates that the corresponding first device has a fault, determining the state of the corresponding first device as an unavailable state and determining the state of the channel associated with the corresponding first device as an unavailable state includes:

[0027] When the fault diagnosis result indicates that the corresponding first device has a fault, determine the state of the corresponding first device as a temporarily unavailable state, determine the state of the channel associated with the corresponding first device as a temporarily unavailable state, and perform a retry of fault diagnosis for the corresponding first device;

[0028] In the case where it is still determined that the corresponding first device has a fault after M retries of fault diagnosis for the corresponding first device, the state of the corresponding first device is determined to be a permanently unavailable state, and the state of the channel associated with the corresponding first device is determined to be a permanently unavailable state, where M is an integer greater than 0.

[0029] In the above solution, scheduling the NCCL tasks in the task queue based on the scheduling priority of each channel includes:

[0030] For each NCCL task in the task queue, based on the scheduling priority of each channel, select one or more channels for the corresponding NCCL task, and schedule one or more first devices associated with the one or more channels to execute the corresponding NCCL task;

[0031] In the case where the target NCCL task fails to execute, schedule the one or more first devices to re-execute the target NCCL task;

[0032] In the case where it still fails after re-executing the target NCCL task N times, re-add the target NCCL task to the task queue to reschedule the target NCCL task, and determine the state of each channel in the one or more channels to be an unavailable state, and the channels in the unavailable state cannot be scheduled, where N is an integer greater than 0.

[0033] An embodiment of the present application further provides an NCCL task scheduling method, which is applied to a second process in an NCCL process group, and includes:

[0034] Monitor the operating condition and / or performance of the first device to obtain first information, where the first information includes information related to the operating condition and / or performance of the first device, and at least the corresponding second process is running on the first device;

[0035] Periodically report the first information to the first process in the NCCL process group for the first process to schedule the NCCL tasks in the task queue; the first process is the main process of the NCCL process group, and the second process includes other processes in the NCCL process group except the first process.

[0036] In the above solution, the method further includes:

[0037] Receiving the seventh information sent by the first process, where the seventh information is used to indicate a second NCCL topology, and the second NCCL topology is the topology of the communication network corresponding to the NCCL process group. The second NCCL topology is the topology obtained by the first process reconstructing the topology and channels of the communication network when a first parameter is greater than or equal to a first threshold. The first parameter represents the ratio of the number of channels in an unavailable state in the first NCCL topology to the total number of channels. The second NCCL topology does not include the first device and channels in the unavailable state in the first NCCL topology.

[0038] An embodiment of the present application further provides an NCCL task scheduling device, including:

[0039] A first processing unit, after being utilized by the first process in the NCCL process group, monitors one or more of the second information, third information, and fourth information of each channel in the first NCCL topology based on the first information periodically reported by each second process in the NCCL process group. The first information includes information related to the operating condition and / or performance of a first device, and at least one corresponding second process runs on the first device. The second information represents whether the corresponding channel has a fault. The third information represents the performance of the corresponding channel. The fourth information represents the load condition of the corresponding channel. The first NCCL topology is the topology of the communication network corresponding to the NCCL process group. The first process is the main process of the NCCL process group. The second processes include other processes in the NCCL process group except the first process.

[0040] A second processing unit, after being utilized by the first process, determines or adjusts the scheduling priority of the corresponding channel based on one or more of the second information, third information, and fourth information of each channel.

[0041] A scheduling unit, after being utilized by the first process, schedules the NCCL tasks in the task queue based on the scheduling priority of each channel.

[0042] An embodiment of the present application further provides an NCCL task scheduling device, including:

[0043] A third processing unit, after being utilized by the second process in the NCCL process group, monitors the operating condition and / or performance of the first device to obtain the first information, where the first information includes information related to the operating condition and / or performance of the first device, and at least one corresponding second process runs on the first device.

[0044] A second communication unit, which is used to periodically report the first information to the first process in the NCCL process group after being utilized by the second process, so as to enable the first process to schedule the NCCL tasks in the task queue; the first process is the master process of the NCCL process group, and the second process includes other processes in the NCCL process group except the first process.

[0045] An embodiment of the present application further provides a second device, which runs at least the first process in the NCCL process group, and the first process is the master process of the NCCL process group; the second device includes: a first processor and a first memory for storing a computer program that can run on the processor.

[0046] Wherein, when the first processor is used to run the computer program, it executes the steps of any of the methods on the first process side.

[0047] An embodiment of the present application further provides a first device, which runs at least the corresponding second process in the NCCL process group, and the second process includes other processes in the NCCL process group except the first process, and the first process is the master process of the NCCL process group; the first device includes: a second processor and a second memory for storing a computer program that can run on the processor.

[0048] Wherein, when the second processor is used to run the computer program, it executes the steps of any of the methods on the second process side.

[0049] An embodiment of the present application further provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of any of the methods on the first process side, or implements the steps of any of the methods on the second process side.

[0050] An embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of any of the methods on the first process side, or implements the steps of any of the methods on the second process side.

[0051] The NCCL task scheduling method, device, related equipment, storage medium and computer program product provided by the embodiments of the present application. The first process in the NCCL process group monitors one or more of the second information, third information, and fourth information of each channel in the first NCCL topology based on the first information periodically reported by each second process in the NCCL process group. The first information includes information related to the operating status and / or performance of the first device, and at least the corresponding second process runs on the first device. The second information represents whether the corresponding channel fails, the third information represents the performance of the corresponding channel, and the fourth information represents the load condition of the corresponding channel; the first NCCL topology is the topology of the communication network corresponding to the NCCL process group, the first process is the main process of the NCCL process group, and the second processes include other processes in the NCCL process group except the first process; based on one or more of the second information, third information, and fourth information of each channel, determine or adjust the scheduling priority of the corresponding channel; based on the scheduling priority of each channel, schedule the NCCL tasks in the task queue. In the solution provided by the embodiments of the present application, each other process (i.e., the second process) in the NCCL process group except the main process (i.e., the first process) periodically reports information related to the operating status and / or performance of its own device (i.e., the first device) (i.e., the first information) to the main process. The main process monitors one or more of the failure conditions (i.e., whether it fails, that is, the second information), performance (i.e., the third information), and load conditions (i.e., the fourth information) of each channel in the corresponding NCCL topology (i.e., the first NCCL topology) based on the reported information, determines or adjusts the scheduling priority of the corresponding channel based on the monitored content, and schedules the NCCL tasks in the task queue based on the scheduling priority of each channel; thus, the main process can monitor one or more of the failure conditions, performance, and load conditions of each channel in real time according to the real-time operating status and / or performance of each device in the NCCL system, and can dynamically adjust the scheduling priority of the corresponding channel according to the monitored content, so as to be able to implement scheduling strategies (which can also be understood as scheduling rules) such as preferentially allocating NCCL tasks to channels with low failure rate, stable performance, and small load, thereby being able to dynamically adjust the load distribution of NCCL tasks among channels, that is, being able to achieve dynamic load adjustment of each channel in the NCCL system, and further being able to improve the overall communication efficiency and reliability of the NCCL system. Description of the Drawings

[0052] Figure 1 It is a schematic flowchart of a method for scheduling NCCL tasks according to an embodiment of the present application;

[0053] Figure 2Schematic flowchart of another NCCL task scheduling method according to an embodiment of the present application;

[0054] Figure 3 Schematic diagram of the NCCL fault detection and diagnosis architecture for an application example of the present application;

[0055] Figure 4 Schematic diagram of the NCCL task scheduling channel architecture for an application example of the present application;

[0056] Figure 5 Schematic flowchart of the NCCL fault tolerance and recovery mechanism for an application example of the present application;

[0057] Figure 6 Schematic diagram of the structure of an NCCL task scheduling device according to an embodiment of the present application;

[0058] Figure 7 Schematic diagram of the structure of another NCCL task scheduling device according to an embodiment of the present application;

[0059] Figure 8 Schematic diagram of the structure of the second device according to an embodiment of the present application;

[0060] Figure 9 Schematic diagram of the structure of the first device according to an embodiment of the present application;

[0061] Figure 10 Schematic diagram of the structure of the NCCL task scheduling system according to an embodiment of the present application. Detailed implementation manners

[0062] The present application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0063] In the related art, NCCL has the characteristics of high performance, easy integration, strong scalability, and high flexibility. Specifically, NCCL is optimized for NVIDIA GPUs and high-speed networks, and can maximize throughput and reduce latency, which makes it possible to achieve efficient data transmission and synchronization in multi-GPU and multi-node systems; at the same time, NCCL provides a simple and easy-to-use application programming interface (API), which can be easily integrated into various deep learning frameworks and HPC applications, such as TensorFlow, PyTorch, Horovod, etc.; moreover, NCCL has good scalability and can support multiple GPUs on a single node to large GPU clusters spanning thousands of nodes; in addition, NCCL supports various network topologies, such as InfiniBand, Ethernet, NVLink networks, etc. with different bandwidth and latency characteristics.

[0064] However, as the scale of distributed collective communication computing tasks (which can also be referred to as NCCL tasks) increases, various types of failures (which can also be understood as problems) may occur in NCCL communication operations in the NCCL system. Exemplarily, the possible failures and impacts in the NCCL system are as follows:

[0065] Failure 1, device failure; specifically, during the execution of NCCL tasks, key hardware such as GPUs and network cards may fail, resulting in the inability of NCCL tasks to proceed normally. These device failures may be caused by reasons such as hardware aging, overheating, and power problems;

[0066] Failure 2, network failure; specifically, during the execution of NCCL tasks, communication problems such as network latency, packet loss, and congestion may be encountered, which may lead to communication failures between nodes and thus affect the normal progress of NCCL tasks;

[0067] Failure 3, software failure; specifically, during the execution of NCCL tasks, software failures such as memory leaks and deadlocks may occur, which may cause NCCL tasks to freeze or crash and thus affect the normal progress of NCCL tasks;

[0068] Failure 4, software configuration problem; specifically, software configuration errors such as drivers and network cards may occur during the development and production process, which may cause corresponding hardware such as GPUs and network cards to be unable to be used normally and thus affect the normal operation of NCCL tasks.

[0069] As can be seen from the above description, in the actual operation of NCCL, the above failures may lead to problems such as NCCL task interruption and performance degradation, thus affecting the normal progress of NCCL tasks. Therefore, related technologies have proposed an NCCL fault tolerance and recovery mechanism, which usually includes the following four solutions:

[0070] Solution 1, manual troubleshooting and recovery solution; specifically, when a failure is encountered, the NCCL task will be stopped. At this time, relevant logs can be manually checked to troubleshoot problems such as devices, communication, and software configuration. After solving the relevant problems, the NCCL task can be restarted to resume operation;

[0071] Solution 2, checkpoint-restart solution; specifically, during the operation of NCCL, specific upper-layer applications (which can be simply referred to as upper-layer applications) can periodically save the status information of the current relevant processes so that fault recovery can be performed based on the saved status in case of a failure;

[0072] Solution 3, an error detection and handling solution; specifically, the upper-layer application can capture exceptions through NCCL and execute an error handling function to handle possible errors, thereby improving the fault tolerance of the program;

[0073] Solution 4, an external monitoring solution; specifically, real-time monitoring can be provided outside NCCL. Specifically, the server status and relevant process information can be monitored to troubleshoot server failures in real time.

[0074] However, the common solutions for the NCCL fault tolerance and recovery mechanism may have the following problems:

[0075] Problem 1, the efficiency of manual troubleshooting is relatively low; specifically, manually checking the logs and troubleshooting problems may take a lot of time and effort, which may lead to a slow recovery speed of NCCL tasks and affect the overall operation efficiency;

[0076] Problem 2, the fault tolerance and recovery logic are not unified; specifically, since different NCCL tasks may require writing different fault response logics, this results in a lack of a unified processing standard for NCCL faults between different NCCL tasks and frameworks, increasing the maintenance and learning costs;

[0077] Problem 3, the fault tolerance and recovery mechanism written by the upper layer has limitations; specifically, writing a custom fault response logic in the upper-layer logic may not be able to fully perform internal fine-grained processing of NCCL faults and can only simply handle NCCL faults at the outer layer, such as restarting, etc.;

[0078] Problem 4, the upper-layer development is difficult; specifically, upper-layer application development may have relatively high requirements for developers' skills and understanding of NCCL, increasing the implementation difficulty.

[0079] In summary, in the related technologies, NCCL does not support internal fault detection and diagnosis, that is, the NCCL system does not support early detection of faults internally and cannot avoid faults in advance. It can only report faults to the upper-layer application for processing when the faults are triggered, and all fault handling depends on the upper-layer application to solve. The upper-layer application is usually outside the NCCL perimeter and cannot finely control the internal logic of NCCL. Usually, it can only simply record the status and restart relevant processes, resulting in relatively low fault tolerance of NCCL itself and a high restart frequency, thus making the communication efficiency and reliability of the NCCL system poor.

[0080] Based on this, in various embodiments of the present application, each other process in the NCCL process group except the main process periodically reports information related to the operating status and / or performance of its own device to the main process. The main process monitors one or more of the fault conditions (i.e., whether a fault occurs), performance, and load conditions of each channel in the corresponding NCCL topology based on the reported information, determines or adjusts the scheduling priority of the corresponding channel based on the monitored content, and schedules the NCCL tasks in the task queue based on the scheduling priority of each channel. In this way, the main process can monitor one or more of the fault conditions, performance, and load conditions of each channel in real time according to the real-time operating status and / or performance of each device in the NCCL system, and can dynamically adjust the scheduling priority of the corresponding channel according to the monitored content, so as to be able to implement scheduling strategies (which can also be understood as scheduling rules) such as preferentially allocating NCCL tasks to channels with low failure rate, stable performance, and light load, thereby being able to dynamically adjust the load distribution of NCCL tasks among channels, that is, being able to achieve dynamic load adjustment of each channel in the NCCL system, and further being able to improve the overall communication efficiency and reliability of the NCCL system.

[0081] It should be noted that, in various embodiments of the present application, "one or more / a plurality of" means at least one / at least one item, and "a plurality of" means at least two / at least two items.

[0082] An embodiment of the present application provides an NCCL task scheduling method, which is applied to a first process in an NCCL process group, as Figure 1 shown, the method includes:

[0083] Step 101: Based on the first information periodically reported by each second process in the NCCL process group, monitor one or more of the second information, third information, and fourth information of each channel (which can be expressed in English as "channel") in the first NCCL topology. The first information includes information related to the operating status and / or performance of the first device, and at least the corresponding second process is running on the first device. The second information represents whether a fault occurs in the corresponding channel, the third information represents the performance of the corresponding channel, and the fourth information represents the load condition of the corresponding channel. The first NCCL topology is the topology of the communication network corresponding to the NCCL process group, the first process is the main process of the NCCL process group, and the second processes include other processes in the NCCL process group except the first process.

[0084] Step 102: Determine or adjust the scheduling priority of the corresponding channel based on one or more of the second information, third information, and fourth information of each channel.

[0085] Step 103: Schedule the NCCL tasks in the task queue based on the scheduling priority of each channel.

[0086] In practical applications, the first process may also be referred to as the main process, etc., and the second process may also be referred to as an ordinary process, etc. The embodiments of the present application do not limit the process names, as long as their functions are implemented. Moreover, the first process may run on the second device, that is, the above NCCL task scheduling method may also be applied to the second device.

[0087] Here, the second device may also be a first device, that is, the first process and one or more corresponding second processes may run on the second device at the same time; or, the second device may not be a first device, that is, only the first process may run on the second device without running the second process. From another perspective, the NCCL system may be deployed with one or more devices, that is, the device cluster corresponding to the NCCL process group may include one or more devices, and these devices may also be referred to as nodes (the second device may also be referred to as the master node, and the first device may also be referred to as an ordinary node). Specifically, it may include devices of types such as servers, network devices, link devices, and GPU devices; one or more corresponding NCCL processes may run on each device. When the first process runs on a device, the device is the second device; when the second process runs on a device, the device is the first device; in other words, the first device and the second device may be the same or different.

[0088] In practical applications, the first NCCL topology may also be referred to as a network topology or a communication topology, etc. The embodiments of the present application do not limit the name of the first NCCL topology, as long as it can represent the topological structure of the communication network corresponding to the NCCL process group. Additionally, the channel may also be referred to as a communication channel or an NCCL task scheduling channel, etc.; the NCCL task may also be referred to as an NCCL communication task, a distributed collective communication computing task, a distributed collective communication operation task, a distributed communication collective operation task, etc.; the embodiments of the present application do not limit the channel name and the task name, as long as their functions are implemented.

[0089] In actual application, each of all other processes (i.e., each of the second processes) in the NCCL process group except the first process can monitor the operating status and / or performance of its own device (i.e., the first device), obtain the first information, and periodically report the first information to the first process. Herein, the second process monitoring the operating status and / or performance of the first device can also be understood as the second process periodically collecting or detecting the first information. Exemplarily, each second process can collect the first information based on a first period and report the first information to the first process based on the first period; correspondingly, the first process can receive the first information reported by each second process based on the first period.

[0090] Herein, the specific value of the first period can be set as needed, such as X seconds or Y days, etc., where both X and Y can be integers greater than 0; and the first process can configure or adjust the value of the first period for each second process as needed (such as the device cluster deployment situation, etc.).

[0091] In addition, the specific manner for each second process to collect the first information can also be set as needed, and the embodiments of the present application do not limit this either. Exemplarily, each second process can use specific NCCL primitives and / or NCCL APIs to implement the collection of the first information. The first information can specifically include status information such as GPUs, network interface cards (NICs), and memory, and can also include performance metric information such as utilization rate and throughput rate, and can also include error log information such as memory exceptions and network packet losses.

[0092] In actual application, the monitoring of one or more of the second information, third information, and fourth information for each channel based on the first information periodically reported by each second process in the NCCL process group means that in each reporting period (such as the first period), the first process can determine one or more of the second information, third information, and fourth information for each channel based on the first information reported by each second process. Among them, when determining the second information, the first process can perform fault diagnosis (which can also be understood as fault detection) on each first device and each channel, so as to be able to realize the availability evaluation (which can also be understood as availability monitoring / detection) of each first device and each channel.

[0093] Based on this, in one embodiment, the monitoring of one or more of the second information, third information, and fourth information for each channel in the first NCCL topology based on the first information periodically reported by each second process in the NCCL process group may include:

[0094] Based on the first information reported by each second process, perform fault diagnosis on each first device in the first NCCL topology to obtain a fault diagnosis result;

[0095] For each first device, based on the fault diagnosis result, determine second information of the channel associated with the corresponding first device (i.e., the channel where the corresponding first device is located); wherein, when the fault diagnosis result indicates that the corresponding first device has a fault, the second information of the channel associated with the corresponding first device indicates that the corresponding channel has a fault.

[0096] Wherein, in practical applications, the second information can be understood as a specific identifier used to identify whether the corresponding channel has a fault, and the specific form of this identifier (which can also be understood as the specific content included in the second information) can be set as needed, and this application embodiment does not limit this. Exemplarily, when the identifier is 0 (i.e., the second information includes 0), it can indicate that the corresponding channel has a fault, and when the identifier is 1 (i.e., the second information includes 1), it can indicate that the corresponding channel has no fault. Correspondingly, it can be understood that the specific form of the fault diagnosis result (which can also be understood as the specific content included in the fault diagnosis result), and the specific manner in which the first process performs fault diagnosis on each first device (which can also be understood as the specific manner of determining the fault diagnosis result) can also be set as needed, and this application embodiment does not limit this either.

[0097] In practical applications, that a first device is associated with a channel can mean that the first device is on the channel, or it can be understood that the first device is a part of the channel. Additionally, it can be understood that when the fault diagnosis result indicates that the corresponding first device has no fault, it is temporarily impossible to determine whether the channel where the first device is located has a fault, that is, whether the channel where the first device is located has a fault may depend on whether other first devices on the channel except this first device have a fault; when all first devices on a channel have no fault, it can be determined that the channel has no fault.

[0098] In practical applications, when it is determined that a channel has a fault, the first process can notify relevant second processes to stop using the channel. Exemplarily, the first process can send eighth information to each second process among all second processes corresponding to all first devices on the channel, and the eighth information is used to indicate that the channel has a fault and / or is used to indicate to stop using the channel. Here, the specific content included in the eighth information can also be set as needed, such as including the identifier of the faulty channel, etc., and this application embodiment does not limit this either.

[0099] In actual application, in the case of a failure of the first device and the associated channels, the first process can record logs. Additionally, alarm-related information (which can be denoted as the fifth information in the following description) can be pre-configured, such as phone numbers for alarm, email addresses, identifiers of specific devices, Internet Protocol (IP) addresses, etc.; in the case where the fault diagnosis result indicates a fault in the corresponding first device, the first process can perform a fault alarm for the corresponding first device and / or the corresponding channel based on the fifth information, so that relevant personnel (such as operation and maintenance personnel, etc.) can perform fault handling (i.e., fault repair) in a timely manner based on the fifth information, thereby further improving the reliability of the NCCL system.

[0100] Based on this, in one embodiment, the method may further include:

[0101] In the case where the fault diagnosis result indicates a fault in the corresponding first device, obtain the preset fifth information, where the fifth information includes alarm-related information;

[0102] Based on the fifth information, send a sixth information, where the sixth information is used to perform a fault alarm for the corresponding first device and / or the corresponding channel.

[0103] Among them, in actual application, the alarm can also be referred to as an alert; the alarm-related information may specifically include one or more of the following contents: phone numbers for alarm, email addresses, identifiers of specific devices, IP addresses, etc. Sending the sixth information based on the fifth information may refer to making a specific call, sending a specific email to a specific email address, sending a specific text message to a specific device, etc.; in other words, the specific content and specific manifestation form of the sixth information can be set as needed, such as phone calls, text messages, emails, pop-up reminders, etc. The embodiments of the present application do not limit this, as long as its function is achieved.

[0104] In one embodiment, the method may further include:

[0105] For each first device, in the case where the fault diagnosis result indicates a fault in the corresponding first device, determine the state of the corresponding first device as an unavailable state, and determine the state of the channel associated with the corresponding first device as an unavailable state. The first device and the channel in the unavailable state cannot be scheduled;

[0106] In the case where a first parameter is greater than or equal to a first threshold, reconstruct the topology and channels of the communication network to obtain a second NCCL topology; the first parameter represents the ratio of the number of channels in the unavailable state in the first NCCL topology to the total number of channels, and the second NCCL topology does not include the first device and channels in the unavailable state in the first NCCL topology;

[0107] Schedule the NCCL tasks in the task queue based on the second NCCL topology;

[0108] Send seventh information to each second process, where the seventh information is used to indicate the second NCCL topology.

[0109] In practical applications, it can be understood that since the first process can implement the status maintenance of the corresponding first device and the status maintenance of the channels associated with the corresponding first device based on the fault diagnosis result, the first process can monitor the status of each first device based on the first information periodically reported by each second process, and at the same time can monitor the status of each channel based on the first information periodically reported by each second process. In addition, it can be understood that the first devices and channels in the unavailable state cannot be scheduled, which means that the first process cannot call the first devices and / or channels in the unavailable state to execute NCCL tasks. In this way, it is possible to prevent the spread of faults and affect other communication tasks, achieve fault prevention and isolation for NCCL, and thus further improve the reliability of the NCCL system.

[0110] In practical applications, the specific value of the first threshold can be set as needed, such as 20% etc., and the embodiments of the present application do not limit this. And the specific content included in the seventh information can also be set as needed as long as its function is achieved. In addition, it can be understood that the second NCCL topology is also the topology of the communication network corresponding to the NCCL process group.

[0111] In one embodiment, in the case where the fault diagnosis result indicates that the corresponding first device has a fault, determining the status of the corresponding first device as the unavailable state and determining the status of the channel associated with the corresponding first device as the unavailable state may include:

[0112] In the case where the fault diagnosis result indicates that the corresponding first device has a fault, determining the status of the corresponding first device as the temporarily unavailable state, determining the status of the channel associated with the corresponding first device as the temporarily unavailable state, and retrying the fault diagnosis for the corresponding first device;

[0113] In the case where it is still determined that the corresponding first device has a fault after retrying the fault diagnosis for the corresponding first device M times, determining the status of the corresponding first device as the permanently unavailable state and determining the status of the channel associated with the corresponding first device as the permanently unavailable state, where M is an integer greater than 0.

[0114] In actual application, the specific value of M can be set as needed, such as 3, and the embodiments of the present application do not limit this. Additionally, it can be understood that the specific manner of the first process for retrying the fault diagnosis of the corresponding first device, and whether the fault diagnosis methods are the same each time can also be set as needed, and the embodiments of the present application do not limit this either. Exemplarily, the first process can reuse relevant first information to perform fault diagnosis on the corresponding first device; and / or, the first process can schedule the corresponding first device to execute a specific task to achieve fault diagnosis of the corresponding first device.

[0115] In actual application, during the process of retrying the M'-th (M' is less than or equal to M) fault diagnosis of the corresponding first device, if it is determined that the first device has not failed, the first process can re-determine the state of the corresponding first device as the available state; if it can be determined that all first devices on the channel associated with the corresponding first device have not failed, the first process can re-determine the state of this channel as the available state. In this way, through retrying the fault diagnosis of the corresponding first device, the first process can achieve fault classification of the corresponding first device, that is, distinguish between temporary faults and permanent faults; and, the first process can achieve state grading of the corresponding first device, that is, set different levels of unavailable states, that is, set temporary unavailable states and permanent unavailable states, so as to avoid the situation of misdiagnosis of faults and be able to timely sense the recovery of faults, thereby further improving the reliability of the NCCL system.

[0116] In actual application, the specific manner and process for the first process to determine the third information and / or the fourth information of each channel based on the first information reported by each second process can be set as needed, and the embodiments of the present application do not limit this. Additionally, when the first process determines or adjusts the scheduling priority of the corresponding channel based on one or more of the second information, third information, and fourth information of each channel, scheduling strategies (which can also be understood as scheduling rules) such as preferentially allocating NCCL tasks to channels with low failure rates, stable performance, and small loads can be implemented, so as to be able to dynamically adjust the load distribution of NCCL tasks among channels, that is, be able to achieve dynamic load adjustment of each channel in the NCCL system, and further improve the overall communication efficiency and reliability of the NCCL system.

[0117] Based on this, in one embodiment, the determining or adjusting the scheduling priority of the corresponding channel may include:

[0118] Using one or more of the following rules (which can be understood as scheduling rules or scheduling strategies) to determine or adjust the scheduling priority of the corresponding channel:

[0119] The scheduling priority of the first channel is higher than that of the second channel. The first channel has not failed, and the second channel has failed.

[0120] The scheduling priority of the third channel is higher than that of the fourth channel. The performance of the third channel is better than that of the fourth channel, that is, the stability of the performance of the third channel is higher than that of the fourth channel.

[0121] The scheduling priority of the fifth channel is higher than that of the sixth channel. The load of the fifth channel is less than that of the sixth channel.

[0122] Specifically, in practical applications, the scheduling priority of the first channel being higher than that of the second channel can be understood as preferentially scheduling the channel that has not failed or has a lower failure rate; the scheduling priority of the third channel being higher than that of the fourth channel can be understood as preferentially scheduling the channel with better performance, that is, preferentially scheduling the channel with higher performance stability; the scheduling priority of the fifth channel being higher than that of the sixth channel can be understood as preferentially scheduling the channel with a smaller load.

[0123] In practical applications, when the first process schedules the NCCL tasks in the task queue, it can automatically perform a specific number of failure retries, thereby further improving the reliability of the NCCL system.

[0124] Based on this, in one embodiment, scheduling the NCCL tasks in the task queue based on the scheduling priority of each channel may include:

[0125] For each NCCL task in the task queue, based on the scheduling priority of each channel, select one or more channels for the corresponding NCCL task, and schedule one or more first devices associated with the one or more channels to execute the corresponding NCCL task;

[0126] In the case where the target NCCL task fails to execute, schedule the one or more first devices to re-execute the target NCCL task;

[0127] In the case where the target NCCL task still fails after being re-executed N times, re-add the target NCCL task to the task queue to re-schedule the target NCCL task, and determine the state of each channel in the one or more channels as an unavailable state. A channel in an unavailable state cannot be scheduled, where N is an integer greater than 0.

[0128] Among them, in practical applications, the specific value of N can be set as needed, such as 3, and the embodiments of the present application do not limit this.

[0129] In practical applications, the first process can also perform predictive fault detection based on the first information periodically reported by each second process. Specifically, based on the first information periodically reported by each second process, the first process can use technologies such as machine learning and data analysis to monitor the operating status of the NCCL system in real time and can predict in advance possible faults of the NCCL system; in this way, countermeasures can be taken before the faults occur, the faults can be avoided in advance, the downtime of the NCCL system can be reduced, and thus the reliability of the NCCL system can be further improved.

[0130] Correspondingly, an embodiment of the present application also provides an NCCL task scheduling method, which is applied to a second process in an NCCL process group, as Figure 2 shown, the method includes:

[0131] Step 201: Monitor the operating condition and / or performance of the first device to obtain first information, where the first information includes information related to the operating condition and / or performance of the first device, and at least one corresponding second process is running on the first device;

[0132] Step 202: Periodically report the first information to the first process in the NCCL process group for the first process to schedule NCCL tasks in the task queue; the first process is the main process of the NCCL process group, and the second process includes other processes in the NCCL process group except the first process.

[0133] Wherein, in practical applications, since at least one corresponding second process is running on the first device, the above NCCL task scheduling method can also be applied to the first device.

[0134] In one embodiment, the method may further include:

[0135] Receiving seventh information sent by the first process, where the seventh information is used to indicate a second NCCL topology, and the second NCCL topology is the topology of the communication network corresponding to the NCCL process group. The second NCCL topology is the topology obtained by the first process reconstructing the topology and channels of the communication network when a first parameter is greater than or equal to a first threshold. The first parameter represents the ratio of the number of channels in an unavailable state in the first NCCL topology to the total number of channels, and the second NCCL topology does not include the first device and channels in the unavailable state in the first NCCL topology.

[0136] In practical applications, the method may further include:

[0137] Receive the eighth information sent by the first process, where the eighth information is used to indicate that a relevant channel has failed and / or to indicate the stop of using the relevant channel, and the relevant channel can be understood as the channel where the first device is located.

[0138] The NCCL task scheduling method provided by the embodiments of this application. The first process in the NCCL process group monitors one or more of the second information, third information, and fourth information of each channel in the first NCCL topology based on the first information periodically reported by each second process in the NCCL process group. The first information includes information related to the operating status and / or performance of the first device, and at least the corresponding second process runs on the first device. The second information represents whether the corresponding channel has failed, the third information represents the performance of the corresponding channel, and the fourth information represents the load condition of the corresponding channel; the first NCCL topology is the topology of the communication network corresponding to the NCCL process group, the first process is the main process of the NCCL process group, and the second processes include other processes in the NCCL process group except the first process; based on one or more of the second information, third information, and fourth information of each channel, determine or adjust the scheduling priority of the corresponding channel; based on the scheduling priority of each channel, schedule the NCCL tasks in the task queue. In the solution provided by the embodiments of this application, each other process (i.e., the second process) in the NCCL process group except the main process (i.e., the first process) periodically reports information related to the operating status and / or performance of its own device (i.e., the first device) (i.e., the first information) to the main process. The main process monitors one or more of the failure conditions (i.e., whether a failure has occurred, that is, the second information), performance (i.e., the third information), and load conditions (i.e., the fourth information) of each channel in the corresponding NCCL topology (i.e., the first NCCL topology) based on the reported information, determines or adjusts the scheduling priority of the corresponding channel based on the monitored content, and schedules the NCCL tasks in the task queue based on the scheduling priority of each channel; in this way, the main process can monitor one or more of the failure conditions, performance, and load conditions of each channel in real time according to the real-time operating status and / or performance of each device in the NCCL system, and can dynamically adjust the scheduling priority of the corresponding channel according to the monitored content, so as to be able to implement scheduling strategies (which can also be understood as scheduling rules) such as preferentially allocating NCCL tasks to channels with low failure rates, stable performance, and small loads, thereby being able to dynamically adjust the load distribution of NCCL tasks among channels, that is, being able to implement dynamic load adjustment of each channel in the NCCL system, and further being able to improve the overall communication efficiency and reliability of the NCCL system.

[0139] The following further describes this application in detail with application examples.

[0140] In this application example, the first process is called the main process, and the second process is called the ordinary process.

[0141] This application example proposes a new built-in fault tolerance and recovery mechanism based on NCCL, which can automatically detect and diagnose potential faults in distributed collective communication computing tasks (abbreviated as tasks in the following description), and at the same time has the capabilities of dynamic load balancing and NCCL communication channel (abbreviated as channel in the following description) adjustment. Through this mechanism, many shortcomings in the related technologies can be overcome, such as low efficiency, inconsistent fault tolerance and recovery logic, and limitations of the fault tolerance and recovery logic written by the upper layer.

[0142] Specifically, the above-mentioned built-in fault tolerance and recovery mechanism based on NCCL is implemented based on Figure 3 the NCCL fault detection and diagnosis architecture shown, and has at least the following six functions:

[0143] Function 1, fault detection and diagnosis function;

[0144] Function 2, fault prevention and isolation function;

[0145] Function 3, dynamic load adjustment function;

[0146] Function 4, failure retry function;

[0147] Function 5, function of reloading communication topology and channels;

[0148] Function 6, logging and alarm function.

[0149] The following will describe in detail the fault detection and diagnosis function (i.e., Function 1).

[0150] In this application example, as Figure 3 shown, each ordinary process in the NCCL system can monitor the device status (which can be understood as the running status) of the node where it is located (i.e., the above-mentioned first device) in real time, regularly collect (i.e., detect) the running status of hardware such as GPUs, NICs, and memories through a timing acquisition module, and regularly collect and report the device status, performance metrics, error logs, etc. (i.e., the above-mentioned first information) to the main process through a timing reporting module. After the main process receives the information reported by the ordinary process (i.e., the above-mentioned first information) through the receiving module, it can cache and record this information in a cache (which can be expressed in English as cache), and evaluate the status of the corresponding device through a timing analysis module, that is, determine whether the corresponding device has failed, that is, determine whether the corresponding device is available.

[0151] Among them, each ordinary process can collect metrics through internal monitoring. That is, each NCCL ordinary process can use a local monitoring thread to periodically detect the status of hardware such as GPUs, NICs, and memory on the node where the process is located. The following is an example of the detection method for each device status.

[0152] Each ordinary process can perform GPU status detection. Exemplarily, the CUDA API cudaGetDeviceProperties can be called to obtain information such as GPU utilization, temperature, and memory usage, and cudaDeviceSynchronize() can be called to detect GPU execution errors. Additionally, GPU error records can be detected by parsing the / var / log / syslog log.

[0153] Each ordinary process can perform NIC status detection. Exemplarily, for ordinary network cards, ethtool can be used to obtain statistical information (such as the number of error packets and dropped packets), nginx / apache, etc. can be used to obtain network traffic and bandwidth utilization, and ping can be used to detect packet loss rate and latency. Additionally, network stack parameters (such as socket buffer size) can be checked to detect TCP connection anomalies / disconnections. For remote direct memory access (RDMA) network cards, ibv_devinfo can be used to obtain RDMA port status and rate, ibstat can be used to obtain RDMA error statistics (such as CRC errors), and perftest can be used to test RDMA-related performance. Additionally, disconnections of queue pairs (QPs) can be detected.

[0154] Each ordinary process can perform memory status detection. Exemplarily, available memory can be obtained from the MemAvailable field by parsing / proc / meminfo. Additionally, the memory usage rate can be calculated using the formula "(total memory - available memory) / total memory". Memory leaks can also be detected using the gperftools malloc debugging library, and memory allocation failure situations can be tracked.

[0155] Each ordinary process can perform PCIE and NVLINK status detection. Exemplarily, lspci can be used to obtain information about the PCIe devices where each GPU is located, and check whether the PCIe link speed and bandwidth are consistent with the GPU's expectations; the error counter, such as correctable errors, can also be viewed in / sys / class / pci_bus; the PCIe bandwidth and latency can also be checked by running GPU benchmarks. Additionally, the nvidia-sminvlink command can be used to check the NVLink link status, detect transmit / receive throughput and error counter; NVLink test programs, such as bandwidthTest, can also be run, and the CUDA API can be called to confirm whether Peer-to-Peer access is normal and detect whether GPU direct memory transfer fails.

[0156] Each ordinary process can also perform other relevant detections. Exemplarily, tools such as top and vmstat can be used to monitor the overall CPU usage rate, pidstat can be used to view the CPU usage of each process, iostat can be used to count disk read / write throughput and utilization rate, and the load average value can be monitored to judge the I / O pressure; system call and API error codes can also be detected, and process signals, such as segmentation fault, can be captured; the network connection status can also be detected using periodic heartbeat packets.

[0157] In this application example, each ordinary process can also collect performance metrics (such as utilization rate, throughput rate, etc.) and error logs (such as memory exceptions, network packet loss, etc.), and report various statuses, performance metrics, and error logs and other information (i.e., the above-mentioned first information) detected and collected to the main process at a specific frequency (such as the above-mentioned first period) for the main process to perform fault diagnosis.

[0158] Specifically, the main process can receive and maintain in memory information such as the status, performance, and error logs of each process in all ordinary processes (i.e., the above-mentioned first information). And for the internal monitoring metrics reported by ordinary processes, the fault / abnormality criteria (i.e., conditions) of each metric can be customized in the main process as needed, that is, the fault / abnormality situations of each metric are preset in advance. For example, a sudden significant drop in throughput rate may indicate a network fault, etc.; the main process can perform fault diagnosis based on these preset fault / abnormality criteria.

[0159] Exemplarily, the above preset fault / abnormality criteria (i.e., conditions) may include:

[0160] For GPU status detection, in the case where one or more of the following conditions are met: GPU utilization rate is greater than 90%, GPU core temperature is greater than 80 °C, GPU memory utilization rate is greater than 90%, and the CUDA API returns an error code, it is determined that there is a fault / abnormality;

[0161] For NIC status detection, in the case where one or more of the following conditions are met: the number of error packets increases by more than 50% compared to the same period yesterday, the number of dropped packets increases by more than 50% compared to the same period yesterday, the network bandwidth utilization rate is lower than 20% of the usual average, the ping packet loss rate exceeds 20%, and the number of abnormal TCP connection disconnections increases by more than 50%, it is determined that there is a fault / abnormality;

[0162] For memory status detection, in the case where one or both of the following conditions are met: the proportion of available memory is lower than 20% and / or the number of memory allocation failures increases by more than 50%, it is determined that there is a fault / abnormality;

[0163] For other related detections, in the case where one or more of the following conditions are met: the average CPU utilization rate exceeds 90%, the single-process CPU utilization rate continuously exceeds 50%, the disk read / write throughput decreases by more than 50%, and important system calls or API error codes appear, it is determined that there is a fault / abnormality.

[0164] In this application example, for each device, the main process can determine whether there is a fault / abnormality in the device according to the above preset fault / abnormality criteria and make a mark; that is, the main process can generate a fault diagnosis result for each device and make a fault / abnormality mark. Moreover, the main process can configure / adjust the status update frequency (which can also be understood as adjusting the status reporting frequency, equivalent to adjusting the above first period). Exemplarily, the main process can reasonably set the status reporting frequency according to the node scale and network conditions. For a deployment scenario with a small number of nodes (such as less than 10 nodes), a shorter reporting period can be set (such as reporting once every 5 seconds); when the number of nodes gradually increases to about 100, the reporting period can be appropriately increased, such as reporting once every 10 seconds; if the number of nodes is very large (such as more than 1000), the reporting period can be further increased, such as reporting once every 30 seconds. Here, customizing the reporting period for various scales is supported. In addition, NCCL can also provide an API to adjust the reporting frequency, that is, the reporting frequency can be dynamically adjusted according to the fault risk situation through the API, such as increasing the reporting frequency when the risk increases.

[0165] The following details the fault prevention and isolation function (i.e., Function 2).

[0166] In this application example, after the above information reporting and fault detection process, the main process can identify unavailable devices based on the received data (i.e., the above first information), and can also set the channels where these devices are located to the unavailable state (i.e., determine the above second information), and set the unavailable devices and unavailable channels as non-schedulable, so as to prevent tasks from being scheduled to faulty channels and avoid causing more serious problems.

[0167] Specifically, first, the main process can perform fault device identification, that is, the main process can determine the faulty devices according to the fault diagnosis results and identify the faulty devices; for example, identify that GPU1 has failed, NIC2 has serious transmission errors, etc. Here, temporary faults and permanent faults can be distinguished. A fault can be first marked as a temporary fault for retry. If the retry fails 3 times, it can be marked as a permanent fault and immediately isolated.

[0168] After that, the main process can perform channel isolation. A channel is a communication path in the NCCL topology, which is formed by connecting communication devices and GPU devices in series through network links such as PCIE, NVLINK links, sockets, and rdma. The main process can lock the status of relevant channels as unavailable according to the information of faulty devices (such as network devices, link devices, GPU devices, etc.), and can update the channel status table to notify each ordinary process to stop using the faulty channels. Here, the channels corresponding to different types of faults (i.e., temporary faults and permanent faults) can be set to different levels of unavailable states (i.e., temporarily unavailable state and permanently unavailable state).

[0169] After that, the main process can perform task flow limiting, that is, the main process can limit the task flow according to the number of unavailable channels and the load capacity of available channels to prevent overload. Among them, as Figure 4 shown, the main process can perform task scheduling through the task scheduling module (which can also be called a scheduler) according to the real-time channel status table to bypass the faulty channels; and the main process can reserve communication bandwidth in advance for retry and fault tolerance; if a task is already in a faulty channel, the task can be reverted to the task queue (which can also be called a task list, expressed in English as task queue) to wait for a new schedule.

[0170] In this application example, the main process can perform fault notification, that is, the main process can send notifications to tasks using faulty devices to prompt possible execution failures. And the main process can provide a fault query interface for returning a list of unavailable devices. In addition, the main process can perform the recovery of automatic isolation, that is, for temporary faults, if the retry is successful, that is, if it can be confirmed that the fault has been recovered, the channel can be reopened.

[0171] The following will describe the dynamic load adjustment function (i.e., function 3) in detail.

[0172] In the related art, the task channel selection method supported by NCCL is to poll and select the channels that can carry tasks in the current channels, which may ultimately result in uneven channel loads. Therefore, in this application example, the main process can dynamically adjust the load distribution of NCCL communication tasks among channels according to the real-time status and performance of the devices, and add scheduling strategies such as preferentially allocating communication tasks to channels with stable performance and low failure rates, so as to improve the overall communication efficiency and reliability of the system.

[0173] First, the main process can evaluate the performance of the channels, that is, determine the above-mentioned third information. Specifically, the main process can continuously evaluate the performance of each channel, such as bandwidth, latency, packet loss rate, etc.; and, the main process can calculate the average value and change trend of the performance metrics within a sliding time window. The calculation method can include: setting the length of the sliding time window, such as 10 minutes; calculating the average value of the performance metrics in the past 10 minutes every 1 minute, such as the average bandwidth, average latency, average packet loss rate, etc. in the past 10 minutes; calculating the change ratio of the performance metrics between two adjacent windows, such as the change amplitude of the current window's average bandwidth compared with the previous window. If the change amplitude exceeds a specific threshold (this threshold can be set according to the actual situation, such as 20%), it means that the performance fluctuates too much, that is, the performance is unstable; after calculating the average metrics and change metrics of each window in chronological order, an index change curve can be drawn, and the stability of each channel can be evaluated through the curve. The channel with a stable change trend has better stability; or, metrics such as mean square error and standard deviation can be used to evaluate the degree of dispersion of the metrics within each window. The smaller the degree of dispersion, the more stable the performance.

[0174] After that, the main process can comprehensively evaluate each channel based on the average value, change trend, and dispersion degree, and use the evaluation results as the assessment of stability and failure rate. Specifically, when performing the stability evaluation, the coefficient of variation (CV) of the performance metrics (such as bandwidth, latency, packet loss rate, etc.) of each channel within a recent period of time (such as 1 hour) can be calculated first. The smaller the CV value, the smoother the change and the better the stability. If the CV exceeds a specific threshold (such as 20%), it is considered unstable. Based on the CV value, a score can be given to the stability of each channel. The scoring formula can be Score_stability = K * (1 / CV), where K represents the coefficient to adjust the score range and can be set according to needs, such as 0 - 100. That is, the stability score (Score_stability) ranging from 0 to 100 can be finally obtained, and the higher the score, the more stable the channel. When performing the failure rate evaluation, the number of failures of the channel within a recent specific period of time (such as 1 day) can be counted first. A channel without failures gets a full score of 100 points; for each additional failure, 10 points are deducted from the full score. Finally, a failure rate score ranging from 0 to 100 can be obtained, and the higher the score, the lower the failure rate. Finally, the two evaluation results can be weighted to obtain a comprehensive score ranging from 0 to 100.

[0175] After that, the main process can sort the channels by priority. Specifically, it can monitor the channel load (i.e., the above fourth piece of information), and first sort the channels according to the performance evaluation results (i.e., the above scores). The weights of the performance metrics can be configured according to needs, and stability or throughput can be given priority. After that, in the case of channel failures or when the channel load is too heavy, the priority of the channel will be lowered. Based on the channel priority, the main process can perform load distribution, that is, new tasks can select channels according to the channel priority list and the load situation (i.e., the priority list can be adjusted according to the load situation); more tasks can be assigned to high-priority channels, and low-priority channels can reduce or suspend scheduling. Alternatively, the main process can distribute the task traffic to multiple channels according to weights to avoid overloading a single channel.

[0176] In this application example, the main process can dynamically adjust the load. Specifically, the main process can monitor the channel status in real time, automatically switch and / or lower the channel priority when a failure occurs; and the main process can detect newly added channels based on changes in the network topology, evaluate their priorities, and then add them to the scheduling. In addition, NCCL can provide an API, and this API can be used to manually adjust the channel priority.

[0177] The following is a detailed description of the failure retry function (i.e., function 4).

[0178] In the related art, since the task failure retry function is not supported, once a network storm or device instability occurs, it may directly cause NCCL to crash. Therefore, in this application example, each task can be retried in case of failure and can cover various device failures. Specifically, when a task is in progress, in case of communication failure or device failures such as GPU, it can be retried first, and the default number of failure retries is set to 3 (an environment variable for the maximum number of failure retries can be provided for custom modification); the failure retry can be implemented in combination with the communication timeout, that is, exceeding the communication timeout means the retry fails; after multiple retries and the number of failure times is greater than the preset number of failure retries, the task being retried can be set (i.e., determined) as failed, and then it can be reverted to the queue to queue up and wait for re-scheduling and execution; the main process can also determine the channel where the task is located as a faulty channel and set it as non-schedulable.

[0179] Among them, the main process can control the retry process, that is, start the retry process and interrupt the current communication task when a failure occurs. Also, the upper limit of the number of retries can be preset, which can be 3 by default, and the number of retries can also be configured through environment variables. In addition, the number of retries that have been performed can be recorded, and combined with the set custom communication timeout (which can be 15 seconds by default) to determine whether the retry fails.

[0180] The main process can also handle faulty channels, that is, determine the channel failure when consecutive retry failures occur and set the channel status as unavailable. Also, the task scheduling module can be notified of the update of the faulty channel, and the temporarily faulty channel can attempt to automatically recover after the custom recovery time (which can be 10 minutes by default).

[0181] The main process can also handle tasks, that is, terminate the task when the retry limit is reached and return a task failure message. Also, the failed task can be re-added to the task queue or added to the retry queue to wait for retry. In addition, a custom failure retry policy can be set for each task.

[0182] The functions of overloading the communication topology and channels (i.e., Function 5) will be described in detail below.

[0183] In the related art, once a failure occurs in NCCL (such as the unavailability of resources like devices or networks), it may directly report an error and cause the program to crash. Therefore, this application example supports NCCL to remove unavailable devices and reload the communication topology and channels when a failure occurs, so as to support NCCL to continue running without restarting. Specifically, as the number of failed devices increases and the number of unavailable channels rises, when the unavailable ratio of the channels reaches a preset threshold (the default can be 20%, and custom settings are supported), the system (specifically, it can be executed by the main process) will remove the unavailable devices and reconstruct the communication topology and channels. Let \(P_{unavailable}\) represent the unavailable ratio of the channels, \(N_{unavailable}\) represent the number of unavailable channels, \(N_{total}\) represent the total number of channels, and \(T_{max}\) represent the maximum unavailable ratio threshold of the channels (such as 20%, that is, the above first threshold). Then, the unavailable channel ratio can be calculated using the formula \(P_{unavailable}=N_{unavailable} / N_{total}*100\%\). If the main process determines that \(P_{unavailable}\) is greater than or equal to \(T_{max}\), the communication topology and channels can be reconstructed, and the reconstructed topology (i.e., the above second NCCL topology) and channels will remove the unavailable devices. Among them, when the main process calculates the unavailable channels, it can continuously count the number of unavailable channels \(N_{unavailable}\) and the total number of channels \(N_{total}\), calculate the unavailable channel ratio using the above formula, and track the change trend of the ratio value to smooth out the instantaneous fluctuations.

[0184] Here, for the above ratio threshold (i.e., the above first threshold), the maximum allowable unavailable channel ratio threshold can be set, and the default can be 20%; moreover, NCCL can provide an API to adjust the threshold, allowing customization of the threshold, and the same or different thresholds can be set for different NCCL cluster environments. In addition, when performing topology reconstruction, the main process can trigger the communication topology reconstruction when the unavailable channel ratio exceeds the threshold, remove the devices and channels marked as unavailable, re-plan the communication topology according to the remaining resources, and optimize resource utilization; after topology reconstruction, the number and numbering of the channels can be recalculated, and all ordinary processes can be notified of the updated communication topology and channels, and each process can re-initialize the communication environment and connections; after topology reconstruction, the restored channels can also be continuously checked and re-integrated into the topology.

[0185] The following describes the log recording and alarm function (i.e., function 6) in detail.

[0186] This application example provides preset alarm call input parameters (i.e., the above-mentioned fifth piece of information). When a device failure or abnormal device status is detected, the system (specifically, it can be executed by the main process) can record the fault-related information (the format of this information can support customization) in the log, and the system can automatically send alarm information (i.e., the above-mentioned sixth piece of information) to notify relevant personnel to perform corresponding processing in a timely manner. Specifically, the status, events, error logs, etc. (i.e., the above-mentioned first piece of information) reported by all ordinary nodes (i.e., the above-mentioned first device) can be uniformly recorded in the log database of the master node (i.e., the above-mentioned second device); this database can provide a query interface, which can be filtered and retrieved by keywords such as time, node, event, etc., and can support log classification, and important logs can be recorded separately. Each log can contain information such as a timestamp, node identifier (such as ID, etc.), and event details. Also, for the pre-configured alarm rules (i.e., the above-mentioned fifth piece of information), different alarm rules for different levels of faults can be specifically set, such as node disconnection, GPU failure, etc.; different levels of faults can notify the same or different receiving groups; the alarm rules can customize the fault detection logic, notification method, etc.; it can also support configuring the interval time for alarm sending to prevent alarm storms. In addition, when sending alarm information, alarm notification methods such as email and SMS can be supported; the contact information of the administrator can be preset, and notifications can be actively sent when an alarm is required. Here, a fault ticket can be generated for each fault event to record the processing process; the administrator can update the status of the fault ticket to indicate whether the problem has been resolved; it can also support generating various fault statistical reports, such as the number of faults, processing duration, etc., so as to help the administrator analyze the system reliability and formulate optimization strategies.

[0187] In this application example, the overall process of the above-mentioned NCCL fault tolerance and recovery mechanism can be as Figure 5As shown in the figure. First, each ordinary process (i.e., the second process mentioned above) can regularly report information such as the status of its own local device (i.e., the first device mentioned above) (i.e., the first information). The main process (i.e., the first process mentioned above) is responsible for storing and scanning this device status information, etc., and monitoring the device operation status in real time and making annotations. After that, when it is found that a device is unavailable, the main process can mark the status of the device as unavailable; if the alarm phone number information (i.e., the fifth information mentioned above) has been entered (i.e., configured), the system (specifically, it can be executed by the main process) can send an alarm notification (i.e., the sixth information mentioned above) to the corresponding phone; at the same time, the main process can mark the channel containing the unavailable device as unavailable (i.e., determine the second information mentioned above), and detect whether the proportion of unavailable channels exceeds a preset threshold (i.e., the first threshold mentioned above); if it exceeds the threshold, the system (specifically, it can be executed by the main process) can start the reconstruction of the communication topology and channels, and the newly constructed topology (i.e., the second NCCL topology mentioned above) and channels will exclude the unavailable devices. After that, within the range of the proportion of unavailable normal devices and channels, the system (specifically, it can be executed by the main process) can select a channel list according to the channel performance (i.e., the third information mentioned above) and / or the channel load (i.e., the fourth information mentioned above), and preferentially use the channels with smaller load. After that, the task can be run on the selected channel; if the task is successfully executed, the process ends; if the task execution fails, then retry; if the retry is successful within the maximum number of retries, the process ends, otherwise the task can be returned to the task list, and the corresponding channel is marked as unavailable.

[0188] The solution provided by this application example has the following advantages:

[0189] 1) Through real-time monitoring and diagnostic methods, this application example automatically detects and processes potential faults, reducing the time and effort required for manual log viewing and problem troubleshooting. It can not only reduce the operation and maintenance burden and cost, but also improve the recovery speed of NCCL tasks, and solve fault problems more quickly during the execution of NCCL tasks, thereby improving the overall operation efficiency of the NCCL system and ensuring business continuity;

[0190] 2) This application example provides unified fault tolerance and recovery logic (i.e., fault handling standards) for different tasks and frameworks, reducing the difficulty of maintenance and collaboration; moreover, through a unified interface to implement fault tolerance and recovery, it can ensure that various tasks can obtain consistent fault tolerance and recovery support when facing faults;

[0191] 3) This application example designs an internal fault tolerance and recovery mechanism for NCCL. Compared with the simple handling of NCCL faults by the upper-layer logic (such as restarting, etc.), it can achieve more refined and effective fault tolerance and recovery;

[0192] 4) This application example enables NCCL to have a fault tolerance and recovery mechanism internally. Developers can more conveniently integrate the fault tolerance and recovery functions described in this application example without developing the NCCL fault recovery logic at the upper layer, thus reducing the development difficulty and the usage and learning costs of NCCL.

[0193] To implement the method on the first process side of the embodiments of the present application (i.e., to implement the method on the second device side of the embodiments of the present application), the embodiments of the present application further provide an NCCL task scheduling device, which is arranged on the second device. The second device runs at least the first process in the NCCL process group, and the first process is the main process of the NCCL process group; as Figure 6 shown, the device includes:

[0194] A first processing unit 601, which is used, after being utilized by the first process, to monitor one or more of the second information, third information, and fourth information of each channel in the first NCCL topology based on the first information periodically reported by each second process in the NCCL process group. The first information includes information related to the operating status and / or performance of the first device, and the first device runs at least the corresponding second process. The second information represents whether the corresponding channel has a fault, the third information represents the performance of the corresponding channel, and the fourth information represents the load condition of the corresponding channel; the first NCCL topology is the topology of the communication network corresponding to the NCCL process group, and the second process includes other processes in the NCCL process group except the first process;

[0195] A second processing unit 602, which is used, after being utilized by the first process, to determine or adjust the scheduling priority of the corresponding channel based on one or more of the second information, third information, and fourth information of each channel;

[0196] A scheduling unit 603, which is used, after being utilized by the first process, to schedule the NCCL tasks in the task queue based on the scheduling priority of each channel.

[0197] Wherein, in one embodiment, the second processing unit 602 is specifically used, after being utilized by the first process, to determine or adjust the scheduling priority of the corresponding channel by using one or more of the following rules:

[0198] The scheduling priority of the first channel is higher than that of the second channel. The first channel has no fault, and the second channel has a fault;

[0199] The scheduling priority of the third channel is higher than that of the fourth channel. The performance of the third channel is better than that of the fourth channel;

[0200] The scheduling priority of the fifth channel is higher than that of the sixth channel, and the load of the fifth channel is less than that of the sixth channel.

[0201] In one embodiment, the first processing unit 601 is specifically configured to perform the following operations after being utilized by the first process:

[0202] Based on the first information reported by each second process, perform fault diagnosis on each first device in the first NCCL topology to obtain a fault diagnosis result;

[0203] For each first device, based on the fault diagnosis result, determine the second information of the channel associated with the corresponding first device; wherein, when the fault diagnosis result indicates that the corresponding first device has a fault, the second information of the channel associated with the corresponding first device indicates that the corresponding channel has a fault.

[0204] In one embodiment, as Figure 6 shown, the device may further include:

[0205] The first communication unit 604 is used to perform the following operations after being utilized by the first process:

[0206] When the fault diagnosis result indicates that the corresponding first device has a fault, obtain preset fifth information, where the fifth information includes alarm-related information;

[0207] Based on the fifth information, send out sixth information, where the sixth information is used to perform fault alarm for the corresponding first device and / or the corresponding channel.

[0208] In one embodiment, the first processing unit 601 is further used to perform the following operations after being utilized by the first process:

[0209] For each first device, when the fault diagnosis result indicates that the corresponding first device has a fault, determine the state of the corresponding first device as an unavailable state, and determine the state of the channel associated with the corresponding first device as an unavailable state. The first device and the channel in the unavailable state cannot be scheduled;

[0210] When the first parameter is greater than or equal to the first threshold, reconstruct the topology and channels of the communication network to obtain a second NCCL topology; the first parameter represents the ratio of the number of channels in the unavailable state in the first NCCL topology to the total number of channels, and the second NCCL topology does not include the first devices and channels in the unavailable state in the first NCCL topology;

[0211] Correspondingly, the scheduling unit 603 is further used to schedule the NCCL tasks in the task queue based on the second NCCL topology after being utilized by the first process;

[0212] The first communication unit 604 is further configured to, after being utilized by the first process, send seventh information to each second process, where the seventh information is used to indicate the second NCCL topology.

[0213] In one embodiment, the first processing unit 601 is further configured to, after being utilized by the first process, perform the following operations:

[0214] In the case where the fault diagnosis result indicates that a corresponding first device has a fault, determine the state of the corresponding first device as a temporarily unavailable state, determine the state of the channel associated with the corresponding first device as a temporarily unavailable state, and retry the fault diagnosis for the corresponding first device;

[0215] In the case where it is still determined that the corresponding first device has a fault after retrying the fault diagnosis for the corresponding first device M times, determine the state of the corresponding first device as a permanently unavailable state, and determine the state of the channel associated with the corresponding first device as a permanently unavailable state, where M is an integer greater than 0.

[0216] In one embodiment, the scheduling unit 603 is further configured to, after being utilized by the first process, perform the following operations:

[0217] For each NCCL task in the task queue, based on the scheduling priority of each channel, select one or more channels for the corresponding NCCL task, and schedule one or more first devices associated with the one or more channels to execute the corresponding NCCL task;

[0218] In the case where the target NCCL task fails to execute, schedule the one or more first devices to re - execute the target NCCL task;

[0219] In the case where it still fails to execute after re - executing the target NCCL task N times, re - add the target NCCL task to the task queue to re - schedule the target NCCL task, and determine the state of each channel in the one or more channels as an unavailable state, and the channels in the unavailable state cannot be scheduled, where N is an integer greater than 0.

[0220] In practical applications, it can be understood that the first communication unit 604 is further configured to, after being utilized by the first process, receive the first information periodically reported by each second process in the NCCL process group.

[0221] Here, the function of the first processing unit 601 is equivalent to that of the timing analysis module in the above application example; the functions of the second processing unit 602 and the scheduling unit 603 are equivalent to those of the task scheduling module in the above application example; the function of the first communication unit 604 is equivalent to that of the receiving module in the above application example.

[0222] In practical applications, the first processing unit 601 and the second processing unit 602 can be implemented by the processor in the NCCL task scheduling device; the scheduling unit 603 can be implemented by the processor in the NCCL task scheduling device in combination with the communication interface; the first communication unit 604 can be implemented by the communication interface in the NCCL task scheduling device.

[0223] To implement the method on the second process side of the embodiment of the present application (i.e., to implement the method on the first device side of the embodiment of the present application), the embodiment of the present application also provides an NCCL task scheduling device, which is set on the first device. The first device runs at least the corresponding second process in the NCCL process group. The second process includes other processes in the NCCL process group except the first process, and the first process is the main process of the NCCL process group. As Figure 7 shown, the device includes:

[0224] A third processing unit 701, configured to monitor the operating status and / or performance of the first device after being used by the second process, and obtain first information, where the first information includes information related to the operating status and / or performance of the first device;

[0225] A second communication unit 702, configured to periodically report the first information to the first process after being used by the second process, so that the first process schedules the NCCL tasks in the task queue.

[0226] Wherein, in one embodiment, the second communication unit 702 is further configured to receive seventh information sent by the first process after being used by the second process. The seventh information is used to indicate a second NCCL topology, and the second NCCL topology is the topology of the communication network corresponding to the NCCL process group. The second NCCL topology is the ratio of the number of channels to the total number of channels of the communication network reconstructed by the first process when a first parameter is greater than or equal to a first threshold. The second NCCL topology does not include the topology obtained by removing the first NCCL topology and channels, and the first parameter characterizes the first device and channels in an unavailable state in the first NCCL topology.

[0227] Here, the function of the third processing unit 701 is equivalent to that of the timing acquisition module in the above application example; the function of the second communication unit 702 is equivalent to that of the timing reporting module in the above application example.

[0228] In practical applications, the third processing unit 701 can be implemented by a processor in the NCCL task scheduling device; the second communication unit 702 can be implemented by a communication interface in the NCCL task scheduling device.

[0229] It should be noted that when the NCCL task scheduling device provided in the above embodiment performs NCCL task scheduling, only the division of the above program modules is used for illustration. In practical applications, the above processing can be allocated to different program modules according to needs, that is, the internal structure of the device is divided into different program modules (such as the timing analysis module, task scheduling module, receiving module in the above application example, and again such as the timing acquisition module, timing reporting module in the above application example) to complete all or part of the processing described above. In addition, the NCCL task scheduling device provided in the above embodiment and the embodiment of the NCCL task scheduling method belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be elaborated here.

[0230] Based on the hardware implementation of the above program modules, and in order to implement the method on the first process side of the embodiment of the present application (that is, to implement the method on the second device side of the embodiment of the present application), the embodiment of the present application also provides a second device, as Figure 8 shown, the second device 800 includes:

[0231] A first communication interface 801 capable of information interaction with other electronic devices (such as the first device, etc.);

[0232] A first processor 802, connected to the first communication interface 801 to achieve information interaction with other electronic devices, and when used to run a computer program, execute the method provided by one or more technical solutions on the first process side above, that is, execute the method provided by one or more technical solutions on the second device side above;

[0233] A first memory 803, on which the computer program is stored.

[0234] Among them, the second device 800 runs at least the first process in the NCCL process group, and the first process is the main process of the NCCL process group.

[0235] Specifically, the first processor 802 is used to perform the following operations after being utilized by the first process:

[0236] Monitor one or more of the second information, third information, and fourth information of each channel in the first NCCL topology based on the first information periodically reported by each second process in the NCCL process group. The first information includes information related to the operating status and / or performance of the first device, and at least the corresponding second process is running on the first device. The second information characterizes whether the corresponding channel has a fault, the third information characterizes the performance of the corresponding channel, and the fourth information characterizes the load condition of the corresponding channel. The first NCCL topology is the topology of the communication network corresponding to the NCCL process group, and the second process includes other processes in the NCCL process group except the first process.

[0237] Determine or adjust the scheduling priority of the corresponding channel based on one or more of the second information, third information, and fourth information of each channel.

[0238] Schedule the NCCL tasks in the task queue based on the scheduling priority of each channel.

[0239] Here, it can be understood that the first communication interface 801 can be used by the first process to receive the first information periodically reported by each second process in the NCCL process group.

[0240] Wherein, in one embodiment, the first processor 802 is further configured to, after being utilized by the first process, determine or adjust the scheduling priority of the corresponding channel by using one or more of the following rules:

[0241] The scheduling priority of the first channel is higher than that of the second channel. The first channel has no fault, and the second channel has a fault.

[0242] The scheduling priority of the third channel is higher than that of the fourth channel. The performance of the third channel is better than that of the fourth channel.

[0243] The scheduling priority of the fifth channel is higher than that of the sixth channel. The load of the fifth channel is less than that of the sixth channel.

[0244] In one embodiment, the first processor 802 is further configured to, after being utilized by the first process, perform the following operations:

[0245] Perform a fault diagnosis on each first device in the first NCCL topology based on the first information reported by each second process to obtain a fault diagnosis result.

[0246] For each first device, based on the fault diagnosis result, determine second information of the channel associated with the corresponding first device; wherein, when the fault diagnosis result indicates that the corresponding first device has a fault, the second information of the channel associated with the corresponding first device indicates that the corresponding channel has a fault.

[0247] In one embodiment, the first communication interface 801 is further configured to perform the following operations after being utilized by the first process:

[0248] When the fault diagnosis result indicates that the corresponding first device has a fault, obtain preset fifth information, where the fifth information includes alarm-related information;

[0249] Based on the fifth information, send sixth information, where the sixth information is used to perform a fault alarm for the corresponding first device and / or the corresponding channel.

[0250] In one embodiment, the first processor 802 is further configured to perform the following operations after being utilized by the first process:

[0251] For each first device, when the fault diagnosis result indicates that the corresponding first device has a fault, determine the state of the corresponding first device as an unavailable state, and determine the state of the channel associated with the corresponding first device as an unavailable state. The first device and the channel in the unavailable state cannot be scheduled;

[0252] When the first parameter is greater than or equal to the first threshold, reconstruct the topology and channels of the communication network to obtain a second NCCL topology; the first parameter represents the ratio of the number of channels in the unavailable state in the first NCCL topology to the total number of channels, and the second NCCL topology does not include the first devices and channels in the unavailable state in the first NCCL topology;

[0253] Schedule the NCCL tasks in the task queue based on the second NCCL topology;

[0254] Correspondingly, the first communication interface 801 is further configured to send seventh information to each second process after being utilized by the first process, where the seventh information is used to indicate the second NCCL topology.

[0255] In one embodiment, the first processor 802 is further configured to perform the following operations after being utilized by the first process:

[0256] When the fault diagnosis result indicates that the corresponding first device has a fault, determine the state of the corresponding first device as a temporarily unavailable state, determine the state of the channel associated with the corresponding first device as a temporarily unavailable state, and retry the fault diagnosis for the corresponding first device;

[0257] In the case that after retrying the M - time fault diagnosis for the corresponding first device, it is still determined that the corresponding first device has a fault, the state of the corresponding first device is determined to be a permanently unavailable state, and the state of the channel associated with the corresponding first device is determined to be a permanently unavailable state, where M is an integer greater than 0.

[0258] In one embodiment, the first processor 802 is further configured to perform the following operations after being utilized by the first process:

[0259] For each NCCL task in the task queue, based on the scheduling priority of each channel, select one or more channels for the corresponding NCCL task, and schedule one or more first devices associated with the one or more channels to execute the corresponding NCCL task;

[0260] In the case that the target NCCL task fails to execute, schedule the one or more first devices to re - execute the target NCCL task;

[0261] In the case that it still fails to execute after re - executing the target NCCL task N times, re - add the target NCCL task to the task queue to re - schedule the target NCCL task, and determine the state of each channel in the one or more channels to be an unavailable state. A channel in the unavailable state cannot be scheduled, where N is an integer greater than 0.

[0262] It should be noted that: The specific processing procedures of the first communication interface 801 and the first processor 802 can be understood with reference to the above - mentioned method, and will not be elaborated here.

[0263] Of course, in actual application, each component in the second device 800 is coupled together through the bus system 804. It can be understood that the bus system 804 is used to realize the connection and communication between these components. The bus system 804 includes, in addition to the data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 8 all kinds of buses are labeled as the bus system 804.

[0264] The first memory 803 in the embodiment of the present application is used to store various types of data to support the operation of the second device 800. Examples of these data include: any computer program for operating on the second device 800.

[0265] The method disclosed in the embodiments of the present application can be applied to or implemented by the first processor 802. The first processor 802 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in the first processor 802. The first processor 802 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The first processor 802 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the first memory 803. The first processor 802 reads the information in the first memory 803 and combines its hardware to complete the steps of the foregoing method.

[0266] In an exemplary embodiment, the second device 800 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontroller units (MCUs), microprocessors, or other electronic components for executing the foregoing method.

[0267] Based on the hardware implementation of the above program module, and in order to implement the method on the second process side of the embodiments of the present application (i.e., to implement the method on the first device side of the embodiments of the present application), the embodiments of the present application further provide a first device, as Figure 9 shown. The first device 900 includes:

[0268] A second communication interface 901 capable of interacting with other electronic devices (such as the second device, etc.) for information.

[0269] A second processor 902, connected to the second communication interface 901 to enable information interaction with other electronic devices, and when used to run a computer program, executes the method provided by one or more of the above-mentioned second process-side technical solutions, that is, executes the method provided by one or more of the above-mentioned first device-side technical solutions;

[0270] A second memory 903, on which the computer program is stored.

[0271] Among them, the first device 900 runs at least the corresponding second process in the NCCL process group, and the second process includes other processes in the NCCL process group except the first process, and the first process is the main process of the NCCL process group.

[0272] Specifically, the second processor 902, after being utilized by the second process, monitors the operating status and / or performance of the first device 900 to obtain first information, and the first information includes information related to the operating status and / or performance of the first device 900;

[0273] The second communication interface 901, after being utilized by the second process, periodically reports the first information to the first process for the first process to schedule NCCL tasks in the task queue.

[0274] Among them, in one embodiment, the second communication interface 901 is further configured to, after being utilized by the second process, receive seventh information sent by the first process, where the seventh information is used to indicate a second NCCL topology, and the second NCCL topology is the topology of the communication network corresponding to the NCCL process group. The second NCCL topology is the topology obtained by the first process reconstructing the topology and channels of the communication network when a first parameter is greater than or equal to a first threshold. The first parameter represents the ratio of the number of channels in an unavailable state in the first NCCL topology to the total number of channels, and the second NCCL topology does not include the first device 900 and channels in the unavailable state in the first NCCL topology.

[0275] It should be noted that: The specific processing procedures of the second communication interface 901 and the second processor 902 can be understood with reference to the above method and will not be elaborated here.

[0276] Of course, in actual application, each component in the first device 900 is coupled together through a bus system 904. It can be understood that the bus system 904 is used to achieve connection communication between these components. The bus system 904 includes not only a data bus but also a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 9 all kinds of buses are labeled as the bus system 904.

[0277] The second memory 903 in the embodiment of the present application is used to store various types of data to support the operation of the first device 900. Examples of such data include: any computer program for operating on the first device 900.

[0278] The method disclosed in the embodiment of the present application above can be applied to the second processor 902 or implemented by the second processor 902. The second processor 902 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the second processor 902 or the instructions in the form of software. The second processor 902 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The second processor 902 can implement or execute each method, step, and logic block diagram disclosed in the embodiment of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the second memory 903. The second processor 902 reads the information in the second memory 903 and combines its hardware to complete the steps of the foregoing method.

[0279] In an exemplary embodiment, the first device 900 can be implemented by one or more ASICs, DSPs, PLDs, CPLDs, FPGAs, general-purpose processors, controllers, MCUs, Microprocessors, or other electronic components for executing the foregoing method.

[0280] It can be understood that the memories (the first memory 803 and the second memory 903) in the embodiments of the present application can be volatile memories or non-volatile memories, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memories.

[0281] To implement the method provided by the embodiments of this application, the embodiments of this application also provide an NCCL task scheduling system, as Figure 10 shown. This system includes: a second device 1001 and one or more first devices 1002.

[0282] Here, it should be noted that: the specific processing procedures of the second device 1001 and the first device 1002 have been described in detail above and will not be elaborated here.

[0283] In an exemplary embodiment, the embodiments of this application also provide a storage medium, namely a computer storage medium, specifically a computer-readable storage medium. For example, it includes a first memory 803 that stores a computer program. The above computer program can be executed by a first processor 802 of a second device 800 to complete the steps of any of the methods on the first process side, that is, to complete the steps of any of the methods on the second device side. Another example is a second memory 903 that stores a computer program. The above computer program can be executed by a second processor 902 of a first device 900 to complete the steps of any of the methods on the second process side, that is, to complete the steps of any of the methods on the first device side. The computer-readable storage medium can be a FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.

[0284] In an exemplary embodiment, the embodiments of this application also provide a computer program product, including a computer program. The computer program can be executed by a first processor 802 of a second device 800 to complete the steps of any of the methods on the first process side, that is, to complete the steps of any of the methods on the second device side; or the computer program can be executed by a second processor 902 of a first device 900 to complete the steps of any of the methods on the second process side, that is, to complete the steps of any of the methods on the first device side.

[0285] It should be noted that: "first", "second", etc. are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence.

[0286] In addition, the technical solutions described in the embodiments of this application can be arbitrarily combined without conflict.

[0287] The above is only a preferred embodiment of this application and is not intended to limit the protection scope of this application.

Claims

1. A NVIDIA collective communication library NCCL task scheduling method, characterized in that: Applies to the first process in an NCCL process group, including: Based on the first information periodically reported by each second process in the NCCL process group, one or more of the second information, the third information, and the fourth information of each channel in the first NCCL topology are monitored, wherein the first information includes relevant information on the operating status and / or performance of the first device, the first device at least runs the corresponding second process, the second information represents whether the corresponding channel fails, the third information represents the performance of the corresponding channel, and the fourth information represents the load of the corresponding channel; the first NCCL topology is the topology of the communication network corresponding to the NCCL process group, the first process is the main process of the NCCL process group, and the second process includes other processes in the NCCL process group except the first process; Determine or adjust the scheduling priority of the corresponding channel based on one or more of the second information, the third information, and the fourth information of each channel; Schedule NCCL tasks in the task queue based on the scheduling priority of each channel.

2. The method according to claim 1, characterized in that The determining or adjusting the scheduling priority of the corresponding channel includes: Use one or more of the following rules to determine or adjust the scheduling priority of the corresponding channel: The scheduling priority of the first channel is higher than the scheduling priority of the second channel, the first channel is not faulty, and the second channel is faulty; The scheduling priority of the third channel is higher than the scheduling priority of the fourth channel, and the performance of the third channel is better than the performance of the fourth channel; The scheduling priority of the fifth channel is higher than the scheduling priority of the sixth channel, and the load of the fifth channel is smaller than the load of the sixth channel.

3. The method according to claim 1, characterized in that The monitoring of one or more of the second information, the third information, and the fourth information of each channel in the first NCCL topology based on the first information periodically reported by each second process in the NCCL process group includes: Based on the first information reported by each second process, perform fault diagnosis on each first device in the first NCCL topology to obtain a fault diagnosis result; For each first device, based on the fault diagnosis result, second information of the channel associated with the corresponding first device is determined; wherein, when the fault diagnosis result indicates that the corresponding first device has a fault, the second information of the channel associated with the corresponding first device indicates that the corresponding channel has a fault.

4. The method according to claim 3, characterized in that The method further comprises: When the fault diagnosis result indicates that the corresponding first device has a fault, obtaining preset fifth information, wherein the fifth information includes alarm-related information; Based on the fifth information, sixth information is issued, and the sixth information is used to issue a fault alarm corresponding to the first device and / or the corresponding channel.

5. The method according to claim 3, characterized in that: The method further comprises: For each first device, when the fault diagnosis result indicates that the corresponding first device has a fault, determining the state of the corresponding first device as an unavailable state, and determining the state of a channel associated with the corresponding first device as an unavailable state, and the first device and the channel in the unavailable state cannot be scheduled; When the first parameter is greater than or equal to the first threshold, reconstructing the topology and channels of the communication network to obtain a second NCCL topology; the first parameter represents the ratio of the number of channels in an unavailable state to the total number of channels in the first NCCL topology, and the second NCCL topology does not include the first device and channel in an unavailable state in the first NCCL topology; Based on the second NCCL topology, scheduling the NCCL tasks in the task queue; Seventh information is sent to each second process, where the seventh information is used to indicate the second NCCL topology.

6. The method according to claim 5, characterized in that The step of determining the state of the corresponding first device as an unavailable state and determining the state of a channel associated with the corresponding first device as an unavailable state when the fault diagnosis result indicates that the corresponding first device has a fault includes: In the case where the fault diagnosis result indicates that the corresponding first device has a fault, determining the state of the corresponding first device as a temporarily unavailable state, determining the state of a channel associated with the corresponding first device as a temporarily unavailable state, and retrying the fault diagnosis for the corresponding first device; If the corresponding first device is still determined to have a fault after M fault diagnosis retries are performed on the corresponding first device, the state of the corresponding first device is determined to be a permanently unavailable state, and the state of the channel associated with the corresponding first device is determined to be a permanently unavailable state, where M is an integer greater than 0.

7. The method according to claim 1, characterized in that The scheduling priority of each channel is based on scheduling the NCCL tasks in the task queue, including: For each NCCL task in the task queue, based on the scheduling priority of each channel, select one or more channels for the corresponding NCCL task, and schedule one or more first devices associated with the one or more channels to execute the corresponding NCCL task; In the case where the target NCCL task fails to execute, scheduling the one or more first devices to re-execute the target NCCL task; If the target NCCL task still fails after re-executing N times, the target NCCL task is re-added to the task queue to reschedule the target NCCL task, and the state of each of the one or more channels is determined to be an unavailable state. A channel in an unavailable state cannot be scheduled, and N is an integer greater than 0.

8. A NCCL task scheduling method, characterized in that: Applicable to the second process in the NCCL process group, including: Monitor the operating status and / or performance of the first device to obtain first information, where the first information includes information related to the operating status and / or performance of the first device, and the first device at least runs a corresponding second process; The first information is periodically reported to the first process in the NCCL process group so that the first process schedules the NCCL tasks in the task queue; the first process is the main process of the NCCL process group, and the second process includes other processes in the NCCL process group except the first process.

9. The method according to claim 8, characterized in that The method further comprises: Receive the seventh information sent by the first process, the seventh information is used to indicate a second NCCL topology, the second NCCL topology is the topology of the communication network corresponding to the NCCL process group, the second NCCL topology is the topology obtained by the first process reconstructing the topology and channels of the communication network when the first parameter is greater than or equal to the first threshold, the first parameter represents the ratio of the number of channels in an unavailable state to the total number of channels in the first NCCL topology, and the second NCCL topology does not include the first device and channel in the unavailable state in the first NCCL topology.

10. An NCCL task scheduling device, characterized in that: include: A first processing unit is used to monitor one or more of the second information, the third information, and the fourth information of each channel in the first NCCL topology based on the first information periodically reported by each second process in the NCCL process group after being used by the first process in the NCCL process group, wherein the first information includes information related to the operating status and / or performance of the first device, the first device at least runs the corresponding second process, the second information represents whether the corresponding channel fails, the third information represents the performance of the corresponding channel, and the fourth information represents the load of the corresponding channel; the first NCCL topology is the topology of the communication network corresponding to the NCCL process group, the first process is the main process of the NCCL process group, and the second process includes other processes in the NCCL process group except the first process; a second processing unit, configured to determine or adjust a scheduling priority of a corresponding channel based on one or more of the second information, the third information, and the fourth information of each channel after being used by the first process; A scheduling unit is used to schedule NCCL tasks in the task queue based on the scheduling priority of each channel after being used by the first process.

11. An NCCL task scheduling device, characterized in that: include: A third processing unit is used to monitor the operation status and / or performance of the first device after being used by the second process in the NCCL process group, and obtain first information, wherein the first information includes information related to the operation status and / or performance of the first device, and the first device at least runs a corresponding second process; The second communication unit is used to periodically report the first information to the first process in the NCCL process group after being used by the second process, so that the first process can schedule NCCL tasks in the task queue; the first process is the main process of the NCCL process group, and the second process includes other processes in the NCCL process group except the first process.

12. A second device, characterized in that: At least a first process in an NCCL process group is running, and the first process is a main process of the NCCL process group; the second device includes: a first processor and a first memory for storing a computer program that can be run on the processor, Wherein, when the first processor is used to run the computer program, the steps of the method described in any one of claims 1 to 7 are executed.

13. A first device, characterized in that: At least a corresponding second process in the NCCL process group is running, the second process includes other processes in the NCCL process group except the first process, and the first process is the main process of the NCCL process group; the first device includes: a second processor and a second memory for storing a computer program that can be run on the processor, Wherein, when the second processor is used to run the computer program, it executes the steps of the method according to claim 8 or 9.

14. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 7, or implements the steps of the method according to claim 8 or 9.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 7, or implements the steps of the method according to claim 8 or 9.