Abnormality detection method and apparatus, and electronic device, distributed computing system and storage medium

By bypassing the CPU in the distributed computing system and directly detecting it on the switch link, the problem of positioning GPU exceptions in a large-scale distributed computing system is solved, and the abnormal unit is quickly isolated, which improves training efficiency and accuracy.

WO2025139723A1PCT designated stage expired Publication Date: 2025-07-03SHANGHAI XIYU JIZHI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/137643
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-31
Filing Date
2024-12-06
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In large-scale distributed computing systems, it is difficult to accurately locate abnormalities in non-CPU computing units such as GPUs and their associated switches and hardware devices, resulting in a drag on training efficiency.

Method used

By initiating a detection communication connection in the computing system, bypassing the CPU to perform correctness and transmission stability detection directly on the switch link, the analysis returns results to determine the abnormal calculation unit, including bandwidth and delay detection, disconnect or shield the abnormal unit.

Benefits of technology

Quickly and accurately locate and isolate the abnormal calculation unit, which improves the efficiency and stability of distributed training, reduces noise interference, and improves the accuracy and speed of neural network training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024137643_03072025_PF_FP_ABST
    Figure CN2024137643_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application is an abnormality detection method for a distributed computing system. The method comprises: initiating a detection communication connection to at least one computing unit of at least some computing nodes in a distributed computing system, wherein the detection communication connection is transmitted by means of a first switch communication link connected to the computing units subjected to detection, and does not pass through CPUs of the computing nodes where the computing units subjected to detection are located; receiving a return result of the detection communication connection from each of the computing units subjected to detection; and analyzing the return result by means of given detection, so as to determine an abnormal computing unit from among the computing units subjected to detection. The present application further relates to an abnormality detection apparatus for a distributed computing system, an electronic device, a distributed computing system, and a storage medium. The solution of the present application can effectively solve the problem of distributed training of a neural network being hindered due to the fact that it is difficult to locate an abnormality in computing units, which are non-central processing units, in a large-scale distributed computing system and an abnormality in switches and hardware devices in associated nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Abnormality detection method and device, electronic device, distributed computing system and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application number 2023118698534, filed with the Chinese Patent Office on December 31, 2023, entitled “Anomaly Detection Method and Device, Electronic Device, Distributed Computing System and Storage Medium,” the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of distributed computing, and in particular to an anomaly detection method and device for a distributed computing system, an electronic device, a distributed computing system, and a storage medium. Background Art

[0004] In the field of artificial intelligence, due to the urgent need for high-performance computing, non-CPU processing units such as GPUs are often used to optimize training and inference processes. These units are designed for parallel processing and can effectively handle complex calculations, significantly improving efficiency compared to CPUs.

[0005] With the development of artificial intelligence (AI), especially with the booming popularity of large language models (LLMs), neural network models are becoming increasingly large, driving the widespread adoption of distributed training methods. In such distributed training environments, a significant portion of computation, communication, and memory access are primarily performed by GPUs, rather than CPUs. This involves intensive communication between GPUs and between GPUs and other heterogeneous components.

[0006] Such distributed computing systems often include multiple computing nodes, each of which may contain multiple non-CPU computing units, such as GPUs, for performing distributed training. When a computing unit, such as a GPU, and its connected components, in such a distributed computing network experiences an anomaly, it is difficult to accurately locate such an abnormal computing unit. If the presence of an abnormal computing unit cannot be quickly identified, the anomalies in these computing units, such as failures, bottlenecks, and downgrades, will significantly reduce the training efficiency of the entire distributed computing system.

[0007] As distributed training is still in its infancy, the industry currently lacks understanding of this issue and lacks corresponding solutions.

[0008] The description of the background technology is intended to help understand the relevant technology in the relevant field, and does not mean that the background technology content is admitted to be prior art. Summary of the Invention

[0009] The embodiments of the present application aim to provide an anomaly detection method and device for a distributed computing system, an electronic device, a distributed computing system, and a storage medium, which can alleviate or solve at least one of the technical problems mentioned above.

[0010] Optionally, the present application provides an anomaly detection method for a distributed computing system. The distributed computing system includes multiple computing nodes, each computing node includes a CPU, a first switch, and one or more computing units, the computing units being non-CPU processing units, wherein the first switch is upstream connected to an upper-level switch or a CPU interface, and the first switch is downstream connected to multiple hardware devices, wherein the multiple hardware devices downstream connected to the first switch include at least one of the computing units.

[0011] Here, the anomaly detection method includes:

[0012] Initiate a detection communication connection to at least one of the computing units of at least some of the computing nodes, wherein the detection communication connection is transmitted via a first switch communication link to which the detected computing unit is connected and does not pass through a CPU of the computing node where the detected computing unit is located;

[0013] receiving a return result of the detection communication connection from each of the detected computing units;

[0014] Through a given detection, the returned result is analyzed to determine an abnormal computing unit among the detected computing units.

[0015] In some embodiments, the detection communication connection does not pass through an upper-level switch of the first switch to which the detected computing unit is connected.

[0016] In some embodiments, the given detection includes correctness detection and transmission stability detection.

[0017] In some embodiments, the transmission stability detection includes bandwidth detection and / or delay detection.

[0018] In some embodiments, the step of analyzing the returned result through a given detection to determine an abnormal computing unit in the detected computing unit includes:

[0019] performing a plurality of correctness checks, and determining a computing unit that fails one or more of the correctness checks, a first switch that fails one or more of the correctness checks, a computing unit connected to a communication link of the first switch, or a computing unit connected to the same first switch as a hardware device that fails one or more of the correctness checks as an abnormal computing unit;

[0020] During the bandwidth detection, determining the connection bandwidth of all the computing units detected in the detection communication, determining the quantile bandwidth based on a given first quantile threshold, and determining a computing unit with a bandwidth less than the quantile bandwidth or less than the threshold bandwidth as an abnormal computing unit, wherein the threshold bandwidth is a first given difference less than the quantile bandwidth; and / or

[0021] During the delay detection, the connection delay of all detected computing units in the detection communication is determined, and the quantile delay is determined based on a given second quantile threshold, and the computing units with a delay greater than the quantile delay or greater than the threshold delay are determined as abnormal computing units, and the threshold delay is slower than the quantile delay by a second given difference.

[0022] In some embodiments, the correctness detection includes memory allocation correctness detection, memory release correctness detection, virtual memory allocation correctness detection and / or virtual memory release correctness detection of the detected computing unit.

[0023] In some embodiments, the correctness test is selected from at least one of the following tests:

[0024] Basic sub-thread fixed buffer memory allocation detection,

[0025] Basic child thread fixed buffer virtual memory allocation detection,

[0026] Basic memory allocation detection,

[0027] Basic small buffer mapping detection,

[0028] Basic misaligned mapping detection,

[0029] Basic virtual memory allocation detection,

[0030] Basic testing with tokens,

[0031] Data verification memory allocation detection,

[0032] Data verification virtual memory allocation detection,

[0033] Access memory allocation failure detection after release,

[0034] Access-after-free virtual memory allocation failure detection,

[0035] Access memory allocation failure detection after GDR is closed,

[0036] Accessing virtual memory allocation failure detection after GDR is closed,

[0037] After the release, access the branch operation fork memory allocation failure detection,

[0038] Access fork virtual memory allocation failure detection after release,

[0039] Fork memory allocation failure detection after GDR mapping,

[0040] Fork virtual memory allocation failure detection after GDR mapping,

[0041] Child process GDR mapping parent process fork memory allocation failure detection,

[0042] Child process GDR mapping parent process fork virtual memory allocation failure detection,

[0043] The child process GDR fixed the parent process's token memory invalidation detection,

[0044] Memory allocation failure detection after fork mapping and release,

[0045] Detection of virtual memory allocation failure after fork mapping and release,

[0046] Double-mapped memory allocation failure detection,

[0047] Double-mapped virtual memory allocation failure detection,

[0048] Unix socket shared file descriptor GDR mapping memory allocation failure detection,

[0049] Unix socket shared file descriptor GDR mapping virtual memory allocation failure detection,

[0050] Unix socket shared file descriptor GDR fixed buffer memory allocation failure detection,

[0051] Unix socket shared file descriptor GDR pinned buffer virtual memory allocation failure detection.

[0052] In some embodiments, the anomaly detection method further includes:

[0053] The computing unit determined to be abnormal, the first switch connected to the computing unit determined to be abnormal, or other hardware devices connected to the same first switch as the computing unit determined to be abnormal are disconnected or shielded.

[0054] In some embodiments, the anomaly detection method includes: newly connecting multiple computing nodes, each newly connected computing node includes one or more computing units; initiating a detection communication connection to at least one of the computing units of at least a portion of the computing nodes includes: initiating the detection communication connection to the computing unit of the newly connected computing node; the anomaly detection method also includes: preventing the computing unit determined to be abnormal from accessing.

[0055] In some embodiments, the first switch and the upper-level switch are PCIe switches.

[0056] Optionally, the present application provides an anomaly detection device for a distributed computing system. The distributed computing system includes multiple computing nodes, each computing node includes a CPU, a first switch, and one or more computing units, the computing units being non-CPU processing units, wherein the first switch is upstream connected to an upper-level switch or a CPU interface, and the first switch is downstream connected to multiple hardware devices, wherein the multiple hardware devices downstream connected to the first switch include at least one of the computing units.

[0057] Here, the abnormality detection device includes:

[0058] a communication initiating unit configured to initiate a detection communication connection to at least one of the computing units of at least a portion of the computing nodes, wherein the detection communication connection is transmitted via a first switch communication link to which the detected computing unit is connected and does not pass through a CPU of the computing node where the detected computing unit is located;

[0059] a receiving unit configured to receive a return result of the detection communication connection from each of the detected computing units; and

[0060] The abnormality determination unit is configured to analyze the returned result through a given detection to determine an abnormal computing unit in the detected computing unit.

[0061] Optionally, the present application provides an electronic device, comprising a processor and a memory storing a computer program, wherein the processor is configured to implement the method described in any embodiment of the present application when running the computer program.

[0062] Optionally, the present application provides a distributed computing system, comprising multiple computing nodes, each computing node comprising a CPU, a first switch, and one or more computing units, wherein the computing units are non-CPU processing units, wherein the first switch is upstream connected to an upper-level switch or a CPU interface, and the first switch is downstream connected to multiple hardware devices, wherein the multiple hardware devices downstream connected to the first switch include at least one of the computing units;

[0063] The distributed computing system further includes one or more electronic devices according to any embodiment of the present application.

[0064] In some embodiments, one or more computing nodes of the plurality of computing nodes serve as the one or more electronic devices.

[0065] In some embodiments, the plurality of hardware devices connected to the first switch further include a network interface card (NIC), and the computing unit includes a GPU.

[0066] In some embodiments, at least some of the computing nodes include two levels of switches, the first level of switches includes the first switch, and the second level of switches includes a second switch to which the first switch is upstream connected, wherein the first switch has two downstream ports, and the two downstream ports of at least some of the first switches are respectively connected to the NIC and the GPU.

[0067] In some embodiments, at least some of the computing nodes include a first-level switch, wherein the first switch is upstream connected to the root complex RC, wherein the first switch has two downstream ports, and at least some of the two downstream ports of the first switch are respectively connected to the NIC and the GPU.

[0068] In a fifth aspect, a storage medium is provided, wherein the storage medium stores a computer program, and the computer program is configured to implement the method described in any embodiment of the present application when executed.

[0069] The anomaly detection method for a distributed computing system of the present application includes: initiating a detection communication connection to at least one computing unit of at least a part of the computing nodes of the distributed computing system, the detection communication connection is transmitted via a first switch communication link to which the detected computing unit is connected and does not pass through the CPU of the computing node where the detected computing unit is located; receiving a return result of the detection communication connection from each of the detected computing units; through a given detection, analyzing the return result to determine the abnormal computing unit among the detected computing units. This solution can effectively solve the current problem of difficulty in locating the anomalies of non-central processing unit computing units and switches and hardware devices in associated nodes in large-scale distributed computing systems, which hinders the distributed training of neural networks.

[0070] Further technical features and effects of this application will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use. It should be noted that the drawings described below only cover some embodiments of the present application. Ordinary technicians in this field can derive other drawings based on these drawings without having to engage in creative work. The purpose of these drawings is to better illustrate the technical details to facilitate understanding of the implementation methods of the present application, including:

[0072] FIG1 shows a topological structure diagram of an exemplary computing node;

[0073] FIG2 shows another topological structure diagram of an exemplary computing node;

[0074] FIG3 shows a schematic structural diagram of a distributed computing system and computing nodes thereof that can be used for distributed training;

[0075] FIG4 shows another schematic structural diagram of a distributed computing system and computing nodes thereof that can be used for distributed training;

[0076] FIG5 shows another schematic structural diagram of a distributed computing system and computing nodes thereof that can be used for distributed training;

[0077] FIG6 is a schematic flow chart of an abnormality detection method according to an embodiment of the present application;

[0078] FIG7 shows a schematic flow chart of bandwidth detection according to an embodiment of the present application;

[0079] FIG8 shows a schematic flowchart of delay detection according to an embodiment of the present application;

[0080] FIG9 shows an exemplary structural diagram of an abnormality detection device according to an embodiment of the present application;

[0081] FIG10 is a schematic diagram showing the structure of an electronic device capable of implementing the method according to an embodiment of the present application;

[0082] FIG11 shows a schematic structural diagram of a distributed computing system according to an embodiment of the present application. DETAILED DESCRIPTION

[0083] The following detailed description of exemplary embodiments of the present application is provided, and illustrations of the exemplary embodiments are shown in the accompanying drawings. When referring to the drawings, unless otherwise specified, identical numbers or symbols in different figures represent identical or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Instead, they are merely examples of some aspects of the apparatus and methods covered by one or more embodiments of this specification, as detailed in the claims of this application.

[0084] As used in this specification, the term "including" and its variations are intended to be broadly inclusive, meaning "including but not limited to" the listed items. Unless otherwise stated, the term "or" means "and / or," the term "based on" means relying on, or at least partially relying on, the terms "an example embodiment" and "an embodiment" refer to at least one example embodiment, and the term "another embodiment" refers to at least one different embodiment. The terms "first," "second," and so on may refer to different or the same items. Other explicit and implicit definitions may be included below.

[0085] Currently, numerous distributed computing systems have been proposed for various scenarios requiring high performance, large computational load, and / or large amounts of data. These include distributed computing scenarios such as neural network training. In these distributed computing systems, multiple computing nodes are connected via a network, and the computing nodes often use non-CPU computing units to perform part, or even most, or all of the computational tasks. Such computing units are, for example, GPUs. Accordingly, in distributed training environments, a significant portion of the computation and related communication and memory access activities are primarily performed by GPUs, rather than CPUs. This involves intensive communication between GPUs and between GPUs and other heterogeneous components.

[0086] With reference to Figures 1 and 2, different topological diagrams of exemplary nodes are shown. As shown in Figure 1, in a computing node, one communication mode between a GPU and other hardware (such as another GPU or an RDMA network interface card RNIC) is to bypass the CPU through a PCIe switch (PCIe Switch) (and then pass through the node host memory) and then transmit to another GPU. As shown in Figure 2, a second communication mode between a GPU and other hardware (such as another GPU or an RDMA NIC) is to bypass the CPU dependency through DMA through the PCIe Switch and transmit directly to the RNIC through the GDR (GPU Direct RDMA) method. As explained above, since the GPU is the main executor of a considerable part of the calculations and related communications and memory access activities in distributed training, the second communication mode is often more advantageous and more important in calculations such as distributed training.

[0087] Furthermore, due to the sheer number of parameters, distributed computing systems used for distributed training of large language models (LLMs), in particular, require a large number of computing nodes and their associated computing units, such as GPUs (e.g., 100 to 10,000), to participate in training. This inevitably requires synchronization and updating of training parameters (such as gradient parameters) and model parameters between different GPUs. The failure of a single GPU, or even simply poor performance such as a bottleneck or degradation, can still significantly impact distributed training. However, due to their large number, it is often difficult to accurately locate these abnormal GPUs or their associated switches or hardware.

[0088] The inventors thus realized the need to provide the ability to quickly locate abnormal non-CPU computing units, such as GPUs, in larger-scale distributed computing systems. However, traditional solutions often focus on fault detection, and traditional fault detection solutions are often executed by the host. Moreover, they are usually only indicator-driven or communication verification is performed at the link layer. These solutions are limited by hardware capabilities and indicator expression capabilities, and these solutions add too much noise and cannot accurately confirm abnormal computing units, such as GPUs, in larger-scale distributed computing systems. In addition, the complexity of PCIe failure scenarios further increases the difficulty of traditional fault detection solutions in solving such problems.

[0089] To this end, embodiments of the present application provide an anomaly detection method for a distributed computing system, which can be implemented by one or more computing nodes of the computing system. The method can be implemented on computing nodes where multiple non-detected computing units are located (referred to as "non-detected"). Optionally, anomaly detection can be performed multiple times, with a different non-detected computing node performing anomaly detection each time; or anomaly detection can be performed once or multiple times, with multiple non-detected computing nodes performing an anomaly detection in parallel.

[0090] In an embodiment of the present application, a distributed computing system may include multiple computing nodes, each computing node including a CPU, a first switch, and one or more computing units. The computing unit is a non-CPU processing unit (such as a GPU). The first switch is connected upstream to an upper-level switch or a CPU interface. The first switch is connected downstream to multiple hardware devices, and the multiple hardware devices connected downstream to the first switch include at least one computing unit. In some embodiments, the first switch and the upper-level switch are PCIe switches.

[0091] Figures 3 to 5 show several schematic structures of a distributed computing system and computing nodes thereof for distributed training, but the computing system to which the anomaly detection solution of the present application can be applied is not limited thereto.

[0092] With reference to Figures 3 to 5, a distributed computing system (not shown in Figures 3 to 5) may include multiple computing nodes 310, 410, 510, wherein the ellipsis in Figures 3 to 5 schematically represents multiple computing nodes. Each computing node includes a CPU (not shown), a first switch 302, 402, 502 and one or more computing units 301, 401, 501, 1101, which are non-CPU processing units. In the embodiments shown in Figures 3 and 4, each computing node has 8 computing units, such as GPUs G0-G7. In the embodiment shown in Figure 5, each computing node has 4 computing units, such as GPUs G0-G3. In some embodiments, the computing unit may be a GPU, but may also be other processing units that can be used for machine learning training or inference or other high-performance computing, such as an NPU.

[0093] As shown in Figures 3 and 4 , first switches 302 and 402 are connected upstream to the upper-level switch. More specifically, compute nodes 310 and 410 may include two levels of switches: the first-level switch includes the aforementioned first switches 302 and 402, and the second-level switch includes the second switches 304 and 404 connected upstream to the first switches. Here, the first switch is the bottom-level switch, and the second switch is the upper-level switch. The second switches 304 and 404 are connected upstream to the root complex (RC) 305 and 405, which serves as the interface between the CPU and the PCIe bus and is also represented by CPU sockets, such as Socket 0 and Socket 1. In the embodiment shown in Figure 3 , the second switch 304 is connected downstream to the first switch 302 and another compute unit, such as a GPU, for example, G1, via two downstream ports. In the embodiment shown in Figure 4 , both downstream ports of the second switch 404 are connected to the first switch. Accordingly, in the embodiment shown in Figure 3 , each compute node has four first switches 302; in the embodiment shown in Figure 4 , each compute node has eight first switches 402.

[0094] 5 , the first switch 502 is upstream connected to the root complex RC 505, which serves as the CPU and PCIe bus interface and is also presented in the form of CPU sockets, such as Socket 0 and Socket 1. Here, there is only a single layer of switches in the computing node, namely the first switch.

[0095] As shown in Figures 3 to 5, the multiple hardware devices connected downstream to the first switches 310, 410, and 510 include at least one computing unit 301, 401, and 501. Furthermore, the hardware devices connected downstream to the first switches may also include network interface cards (NICs) 303, 403, and 503, which in some embodiments may be RDMA network interface cards (RNICs). As shown in Figures 3 to 5, the first switches 310, 410, and 510 have two downstream ports. At least some of the first switches 310, 410, and 510, and optionally all of the first switches, have two downstream ports (downlinks 306, 406, and 506) connected to a computing unit, such as a GPU, and a NIC, such as an RNIC, respectively. Accordingly, in the embodiment shown in Figure 3, each computing node has four RNICs, namely, R0-R3; in the embodiment shown in Figure 4, each computing node has eight RNICs, namely, R0-R7; and in the embodiment shown in Figure 5, each computing node has four RNICs, namely, R0-R3.

[0096] The first switch and the possible second switch in the distributed computing system shown in Figures 3 to 5 are switches within a node, not switch devices used to connect two nodes. Optionally, the first switch and the possible second switch are PCIe switches. Accordingly, the first switch and the possible second switch in the embodiments of the present application can have any suitable shape and structure, for example, they can be in the form of a PCIe chip and corresponding slots, and this application does not limit this.

[0097] Next, the anomaly detection method according to the embodiment of the present application is described. In this article, "abnormal" will be interpreted broadly, including but not limited to failures, as well as bottlenecks, degradations and poor performance that are not failures (at this time, the relevant components can still work). Further, in some embodiments of the present application, the detection scheme explicitly includes performing bottleneck, degradation and poor performance detection on the computing node and the related first switch and hardware and related communication links that are not failures or errors (at this time, the relevant components can still work), which is particularly beneficial to the distributed training of the neural network model.

[0098] FIG6 shows an anomaly detection method according to an embodiment of the present application, which may include:

[0099] S610: Initiate a detection communication connection to at least one computing unit of at least a portion of the computing nodes.

[0100] In a further embodiment of the present application, the detection communication connection is transmitted via a first switch communication link to which the detected computing unit is connected and does not pass through a CPU of a computing node where the detected computing unit is located.

[0101] In a further embodiment, the detection communication connection does not pass through the upper-level switch of the first switch to which the detected computing unit is connected. In a further embodiment, the downlink port (downlink) of the first switch is respectively connected to computing units, such as a GPU and a NIC, and the detection communication connection is established so that the detection communication connection is established from the NIC, such as RNIC, connected to the first switch corresponding to the detected computing unit (GPU), through a downlink of the first switch corresponding to the detected computing unit, through the first switch corresponding to the detected computing unit, and through another downlink of the first switch corresponding to the detected computing unit to reach the detected computing unit. Taking the embodiments shown in Figures 3 to 5 as an example, if GPU G0 is the detected computing unit, the detection communication connection passes through NICs 303, 403, 503, such as R0; downlinks 306, 406, 506 on the left side of the figure; first switches 302, 402, 502 on the left side of the figure, such as a PCIe switch; downlinks 306, 406, 506 on the right side of the first switch; and reaches the detected computing unit 301, 401, 502, such as G0. As mentioned above, anomaly detection is performed by one or more computing nodes in the computing system where the computing unit not being detected is located, and more specifically, a detection communication connection is initiated. In some embodiments of the present application, the detection may be performed only on the PCIe link of the computing unit being detected.

[0102] In some embodiments of the present application, the gdr_copy tool may be used to verify the correctness, performance, and GDR capabilities of a PCIe link.

[0103] S620: Receive a return result of detecting the communication connection from each of the detected computing units.

[0104] S630: Through a given detection, analyzing the returned results to determine abnormal computing units among the detected computing units.

[0105] In the embodiment of the present application, the given detection includes correctness detection and transmission stability detection.

[0106] In some embodiments of the present application, the anomaly detection is implemented at the application layer and can be implemented through other computing nodes in the computing system, which can help eliminate noise and quickly locate abnormal computing units.

[0107] In an embodiment of the present application, performing the correctness check for analysis in step S630 may include performing multiple correctness checks and determining as an abnormal computing unit a computing unit that fails one or more of the correctness checks, a first switch that fails one or more of the correctness checks, a computing unit connected to a communication link of the first switch, or a computing unit connected to the same first switch as a hardware device that fails one or more of the correctness checks.

[0108] According to a further embodiment, the correctness check includes a memory allocation correctness check, a memory release correctness check, a virtual memory allocation correctness check and / or a virtual memory release correctness check of the checked computing unit.

[0109] In some embodiments of the present application, when performing correctness detection for analysis in step S630, correctness verification can be performed according to the scenario, and if an error occurs, the hardware behavior can be clearly located.

[0110] Optionally, step S630 includes: determining different assigned tasks for the tested computing unit; performing a correctness check on the tested computing unit; and determining whether the tested computing unit is abnormal based on the assigned tasks and the correctness check results. For example, multiple different computing units may be assigned different tasks. If different related correctness check issues are detected, the computing unit, such as a GPU, involved in the correctness issue can be quickly determined based on the association between the tasks and the correctness checks.

[0111] In some embodiments, the correctness test is selected from at least one of the following tests, and optionally, includes all of the following tests:

[0112] Basic sub-thread fixed buffer memory allocation detection,

[0113] Basic child thread fixed buffer virtual memory allocation detection,

[0114] Basic memory allocation detection,

[0115] Basic small buffer mapping detection,

[0116] Basic misaligned mapping detection,

[0117] Basic virtual memory allocation detection,

[0118] Basic testing with tokens,

[0119] Data verification memory allocation detection,

[0120] Data verification virtual memory allocation detection,

[0121] Access memory allocation failure detection after release,

[0122] Access-after-free virtual memory allocation failure detection,

[0123] Access memory allocation failure detection after GDR is closed,

[0124] Accessing virtual memory allocation failure detection after GDR is closed,

[0125] After the release, access the branch operation fork memory allocation failure detection,

[0126] Access fork virtual memory allocation failure detection after release,

[0127] Fork memory allocation failure detection after GDR mapping,

[0128] Fork virtual memory allocation failure detection after GDR mapping,

[0129] Child process GDR mapping parent process fork memory allocation failure detection,

[0130] Child process GDR mapping parent process fork virtual memory allocation failure detection,

[0131] The child process GDR fixed the parent process's token memory invalidation detection,

[0132] Memory allocation failure detection after fork mapping and release,

[0133] Detection of virtual memory allocation failure after fork mapping and release,

[0134] Double-mapped memory allocation failure detection,

[0135] Double-mapped virtual memory allocation failure detection,

[0136] Unix socket shared file descriptor GDR mapping memory allocation failure detection,

[0137] Unix socket shared file descriptor GDR mapping virtual memory allocation failure detection,

[0138] Unix socket shared file descriptor GDR fixed buffer memory allocation failure detection,

[0139] Unix socket shared file descriptor GDR pinned buffer virtual memory allocation failure detection.

[0140] In a further embodiment, the aforementioned memory may be CUDA memory.

[0141] In another embodiment of the present application, the transmission stability detection includes bandwidth detection, delay detection, or both.

[0142] In an embodiment of the present application, when performing bandwidth detection for analysis in step S630, it may include: determining the connection bandwidth of all computing units detected in the detection communication, determining it as the quantile bandwidth based on a given first quantile threshold, and determining a computing unit that is less than the quantile bandwidth or less than the threshold bandwidth as an abnormal computing unit, and the threshold bandwidth is a first given difference smaller than the quantile bandwidth.

[0143] Referring to FIG. 7 , when performing bandwidth testing, the bandwidth of all tested computing units can be determined in step S710 . In some specific embodiments, multiple different types of bandwidth can be obtained, such as read bandwidth and write bandwidth; and / or bandwidth can be obtained multiple times, for example, to determine the bandwidth corresponding to repeated writes of a given size, such as files of the same size, increasing size, or decreasing size. Accordingly, the return result obtained in step S620 can contain relevant information. For example, the cuMemAlloc function can be used to allocate memory in the GPU's global memory. Through device pointers and mapped addresses, the electronic device initiating the testing communication connection can directly access the memory of the tested GPUs in other computing nodes, primarily focusing on measuring the bandwidth of data writes and reads. The results show the efficiency of GPU memory operations under different test conditions. Continuing with FIG. 7 , the quantile bandwidth can be determined in step S720 based on the bandwidth of all tested computing units. In one exemplary embodiment, the 90th percentile bandwidth can be selected, but other quantiles are also contemplated.

[0144] Next, in step S730, abnormal computing nodes, such as abnormal GPUs, bottlenecks, or degraded GPUs, can be identified in the detected computing nodes, such as GPUs, based on the quantile bandwidth. In some embodiments, an abnormality, such as a bottleneck or degradation, can be determined by setting a threshold bandwidth that is lower than the quantile bandwidth by a first given difference. The first given difference can be a percentage threshold, such as 10% to 30%. In other embodiments, the quantile bandwidth can be directly used as the threshold bandwidth, in which case the quantile can be selected to be relatively low, such as 10-30 percentiles.

[0145] In an embodiment of the present application, when delay detection is performed in step S630 for analysis, it may include: determining the connection delay of all detected computing units in the detection communication, determining it as a quantile delay based on a given second quantile threshold, and determining a computing unit that is greater than the quantile delay or greater than the threshold delay as an abnormal computing unit, and the threshold delay is slower than the quantile delay by a second given difference.

[0146] Referring to FIG8 , when performing latency detection, the latency of all detected computing units may be determined in step S810 . In some embodiments, the latency may be obtained multiple times. Accordingly, the returned result obtained in step S620 may include relevant information. The returned result may include a benchmark that measures the performance of the GPU Direct RDMA (GDR) memory copy function, optionally with iterative tests for different data sizes.

[0147] 8 , the quantile bandwidth may be determined in step S820 based on the delays of all detected computing units. In one exemplary embodiment, the 90th percentile delay may be selected, but other quantiles are also contemplated.

[0148] Next, in step S830, abnormal computing nodes, such as abnormal GPUs, bottlenecks, or degraded GPUs, can be identified in the detected computing nodes, such as GPUs, based on the quantile delay. In some embodiments, an abnormality, such as a bottleneck or degradation, can be determined by setting a threshold delay that is lower than the quantile delay by a second given difference. The second given difference can be a percentage threshold, such as 10% to 30%. In other embodiments, the quantile delay can be directly used as the threshold delay, in which case the quantile can be selected to be relatively low, such as 10-30 percentiles.

[0149] In the embodiment of the present application, the anomaly detection scheme can be implemented in the deployed computing nodes of the computing system, for example, before executing the computing task. Accordingly, when an abnormal computing unit is detected, the abnormal computing unit can be disconnected or shielded from the computing system.

[0150] In an embodiment of the present application, the above-mentioned abnormality detection method also includes: disconnecting or shielding the computing unit determined to be abnormal, the first switch connected to the computing unit determined to be abnormal, or other hardware devices connected to the same first switch as the computing unit determined to be abnormal.

[0151] In another embodiment, the anomaly detection scheme can also be applied when a new computing node is connected to the computing system. In this embodiment, the anomaly detection method includes: connecting multiple computing nodes, each of which includes one or more computing units; accordingly, step S610 may include: initiating a detection communication connection to the computing units of the newly connected computing nodes; optionally, the method further includes: preventing the access of computing units determined to be abnormal.

[0152] By way of explanation and not limitation, in the distributed training of large language models (LLMs), efficient parameter synchronization and updates across GPUs are important to maintain model consistency and accuracy. Each GPU in the network may be processing different parts of the model or different batches of data, but they all need to share insights to update the model uniformly. This process is highly dependent on the speed and efficiency of the communication network, which typically involves PCIe switches and RDMA devices. Therefore, even if the PCIe switches or RDMA devices are not faulty, any performance degradation or bottleneck or degradation can cause significant problems in distributed training, including parameter synchronization delays. Specifically, for LLMs, all nodes in the network may need to synchronize billions of parameters regularly. Even minor delays or bandwidth limitations can result in training with outdated parameters, leading to inconsistencies and potentially reducing the overall accuracy of the model or increasing training time. In this regard, the embodiments of the present application provide an anomaly detection scheme for distributed computing systems that can specifically address the current problem of difficulty in locating anomalies in non-central processing unit computing units and associated switches and hardware devices within nodes in large-scale distributed computing systems, which hinders the distributed training of neural networks. For example, on the one hand, correctness verification can be performed based on the scenario, and if an error occurs, the hardware behavior can be clearly located. On the other hand, stability verification can be performed based on performance, and if an error / performance degradation occurs, the PCIe fault / anomaly can be located. In addition, although this document describes the anomaly detection method of the embodiments of the present application as being used in a computing system for distributed training of neural networks, it is conceivable that the anomaly detection scheme of the embodiments of the present application can be used in other application scenarios, such as other distributed computing systems with high performance, large computing volume, and / or large data volume computing.

[0153] In another embodiment of the present application, an anomaly detection device for a distributed computing system may be provided, which may be implemented, for example, at the software level. The anomaly detection device may be used to detect distributed computing systems described elsewhere herein, such as the distributed computing systems described above with reference to the anomaly detection method, such as those shown in Figures 3 to 5 , which are not described in detail here.

[0154] Figure 9 shows an anomaly detection device 900 according to an embodiment of the present application. The anomaly detection device 900 may include a communication initiating unit 910, a receiving unit 920, and an anomaly determination unit 930. In this embodiment, the communication initiating unit 910 may be configured to initiate a detection communication connection to at least one of the computing units of at least a portion of the computing nodes. Here, the detection communication connection is transmitted via the first switch communication link to which the detected computing unit is connected and does not pass through the CPU of the computing node where the detected computing unit is located. The receiving unit 920 may be configured to receive a return result of the detection communication connection from each of the detected computing units. The anomaly determination unit 930 may be configured to analyze the return result through a given detection to determine an abnormal computing unit among the detected computing units.

[0155] In a further embodiment of the present application, the abnormality determination unit 930 may include a correctness detection subunit, a bandwidth detection subunit and / or a delay detection subunit, and these subunits may be correspondingly configured to perform correctness detection, bandwidth detection and / or delay detection as described above, which will not be repeated here.

[0156] The method and its steps, sub-steps and features in the embodiments of the present application can be combined with the device of the embodiments of the present application in a non-contradictory manner, and the device and its components, modules, units and features in the embodiments of the present application can be combined with the method of the embodiments of the present application in a non-contradictory manner.

[0157] In some embodiments of the present application, an electronic device is provided, which includes a processor and a memory storing a computer program, and the processor is configured to implement the method of any embodiment of the present application when running the computer program.

[0158] FIG10 shows a schematic diagram of an electronic device 1000 that can be used to implement the method or realize the embodiment of the present application. In some embodiments, the number of electronic devices may be more or less than the number shown. In some embodiments, a single or multiple electronic devices may be used for implementation. In some embodiments, cloud-based or distributed electronic devices may also be used for implementation.

[0159] As shown in Figure 10, the electronic device 1000 includes a processor 1010 and a memory 1020. The processor is used to execute programs stored in the memory, which can implement the methods, steps or functions described in the above embodiments when executed by a computer. The processor 1010 may include various types of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. The processor 1010 and the memory 1020 are interconnected via a bus 1030. An input / output (I / O) interface and the like can also be connected to the bus 1030.

[0160] In a further embodiment of the present application, a distributed computing system incorporating the above electronic device and / or anomaly detection device may also be included.

[0161] 11 and in combination with FIG3 to FIG5 , a distributed computing system 1100 (not shown in FIG3 to FIG5 ) may include multiple computing nodes 310, 410, 510, 1110, each of which includes a CPU (not shown), a first switch 302, 402, 502, 1102, and one or more computing units 301, 401, 501, 1101. In the illustrated embodiment, the computing unit is a non-CPU processing unit, and optionally, may be a GPU, such as G0...G7 shown in FIG3 to FIG5 and FIG11 , but may also be other processing units that can be used for machine learning training or inference or other high-performance computing, such as an NPU.

[0162] As shown in Figures 3 and 4, first switches 302 and 402 are connected upstream to the upper-level switch. More specifically, compute nodes 310 and 410 may include two levels of switches: the first-level switch includes the aforementioned first switch 302 and 402, and the second-level switch includes a second switch 304 and 404 connected upstream to the first switch. Here, the first switch is a bottom-level switch, and the second switch is an upper-level switch. The second switch 304 and 404 are connected upstream to the root complex (RC), which serves as the interface between the CPU and the PCIe bus and is also represented by CPU sockets, such as Socket 0 and Socket 1. In the embodiment shown in Figure 3, the second switch 304 is connected downstream to the first switch 302 and another compute unit, such as a GPU, for example, G1, via two downstream ports. In the embodiment shown in Figure 4, both downstream ports of the second switch 404 are connected to the first switch.

[0163] 5 , the first switch 502 is upstream connected to the root complex RC, which serves as the CPU and PCIe bus interface, and is also presented in the form of CPU sockets, such as Socket 0 and Socket 1. Here, there is only a single layer of switches in the computing node, namely the first switch.

[0164] As shown in Figures 3 to 5, the first switches 310, 410, and 510 are connected downstream to multiple hardware devices, including at least one computing unit 301, 401, and 501. In addition, the hardware devices connected downstream to the first switches may also include network interface cards (NICs), optionally RDMA network interface cards (RNICs), such as R0, R1, R2, and R3 shown in Figures 3 to 5 and 11. As shown in Figures 3 to 5, the first switches 310, 410, 510, and 1110 have two downstream ports. At least some of the first switches 310, 410, 510, and 1110, and optionally all of the first switches, have two downstream ports connected to a NIC, such as an RNIC, such as R0, and a GPU, such as G0.

[0165] Referring to Figure 11, the distributed computing system 1100 also includes one or more electronic devices 1120 according to an embodiment of the present application, which is, for example, the electronic device 1000 shown in Figure 10. The electronic device 1120 can initiate anomaly detection and log in to the detected computing node, more specifically, through the network, corresponding hardware devices, such as NIC and corresponding first switch, such as PCIe switch, to initiate a cross-machine detection communication connection to the detected computing unit, such as GPU. In some embodiments, one or more computing nodes in the distributed computing system will serve as one or more electronic devices. By using one or more computing nodes of the computing system as electronic devices for implementing the anomaly detection method of the embodiment of the present application, an additional benefit can be provided, namely, anomaly detection is initiated by "other" computing nodes within the computing system to detect the detected computing node, better simulating the calculation of the computing system, such as data transmission and synchronization during the distributed training process of a large-scale neural network. In some embodiments, when anomaly detection is implemented by multiple electronic devices, it can be detected simultaneously or separately. In some embodiments, when performing anomaly detection, all computing nodes in the computing system except the computing nodes used as electronic devices are detected.

[0166] Although not shown in the figures, an embodiment of the present application further provides a storage medium storing computer programs configured to execute a method related to any embodiment when run.

[0167] It should be understood that all operations in the method described above are merely exemplary, and the present disclosure is not limited to any operation in the method or the order of these operations, but should cover all other equivalent transformations under the same or similar concept.

[0168] It should also be understood that all modules in the above-described device can be implemented in various ways. These modules can be implemented as hardware, software, or a combination thereof. In addition, any module in these modules can be further divided into submodules or combined together in function.

[0169] Processor has been described in conjunction with various devices and methods.These processors can be implemented using electronic hardware, computer software or its arbitrary combination.Whether these processors are implemented as hardware or software will depend on specific application and the overall design constraint imposed on the system.As an example, the processor provided in this disclosure, any part of the processor or any combination of processors can be implemented as microprocessor, microcontroller, digital signal processor (DSP), field programmable gate array (FPGA), programmable logic device (PLD), state machine, gate logic, discrete hardware circuit and other suitable processing components configured for performing the various functions described in this disclosure.The function of the processor provided in this disclosure, any part of the processor or any combination of processors can be implemented as software performed by microprocessor, microcontroller, DSP or other suitable platform.

[0170] Software should be broadly considered to mean instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, running threads, processes, functions, etc. Software can reside in a computer-readable medium. A computer-readable medium can include, for example, a memory, which can be, for example, a magnetic storage device (e.g., a hard disk, a floppy disk, a magnetic stripe), an optical disk, a smart card, a flash memory device, a random access memory (RAM), a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a register, or a removable disk. Although the memory is shown as being separate from the processor in various aspects provided in the present disclosure, the memory can also be located inside the processor (e.g., a cache or register).

[0171] The above description is provided to enable any person skilled in the art to implement the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents of the elements of the various aspects described in this disclosure that are known or to be known to those skilled in the art are expressly incorporated herein by reference and are intended to be covered by the claims. Industrial Applicability

[0172] The above solution can effectively solve the current problem of difficulty in locating abnormalities in non-central processing unit computing units and associated switches and hardware devices within nodes in large-scale distributed computing systems, which hinders the distributed training of neural networks.

Claims

1. An anomaly detection method for a distributed computing system, characterized in that, The distributed computing system includes a plurality of computing nodes, each computing node including a CPU, a first switch, and one or more computing units, where the computing units are processing units other than the CPU. The first switch is connected upstream to a higher-level switch or a CPU interface, and the first switch is connected downstream to a plurality of hardware devices, and the plurality of hardware devices connected downstream by the first switch includes at least one of the computing units; The anomaly detection method includes: Initiating a detection communication connection to at least one of the computing units of at least some of the computing nodes, where the detection communication connection is transmitted via the communication link of the first switch to which the computing unit under detection is connected and does not pass through the CPU of the computing node where the computing unit under detection is located; Receiving the return results of the detection communication connection from each of the computing units under detection; and Analyzing the return results through a given detection to determine the anomalous computing units among the computing units under detection.

2. The anomaly detection method according to claim 1, wherein The detection communication connection does not pass through the higher-level switch of the first switch to which the computing unit under detection is connected.

3. The anomaly detection method according to claim 1 or 2, characterized in that The given detection includes correctness detection and transmission stability detection.

4. The anomaly detection method according to claim 3, wherein, The transmission stability detection includes bandwidth detection and / or latency detection.

5. The anomaly detection method according to claim 4, wherein The analyzing the return results through a given detection to determine the anomalous computing units among the computing units under detection includes: Through multiple correctness detections, determining as anomalous computing units those computing units that fail one or more of the correctness detections, the first switches that fail one or more of the correctness detections, the computing units connected to the communication links of the first switches, or the computing units connected to the same first switch as the hardware devices that fail one or more of the correctness detections; Through the bandwidth detection, determining the connection bandwidths of all the computing units under detection in the detection communication, determining the quantile bandwidth based on a given first quantile threshold, and determining as anomalous computing units those computing units with a bandwidth less than the quantile bandwidth or less than a threshold bandwidth, where the threshold bandwidth is smaller than the quantile bandwidth by a first given difference; and / or Through the latency detection, determining the connection latencies of all the computing units under detection in the detection communication, determining the quantile latency based on a given second quantile threshold, and determining as anomalous computing units those computing units with a latency greater than the quantile latency or greater than a threshold latency, where the threshold latency is slower than the quantile latency by a second given difference.

6. The anomaly detection method according to any one of claims 3-5, characterized in that, The correctness detection includes memory allocation correctness detection, memory release correctness detection, virtual memory allocation correctness detection, and / or virtual memory release correctness detection of the computing units under detection.

7. The anomaly detection method according to claim 6, wherein The correctness detection is selected from at least one of the following detections: Basic sub-thread fixed buffer memory allocation detection, Basic sub-thread fixed buffer virtual memory allocation detection, Basic memory allocation detection, Basic small buffer mapping detection, Basic misaligned mapping detection, Basic virtual memory allocation detection, Basic token test detection, Data verification memory allocation detection, Data verification virtual memory allocation detection, Memory allocation invalidation detection after release Virtual memory allocation invalidation detection after release Memory allocation invalidation detection after GDR is closed Virtual memory allocation invalidation detection after GDR is closed Memory allocation invalidation detection for the fork operation of the branch operation after release Virtual memory allocation invalidation detection for the fork operation after release Memory allocation invalidation detection for fork after GDR mapping Virtual memory allocation invalidation detection for fork after GDR mapping Memory allocation invalidation detection for the fork of the parent process mapped by the child process's GDR Virtual memory allocation invalidation detection for the fork of the parent process mapped by the child process's GDR Memory invalidation detection for the token-bearing memory of the parent process fixed by the child process's GDR Memory allocation invalidation detection after fork mapping and release Virtual memory allocation invalidation detection after fork mapping and release Double mapping memory allocation invalidation detection Double mapping virtual memory allocation invalidation detection Memory allocation invalidation detection for the GDR mapping of the shared file descriptor of the Unix socket Virtual memory allocation invalidation detection for the GDR mapping of the shared file descriptor of the Unix socket Memory allocation invalidation detection for the GDR-fixed buffer of the shared file descriptor of the Unix socket Virtual memory allocation invalidation detection for the GDR-fixed buffer of the shared file descriptor of the Unix socket.

8. The anomaly detection method according to any one of claims 1-7, characterized in that, The abnormal detection method further includes: Disconnecting or shielding the computing unit determined to be abnormal, the first switch connected to the computing unit determined to be abnormal, or other hardware devices connected to the same first switch as the computing unit determined to be abnormal.

9. The anomaly detection method according to any one of claims 1-8, characterized in that, The method includes: newly accessing a plurality of the computing nodes, each newly accessed computing node including one or more of the computing units; The step of initiating a detection communication connection to at least one of the computing units of at least some of the computing nodes includes: initiating the detection communication connection to the computing units of the newly accessed computing nodes; The method further includes: preventing the computing unit determined to be abnormal from accessing.

10. The anomaly detection method according to any one of claims 1-9, characterized in that, The first switch and the upper-level switch are PCIe switches (PCIeSwitch).

11. An anomaly detection device for a distributed computing system, characterized in that, The distributed computing system includes a plurality of computing nodes, each computing node including a CPU, a first switch, and one or more computing units, where the computing unit is a non-CPU processing unit, and where the first switch is connected upstream to an upper-level switch or a CPU interface, and the first switch is connected downstream to a plurality of hardware devices, and the plurality of hardware devices connected downstream by the first switch includes at least one of the computing units; The abnormal detection device includes: A communication initiation unit configured to initiate a detection communication connection to at least one of the computing units of at least some of the computing nodes, where the detection communication connection is transmitted via a communication link of the first switch connected to the computing unit to be detected and does not pass through the CPU of the computing node where the computing unit to be detected is located; A receiving unit configured to receive the return result of the detection communication connection from each of the computing units to be detected; and Anomaly determination unit, configured to analyze the return result through a given detection to determine an abnormal computing unit in the detected computing unit.

12. An electronic device, characterized in that, Comprising a processor and a memory storing a computer program, the processor being configured to implement the method according to any one of claims 1 to 10 when running the computer program.

13. A distributed computing system, characterized in that, Comprising a plurality of computing nodes, each computing node comprising a CPU, a first switch, and one or more computing units, the computing units being non-CPU processing units, wherein the first switch is connected upstream to a higher-level switch or a CPU interface, and the first switch is connected downstream to a plurality of hardware devices, and the plurality of hardware devices connected downstream by the first switch includes at least one of the computing units; The distributed computing system further includes one or more electronic devices according to claim 12.

14. The distributed computing system according to claim 13, wherein One or more of the plurality of computing nodes serve as the one or more electronic devices.

15. The distributed computing system according to claim 13 or 14, characterized in that, The plurality of hardware devices connected to the first switch further includes a network interface card NIC, and the computing unit includes a GPU.

16. The distributed computing system according to claim 15, wherein At least some of the computing nodes include a two-level switch, the first-level switch including the first switch, and the second-level switch including a second switch to which the first switch is connected upstream, wherein the first switch has two downstream ports, and two downstream ports of at least some of the first switches are respectively connected to the NIC and the GPU.

17. The distributed computing system according to claim 15 or 16, characterized in that, At least some of the computing nodes include a one-level switch, the first switch being connected upstream to a root complex RC, wherein the first switch has two downstream ports, and two downstream ports of at least some of the first switches are respectively connected to the NIC and the GPU.

18. A storage medium, characterized in that, The storage medium stores a computer program, the computer program being configured to implement the method according to any one of claims 1 to 10 when being run.

Citation Information

Patent Citations

  • Heterogeneous thousand-kernel high-throughput processing system based on CPU and GPU and modification method of system

    CN107122162A

  • Task processing method and device based on distributed system, equipment and medium

    CN115686831A

  • Abnormality detection method and device, electronic equipment, distributed computing system and storage medium

    CN118012680A

  • Reconfigurble CPU / GPU interconnect to mitigate power / thermal throttling

    US20200166977A1