Fault positioning method and device in set communication, storage medium and equipment
By analyzing the communication logs of the collection communication members and determining the strategy based on the fault type, the problem of difficulty in collective communication failure positioning in distributed large-scale model training is solved, fast and effective fault positioning is achieved, and training efficiency is improved.
Patent Information
- Application Number
- CN202510125491.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-13
AI Technical Summary
In distributed large-scale model training, fault location is difficult during the collective communication process, and the existing technology takes time to affect training efficiency.
By obtaining the communication log of each communication member, determining the fault location strategy based on the fault type of the alarm event, and analyzing the communication log to locate the root cause of the fault.
It realizes rapid positioning of the causes of failures in collective communication, reduces the fault location time, and improves the efficiency of large-scale model training.
Smart Images

Figure CN119996173A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method, device, storage medium and equipment for locating a fault in collective communication. Background Art
[0002] The rapid development of large language models has driven the progress of machine learning and artificial intelligence. For example, GPT-3 and Llama are close to human level in text generation, question answering, and code writing. However, the process of training these models is very complex and resource-intensive. For example, training Llama 3.1405B requires a cluster of 16,384 NVIDIA H100 80GB GPUs. Large-scale distributed computing makes collective communication the key to connecting multiple nodes.
[0003] Distributed large model training involves a wide range of knowledge and complex link structures. Once any link has a problem, the entire training task may fail. Among them, the failures that occur during collective communication are particularly prominent. Such failures are often difficult to locate, and the causes of these failures are varied. Therefore, how to locate failures in collective communication is an urgent problem to be solved. Summary of the invention
[0004] The present specification provides a method, apparatus, storage medium and device for locating faults in collective communications to at least partially solve the above-mentioned problems existing in the prior art.
[0005] This manual adopts the following technical solutions:
[0006] This specification provides a method for locating a fault in collective communication, including:
[0007] In response to an alarm event of collective communication, a communication log of each communication member in the collective communication is obtained; the communication log includes a collective communication log corresponding to the collective communication and a point-to-point communication log corresponding to each point-to-point communication included in the collective communication;
[0008] Determine a fault location strategy according to the fault type in the alarm event;
[0009] The determined fault location strategy is used to analyze the collective communication log and the point-to-point communication log contained in the communication log of each communication member, so as to locate the root cause of the alarm event in each communication member.
[0010] This specification provides a fault location device in collective communication, including:
[0011] A log acquisition module, configured to obtain, in response to an alarm event of collective communication, a communication log of each communication member in the collective communication; the communication log includes a collective communication log corresponding to the collective communication and a point-to-point communication log corresponding to each point-to-point communication included in the collective communication;
[0012] A strategy selection module, used to determine a fault location strategy according to the fault type in the alarm event;
[0013] The analysis module is used to analyze the collective communication log and the point-to-point communication log contained in the communication log of each communication member by adopting the determined fault location strategy, so as to locate the root cause of the alarm event in each communication member.
[0014] This specification provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the fault location method in the collective communication described above is implemented.
[0015] The present specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned fault location method in collective communication when executing the program.
[0016] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0017] In the fault location method in collective communication provided in this specification, each communication member in the collective communication records its own communication log, and the communication log contains the collective communication log corresponding to the collective communication and the point-to-point communication log corresponding to each point-to-point communication contained in the collective communication. When the host receives an alarm event of the collective communication, the communication log of each communication member is obtained, and then the fault location strategy is determined according to the fault type indicated in the alarm event. The determined fault location strategy is used to analyze the collective communication log and the point-to-point communication log in the communication log of each communication member, so as to locate the root cause of the alarm event in each communication member. The above method can effectively and quickly locate the cause of the fault in the collective communication process. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification.
[0019] In the figure:
[0020] Figure 1 A schematic diagram of the collective communication principle provided in the embodiments of this specification;
[0021] Figure 2 A schematic diagram of a flow chart of fault location in collective communication provided in this specification;
[0022] Figure 3 A schematic diagram of collective communication implemented using the ring algorithm provided in this specification;
[0023] Figure 4 A schematic diagram of collective communication implemented using a tree structure algorithm provided in this specification;
[0024] Figure 5 A schematic diagram of a structure of a fault location device in collective communication provided in this specification;
[0025] Figure 6 A method corresponding to the Figure 2 Schematic diagram of electronic equipment. DETAILED DESCRIPTION
[0026] In the collective communication process, the host (i.e., host) generally sends the kernel function required for this collective communication to the communication members (i.e., ranks) participating in the collective communication, and each communication member calls the kernel function to perform the collective communication operation. In the process of training a large model, the host is generally a CPU, and the communication members are generally multiple GPUs or multiple embedded neural network processors (NPUs) or parallel processing units (PPUs). A GPU or an NGU or a PPU is used as a communication member. The following text only takes the GPU as an example. The kernel function includes a variety of functions, such as broadcast, reduce, all-reduce, scatter, and gather.
[0027] Figure 1 The schematic diagram of the collective communication principle provided in the embodiment of this specification is that when the GPU calls and executes the kernel function, the data required for the collective communication can be processed concurrently through multiple channels (i.e., channels). Since these GPUs may be located on the same server or on different servers, when the GPUs located on the same server perform collective communication on the same channel, they can use the nvlink intra-machine transmission method for communication, such as Figure 1 When GPUs located on different servers communicate collectively on the same channel, they can use the qp (queue pair) transmission of the RDMA network between machines, such as Figure 1 r0 and r1 in.
[0028] The following uses the kernel function as an aggregation (all-reduce) function as an example to illustrate the role of collective communication.
[0029] Assume that the communication members participating in the collective communication include rank0, rank1, and rank2, and these three communication members have arrays (d 00 , d 01 )、(d 10 , d 11 )、(d 20 , d 21 ), then when the kernel function is an all-reduce function, after the three communication members call the all-reduce function and execute it respectively, the arrays owned by the three communication members will become (d 00 +d 10 +d 20 , d 01 +d 11 +d 21 ). This completes a collective communication with the all-reduce function.
[0030] It can be seen that to realize the above collective communication function, a collective communication must contain multiple point-to-point communications between rank0, rank1, and rank2. Different collective communication implementation methods (such as collective communication implemented by the ring algorithm or collective communication implemented by the tree structure algorithm) affect the number of point-to-point communications between each communication member in a collective communication. Therefore, once a fault occurs during the collective communication process, it is extremely difficult to locate the fault. The current existing technology generally takes more than 4 hours or even several days to locate the fault generated in the collective communication, which will seriously affect the efficiency of large model training.
[0031] In the fault locating method provided in the embodiments of this specification, each communication member records its own communication log during the collective communication process. The communication log includes the collective communication log and the point-to-point communication log included in the collective communication. When the host discovers an alarm event, the fault is located by analyzing the communication log of each communication member, so as to achieve the purpose of quickly locating the fault in the collective communication.
[0032] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.
[0033] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0034] Figure 2 A schematic diagram of a flow chart of fault location in collective communication provided in this specification includes the following steps:
[0035] S100: In response to an alarm event of collective communication, obtaining a communication log of each communication member in the collective communication.
[0036] In the embodiment of this specification, the fault location method can be applied to the host, that is, the CPU. When a collective communication fails, an alarm event of the collective communication will be generated. When the host finds the alarm event of the collective communication, it can obtain the communication log of each communication member participating in the collective communication. The communication log contains the collective communication log corresponding to the collective communication and the point-to-point communication log corresponding to each point-to-point communication included in the collective communication.
[0037] It should be noted that this specification does not limit the generation mechanism of alarm events.
[0038] S102: Determine a fault location strategy according to the fault type in the alarm event.
[0039] Generally, there are two types of faults included in the alarm events of collective communication, one is the hang fault type, and the other is the slow communication fault type. The so-called hang fault means that the collective communication cannot continue to execute after a certain stage. The so-called slow communication fault means that the collective communication can be executed, but the execution process is very slow. In the embodiments of this specification, corresponding fault location strategies are set for two different types of collective communication faults.
[0040] S104: Analyze the collective communication log and the point-to-point communication log contained in the communication log of each communication member by using the determined fault location strategy, so as to locate the root cause of the alarm event in each communication member.
[0041] Since the external manifestations and root causes of the two fault types are different, they will be described in detail below.
[0042] 1. For the hanging fault type.
[0043] When collective communication hangs, there are usually three main root causes:
[0044] 1. A communication member who should have entered collective communication did not enter collective communication, causing other communication members who have entered collective communication to always wait for the communication member who did not enter collective communication;
[0045] 2. All communication members that should have entered collective communication have entered collective communication, but the collective communication operation performed by one of the communication members is inconsistent with the collective communication operations performed by other communication members; for example, rank0, rank1, and rank2 have all joined the collective communication, and the collective communication operations that should have been performed are all-reduce, while the collective communication operation performed by rank2 is not all-reduce, but reduce, or the collective communication operation that should have been performed is an all-reduce operation on a 4-dimensional array, but the array owned by rank2 is only 2-dimensional;
[0046] 3. All communication members that should enter collective communication have entered collective communication, and the collective communication operations performed are also consistent. However, an occasional error occurs when a communication member performs a collective communication operation. This may be caused by an error in the hardware or driver of the communication member.
[0047] For the above three situations of hang-ups, the fault location method provided in the embodiments of this specification needs to be fully covered.
[0048] Specifically, for the first type of hang-up situation, the collective communication log in the communication log recorded by each communication member needs to include the start time and end time of the collective communication. This is because as long as there is a communication member that should have entered the collective communication but did not enter the collective communication, the start time and end time of this collective communication will not exist in the collective communication log recorded by the communication member. Therefore, for the first type of hang-up situation, the fault location strategy may include: determining the target communication member among each communication member (as for how to determine the target communication member will be described later, and will not be discussed here for the time being), comparing the collective communication start time and the collective communication end time in the collective communication logs of different target communication members, so as to determine the target communication member that did not enter the collective communication as an abnormal communication member. That is, by comparing the start time and end time of this collective communication in the collective communication logs of different target communication members, if there is one or several target communication members whose collective communication logs do not have the start time and end time of this collective communication, then these target communication members are abnormal communication members that should have joined the collective communication but did not join.
[0049] For the above-mentioned second hanging situation, the collective communication log in the communication log recorded by each communication member needs to include the operation information of the collective communication, and the operation information of the collective communication includes but is not limited to the kernel function called and executed by the communication member and the input parameters of the kernel function. This is because if all communication members enter the collective communication, but the collective communication operation performed by one or several communication members is inconsistent with other communication members, then the operation information of this collective communication recorded by these communication members will inevitably be inconsistent with other communication members. Therefore, for the second hanging situation, the fault location strategy may include: determining the target communication member among each communication member, and comparing the operation information of the collective communication in the collective communication logs of different target communication members to determine the target communication member whose operation is different from that of other target communication members as an abnormal communication member. That is, by comparing the operation information of this collective communication in the collective communication logs of different target communication members, if there is one or several target communication members whose operation information of this collective communication in the collective communication log is inconsistent with other communication members, then these target communication members are abnormal communication members.
[0050] For the third hanging situation mentioned above, the point-to-point communication log in the communication log recorded by each communication member needs to include the number of point-to-point communications. This is because if all communication members enter the collective communication and the collective communication operations performed are also consistent, but one or several communication members have occasional errors when performing collective communication operations, then the number of point-to-point communications performed by these communication members in this collective communication will definitely be abnormal (this is because the collective communication itself is implemented by point-to-point communication). Therefore, for the third hanging situation, the fault location strategy may include: determining the target communication member among each communication member, and comparing the number of point-to-point communications in the point-to-point communication logs of different target communication members to determine the target communication member where the point-to-point communication abnormality occurs as the abnormal communication member. That is, compare the number of point-to-point communications performed by different target communication members in this collective communication. If there is one or several target communication members whose point-to-point communication times are inconsistent with other communication members, then these target communication members are abnormal communication members.
[0051] Of course, for the above three hanging situations, the collective communication log in the communication log may include the start time of collective communication, the end time of collective communication, and the operation information of collective communication, and the point-to-point communication log may include the number of point-to-point communications. A preferred method is to analyze the first hanging situation first, then the second hanging situation, and finally the third hanging situation in order of priority. That is, first adopt the method of analyzing the first deadlock situation, compare the start time and the end time of the collective communication, and determine whether there is a target communication member that has not entered the collective communication. If so, directly determine that the target communication member that has not entered the collective communication is the root cause of the deadlock failure. If not, adopt the method of analyzing the second deadlock situation, continue to compare the operation information of the collective communication, and determine whether there is a target communication member whose operation is inconsistent with other target communication members. If so, determine that the target communication member whose operation is inconsistent with other target communication members is the root cause of the deadlock failure. If not, adopt the method of analyzing the third deadlock situation, continue to compare the number of point-to-point communications, and determine whether there is a target communication member that encounters occasional errors in executing the collective communication operation. If so, the target communication member is the root cause of the deadlock failure.
[0052] 2. For the slow communication fault type.
[0053] When collective communication is slow, there are usually two main reasons:
[0054] 1. The calculation process of a certain communication member is too slow (the process of collective communication is generally that the communication member completes the calculation process before executing the collective communication process), resulting in its entry into collective communication too late. Other communication members need to wait for the communication member to complete the calculation before entering collective communication to start collective communication;
[0055] 2. The point-to-point communication process of a communication member is too slow (possibly due to a problem with the network of the communication member), resulting in a slow overall collective communication.
[0056] For the slow communication in the above two situations, the fault location method provided in the embodiments of this specification also needs to achieve comprehensive coverage.
[0057] Specifically, for the first slow communication situation, the collective communication log in the communication log recorded by each communication member needs to include the communication duration of the collective communication. This is because the communication members participating in the collective communication end almost at the same time in this collective communication, and if one or several communication members enter the collective communication too late, then the communication duration of this collective communication recorded by these communication members in the collective communication log will be shorter than that of other communication members. Therefore, for the first slow communication situation, the fault location strategy may include: determining the target communication member among the communication members, comparing the communication duration of the collective communication in the collective communication logs of different target communication members, so as to determine the target communication member whose execution duration of the collective communication is later than the set duration as an abnormal communication member. That is, the communication duration of this collective communication of each target communication member can be determined according to the start time and end time of the collective communication of each target communication member, and the communication duration of different target communication members in this collective communication can be compared. If there is one or several target communication members whose communication duration is shorter than that of other communication members, then these target communication members are abnormal communication members. The set duration can be determined based on the communication duration of the collective communication of all target communication members, such as clustering the communication duration of the collective communication of all target communication members to obtain several clusters, and determining the cluster center of the cluster with the most members as the standard communication duration. If there is a communication member whose collective communication duration is not in the cluster with the most members and is less than the standard communication duration, it can be determined that the communication member is the root cause of the slowness of the entire collective communication process due to the slow calculation process.
[0058] For the second situation of slow communication mentioned above, the point-to-point communication log in the communication log recorded by each communication member needs to include the execution speed of the point-to-point communication included in the collective communication. The execution speed of the point-to-point communication can be characterized by the rate of change of the number of point-to-point communications included in the collective communication. The higher the rate of change, the faster the execution speed of the point-to-point communication, and vice versa. The execution speed of the point-to-point communication is slower. This is because if the execution speed of the point-to-point communication of one or several communication members is too slow, then the number of point-to-point communications executed by these communication members per unit time will inevitably be less than that of other communication members. For example, a normal communication member can perform 8 point-to-point communications within a period of time, while a communication member with a slow point-to-point communication process can only perform 1 or 2 point-to-point communications in the same time, resulting in the entire collective communication being too slow as a whole. Therefore, for the second situation of slow communication, the fault location strategy may include: determining the target communication member among the communication members, and comparing the execution speed of the point-to-point communication in the point-to-point communication logs of different target communication members to determine the target communication member with the point-to-point communication abnormality as the abnormal communication member. That is, compare the change rate of the number of point-to-point communications of different target communication members in this collective communication. If there is one or several target communication members whose change rate of the number of point-to-point communications is smaller than that of other communication members, then these target communication members are abnormal communication members. Among them, the change rate of the number of point-to-point communications is used to indicate the number of point-to-point communications performed by the communication member in a unit time.
[0059] Furthermore, if a communication member wants to record the number of times it performs point-to-point communication within a unit of time, it is generally necessary to record the time of each execution of point-to-point communication, or to record the timestamp of each point-to-point communication in the point-to-point communication log. However, since the communication member in the embodiments of this specification is generally a GPU, it is not easy for a GPU running at high speed and high concurrency to perform an operation such as recording a timestamp. Therefore, in the embodiments of this specification, a counter can be set in the communication member to record the number of point-to-point communications that occur within the communication member itself. The host then collects the value recorded by the counter according to a set period. Then, the host can obtain the rate of change of the number of times the communication member performs point-to-point communication based on the set period and the value of the counter collected each time. For example, if the period is set to 0.5ms, the host will collect the value recorded by the counter every 0.5ms. Assuming that the first value collected for rank0 is 4, and the second value collected is 8, it means that rank0 has performed 8 point-to-point communications within 1ms, and the first value collected for rank1 is 1, and the second value collected is 2, and so on. The value collected for the 8th time is 8, which means that rank1 has performed 8 point-to-point communications within 4ms, that is, rank1 can only perform 2 point-to-point communications within 1ms, which is significantly lower than rank0.
[0060] When the host issuing the kernel function is the CPU and the communication member participating in the collective communication is the GPU, the GPU can write the number of its own point-to-point communications in the CPU's page-locked memory. Correspondingly, when the CPU obtains the change rate of the number of GPU point-to-point communications in the communication log, it can obtain the number of point-to-point communications written by the GPU when executing collective communication from its own page-locked memory according to the set period.
[0061] Considering that in actual application scenarios, each communication member may participate in different collective communications multiple times during the entire large model training process, and for the second situation of slow communication, it is necessary to compare the change rate of the number of point-to-point communications of each communication member in the same collective communication. Therefore, it is necessary to associate the data written in the pinned memory by different communication members for the same collective communication. Therefore, in the embodiment of this specification, a monitoring identifier (traceid) is introduced to associate the data written in the pinned memory by different communication members for the same collective communication.
[0062] The data format of the trace id can be in any format as long as each collective communication corresponds to a unique trace id, as shown in Table 1.
[0063]
[0064] Table 1
[0065] In the data structure of the trace id shown in Table 1 above, the communication domain id indicates the communication domain where each communication member monitored by the host is located. Only the communication members in the same communication domain can participate in collective communication. One communication domain uniquely corresponds to one communication domain id. The collective communication identifier is used to uniquely identify a collective communication. The extension bit can be used as a reserved position to record other information.
[0066] Correspondingly, the data written by the communication member into the pinned memory also needs to include the above trace id, as shown in Table 2 and
[0067] As shown in Table 3.
[0068]
[0069] Table 2
[0070] Table 2 is an exemplary data header provided in an embodiment of this specification, wherein KernelIndex (hereinafter referred to as KID) is used to index n blocks of data in the data body part. Mode is used to indicate the data recording mode, such as not recording any data, only recording the number of point-to-point communications in a certain collective communication, recording the number of point-to-point communications in all collective communications, etc. Counter is used to indicate the number of collective communications. DetectChannel is used to indicate the channels used by each communication member to perform collective communication, and trace id is the trace id shown in Table 1.
[0071]
[0072]
[0073] Table 3
[0074] Following Table 2, Table 3 divides the body of the data into n blocks of data, each of which records the number of times the communication members perform point-to-point communication on different channels.
[0075] Through the above method, the embodiments of this specification can cover 3 situations where a hang failure occurs in collective communication and 2 situations where a communication slow failure occurs.
[0076] Since the algorithms for implementing collective communication mainly include the ring algorithm and the tree structure algorithm, and the analysis methods for the above-mentioned five situations provided in the embodiments of this specification all require comparison of the communication logs recorded by different communication members, therefore, for different algorithms for implementing collective communication, it is also necessary to adaptively determine which communication members' communication logs need to be compared.
[0077] Specifically, for collective communication implemented using the ring algorithm, the number of point-to-point communications performed by each communication member and the rate of change of the number should theoretically be equal and consistent, such as Figure 3 Therefore, in the above analysis of the five fault scenarios, if the algorithm used in the collective communication is the ring algorithm, all communication members are taken as target communication members for comparative analysis.
[0078] For collective communication implemented by tree structure algorithm, only the number of point-to-point communications performed by communication members at the same level of the tree structure and the rate of change of the number should be equal and consistent in theory, such as Figure 4 As shown in the figure, the communication members in the dashed box are located at the same layer of the tree structure. Therefore, in the above analysis of the five fault scenarios, if the algorithm used for collective communication is a tree structure algorithm, the communication members at the same layer of the tree structure will be used as target communication members for comparison and analysis, and cross-layer comparison is not allowed.
[0079] Through the above method, most of the root causes of hangs and slow communication failures in collective communications can be effectively covered, and the root causes of failures can be accurately located or demarcated in collective communications composed of a large number of communication members and complex links.
[0080] The above is one or more implementations of the method for locating a fault in collective communication of this specification. Based on the same idea, this specification also provides a corresponding device for locating a fault in collective communication, such as Figure 5 shown.
[0081] Figure 5 A schematic diagram of a structure of a fault location device in collective communication provided in this specification includes:
[0082] The log acquisition module 501 is used to obtain the communication log of each communication member in the collective communication in response to the alarm event of the collective communication; the communication log includes the collective communication log corresponding to the collective communication and the point-to-point communication log corresponding to each point-to-point communication included in the collective communication;
[0083] A strategy selection module 502, configured to determine a fault location strategy according to the fault type in the alarm event;
[0084] The analysis module 503 is used to analyze the collective communication log and the point-to-point communication log contained in the communication log of each communication member by using the determined fault location strategy, so as to locate the root cause of the alarm event in each communication member.
[0085] Optionally, the fault type includes a hang fault type;
[0086] The analysis module 503 is specifically configured to adopt a fault location strategy corresponding to the hang fault type to compare the communication logs of each communication member to determine the abnormal communication member.
[0087] Optionally, the collective communication log includes: collective communication start time, collective communication end time, and collective communication operation information;
[0088] The point-to-point communication log includes: the number of point-to-point communications included in the collective communication;
[0089] The analysis module 503 is specifically used to determine the target communication member among the communication members; and compare at least one of the start time of the collective communication, the end time of the collective communication, the operation information of the collective communication in the collective communication log of different target communication members, and the number of point-to-point communications in the point-to-point communication log.
[0090] Optionally, the analysis module 503 is specifically used to compare the collective communication start time and the collective communication end time in the collective communication logs of different target communication members to determine the target communication members who have not entered the collective communication as abnormal communication members; and / or, to compare the collective communication operation information in the collective communication logs of different target communication members to determine the target communication members whose operations are different from those of other target communication members as abnormal communication members; and / or, to compare the number of point-to-point communications in the point-to-point communication logs of different target communication members to determine the target communication members where point-to-point communication abnormalities occur as abnormal communication members.
[0091] Optionally, the fault type includes a slow communication fault type;
[0092] The analysis module 503 is specifically configured to adopt a fault location strategy corresponding to the slow communication fault type to compare the communication logs of each communication member to identify an abnormal communication member.
[0093] Optionally, the collective communication log includes: the communication duration of the collective communication;
[0094] The point-to-point communication log includes: the execution speed of the point-to-point communication included in the collective communication;
[0095] The analysis module 503 is specifically used to determine a target communication member among the communication members; and compare at least one of the communication duration of the collective communication in the collective communication log and the execution speed of the point-to-point communication in the point-to-point communication log of different target communication members.
[0096] Optionally, the analysis module 503 is specifically used to compare the communication duration of collective communications in the collective communication logs of different target communication members to determine the target communication member whose communication duration of collective communication is less than the set time as an abnormal communication member; and / or to compare the execution speed of point-to-point communication in the point-to-point communication logs of different target communication members to determine the target communication member where the point-to-point communication abnormality occurs as an abnormal communication member.
[0097] Optionally, the analysis module 503 is specifically used to take all communication members as target communication members if the algorithm adopted by the collective communication is a ring algorithm; if the algorithm adopted by the collective communication is a tree structure algorithm, take communication members located at the same layer of the tree structure as target communication members.
[0098] Optionally, the method is applied to a host, the host comprising a CPU;
[0099] The communication member includes at least one of a GPU, an NPG, and a PPU;
[0100] The point-to-point communication log includes: the execution speed of the point-to-point communication included in the collective communication;
[0101] The log acquisition module 501 is specifically used to acquire the execution speed of the point-to-point communication of each communication member in the collective communication from a preset page-locked memory.
[0102] Optionally, the log acquisition module 501 is specifically used to read the number of times each communication member performs point-to-point communication in the collective communication written by each communication member from the page-locked memory according to a set period; and obtain the execution speed of the point-to-point communication of each communication member in the collective communication based on the set period and the number of times read.
[0103] This specification also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 2 A fault location method in collective communication is provided.
[0104] This manual also provides Figure 6 The one shown corresponds to Figure 2 A schematic diagram of the electronic device. Figure 6 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 2 Of course, in addition to the software implementation, this specification does not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0105] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.
Claims
1. A method for locating a fault in collective communication, the method comprising: In response to an alarm event of the collective communication, obtaining a communication log of each communication member in the collective communication; The communication log includes a collective communication log corresponding to the collective communication and a point-to-point communication log corresponding to each point-to-point communication included in the collective communication; Determine a fault location strategy according to the fault type in the alarm event; The determined fault location strategy is used to analyze the collective communication log and the point-to-point communication log contained in the communication log of each communication member, so as to locate the root cause of the alarm event in each communication member.
2. The method according to claim 1, wherein the fault type comprises a hang fault type; The determined fault location strategy is used to analyze the collective communication log and the point-to-point communication log contained in the communication log of each communication member, specifically including: The communication logs of each communication member are compared by using a fault location strategy corresponding to the hang fault type to determine the abnormal communication member.
3. The method according to claim 2, wherein the aggregate communication log comprises: Collective communication start time, collective communication end time, and collective communication operation information; The point-to-point communication log includes: the number of point-to-point communications included in the collective communication; Adopt the fault location strategy corresponding to the hang fault type and compare the communication logs of each communication member, specifically including: determining a target communication member among the communication members; At least one of the start time of collective communication, the end time of collective communication, and the operation information of collective communication in the collective communication logs of different target communication members and the number of point-to-point communications in the point-to-point communication logs is compared.
4. The method according to claim 3, comprising comparing at least one of the collective communication start time, the collective communication end time, the collective communication operation information in the collective communication logs of different target communication members, and the number of point-to-point communications in the point-to-point communication logs, and specifically comprising: Comparing the collective communication start time and the collective communication end time in the collective communication logs of different target communication members to determine the target communication members who have not entered the collective communication as abnormal communication members; and / or Comparing the collective communication operation information in the collective communication logs of different target communication members to identify the target communication member whose operation is different from that of other target communication members as an abnormal communication member; and / or The number of point-to-point communications in the point-to-point communication logs of different target communication members is compared to determine the target communication member where the point-to-point communication abnormality occurs as the abnormal communication member.
5. The method according to claim 1, wherein the fault type comprises a slow communication fault type; The determined fault location strategy is used to analyze the collective communication log and the point-to-point communication log contained in the communication log of each communication member, specifically including: A fault location strategy corresponding to the slow communication fault type is adopted to compare the communication logs of each communication member to determine the abnormal communication member.
6. The method of claim 5, wherein the aggregate communication log comprises: The communication duration of collective communication; The point-to-point communication log includes: the execution speed of the point-to-point communication included in the collective communication; Adopt the fault location strategy corresponding to the communication slow fault type and compare the communication logs of each communication member, specifically including: determining a target communication member among the communication members; At least one of the communication duration of the collective communication in the collective communication log and the execution speed of the point-to-point communication in the point-to-point communication log of different target communication members is compared.
7. The method according to claim 6, wherein at least one of the communication duration of the collective communication in the collective communication logs of different target communication members and the execution speed of the point-to-point communication in the point-to-point communication logs is compared, specifically comprising: Comparing the communication durations of collective communications in collective communication logs of different target communication members, to determine the target communication member whose collective communication duration is less than the set time as an abnormal communication member; and / or The execution speeds of the point-to-point communications in the point-to-point communication logs of different target communication members are compared to determine the target communication member where the point-to-point communication anomaly occurs as the abnormal communication member.
8. The method according to claim 3 or 6, wherein determining a target communication member from among the communication members comprises: If the algorithm adopted by the collective communication is the ring algorithm, all communication members are regarded as target communication members; If the algorithm adopted by the collective communication is a tree structure algorithm, the communication members located at the same layer of the tree structure are taken as target communication members.
9. The method according to claim 1, wherein the method is applied to a host, wherein the host comprises a CPU; The communication member includes at least one of a GPU, an NPG, and a PPU; The point-to-point communication log includes: the execution speed of the point-to-point communication included in the collective communication; Obtaining a communication log for each communication member in the collective communication, specifically including: The execution speed of the point-to-point communication of each communication member in the collective communication is obtained from a preset page-locked memory.
10. The method according to claim 9, wherein obtaining the execution speed of the point-to-point communication of each communication member in the collective communication from the preset page-locked memory comprises: Reading the number of times each communication member performs point-to-point communication in the collective communication from the page-locked memory according to a set period; The execution speed of the point-to-point communication of each communication member in the collective communication is obtained based on the set cycle and the number of times read.
11. A fault location device in collective communication, comprising: A log acquisition module, configured to acquire a communication log of each communication member in the collective communication in response to an alarm event of the collective communication; The communication log includes a collective communication log corresponding to the collective communication and a point-to-point communication log corresponding to each point-to-point communication included in the collective communication; A strategy selection module, used to determine a fault location strategy according to the fault type in the alarm event; The analysis module is used to analyze the collective communication log and the point-to-point communication log contained in the communication log of each communication member by adopting the determined fault location strategy, so as to locate the root cause of the alarm event in each communication member.
12. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 10.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 10 when executing the program.