Communication algorithm selection method and device, nonvolatile storage medium and electronic equipment

By generating traffic test cases corresponding to the computing cluster, simulating real traffic and selecting the communication algorithm with the shortest execution time, the problem of insufficient accuracy and flexibility of existing testing methods is solved, and efficient communication algorithm selection and cluster performance optimization in practical application scenarios are realized.

CN120301801APending Publication Date: 2025-07-11CHINA TELECOM CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510422483.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing collective communication testing methods lack accuracy, versatility and flexibility, and cannot fully evaluate the performance of the collective communication performance of intelligent computing clusters in different scenarios, making it difficult to choose the optimal communication algorithm.

Method used

By generating traffic test cases corresponding to the computing cluster, simulating real traffic based on topology and parallel policies, configuring candidate communication algorithms, executing traffic test cases, and selecting the target communication algorithm with the shortest execution time, while monitoring traffic differences to detect topological failures.

Benefits of technology

It realizes the accurate evaluation and selection of the optimal communication algorithm in practical application scenarios, improves the accuracy and efficiency of testing, adapts to multiple heterogeneous clusters, and optimizes cluster performance and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120301801A_ABST
    Figure CN120301801A_ABST
Patent Text Reader

Abstract

The invention discloses a communication algorithm selection method and device, a nonvolatile storage medium and electronic equipment. The method comprises the steps that a traffic test case corresponding to a computing cluster is generated according to configuration information of the computing cluster, the configuration information comprises topological structure information of the computing cluster, and the traffic test case comprises input traffic size information and output traffic size information corresponding to one or more interfaces in a topological structure; a candidate communication algorithm is configured in the computing cluster, and the candidate communication algorithm is used for indicating one or more interfaces to communicate according to a preset rule; executing the traffic test case in the computing cluster, and obtaining the execution duration of each candidate communication algorithm when executing the traffic test case; and selecting a target communication algorithm from the candidate communication algorithms according to the execution duration of each candidate communication algorithm. According to the invention, the technical problem that the communication algorithm is difficult to select due to the lack of authenticity of the existing set communication test is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communications, and in particular, to a method and apparatus for selecting a communication algorithm, a non-volatile storage medium, and an electronic device. Background Art

[0002] With the rapid development of artificial intelligence technology, the close combination of large models and intelligent computing power is driving all industries to transform towards "AI+", comprehensively reconstructing production processes and lifestyles. Against this background, the parameter scale of large models has been continuously leaping from tens of billions and hundreds of billions to trillions and hundreds of trillions, and the demand for intelligent computing power resources is increasing day by day. Large-scale intelligent computing clusters have become the standard infrastructure to support the training of mainstream large models, and leading domestic and international enterprises have all built or plan to build intelligent computing clusters with tens of thousands of cards and more than tens of thousands of cards.

[0003] However, the linear increase in the scale of the intelligent computing cluster does not directly bring a linear increase in the effective computing power of the cluster. The performance of collective communication is an important indicator to measure the performance of a large-scale intelligent computing cluster. Collective communication algorithms can intelligently utilize the network topology structure and, through communication algorithms such as RING, TREE, and COLLNET, achieve data or gradient exchange and aggregation among multiple GPU cards, and achieve high network throughput between GPUs with the RDMA high-performance protocol technology to meet the data synchronization requirements during large model training. Its performance directly affects the computing efficiency and stability of the entire cluster.

[0004] However, existing collective communication testing methods often lack accuracy, generality, and flexibility, and cannot comprehensively evaluate the performance of collective communication in an intelligent computing cluster under different scenarios, making it difficult to select the optimal communication algorithm.

[0005] No effective solution has been proposed for the above problems. Summary of the Invention

[0006] Embodiments of the present application provide a method and apparatus for selecting a communication algorithm, a non-volatile storage medium, and an electronic device to at least solve the technical problem of difficult selection of communication algorithms due to the lack of authenticity in existing collective communication testing.

[0007] According to one aspect of the embodiments of the present application, a communication algorithm selection method is provided, including: generating a traffic test case corresponding to a computing cluster according to the configuration information of the computing cluster, where the configuration information includes the topology structure information of the computing cluster, and the traffic test case includes the input traffic size information and the output traffic size information corresponding to one or more interfaces in the topology structure; configuring candidate communication algorithms in the computing cluster, where the candidate communication algorithms are used to instruct one or more interfaces to communicate according to preset rules; executing the traffic test case in the computing cluster, and obtaining the execution duration of each candidate communication algorithm when executing the traffic test case; selecting a target communication algorithm from the candidate communication algorithms according to the execution duration of each candidate communication algorithm, where the execution duration corresponding to the target communication algorithm is less than or equal to the execution duration corresponding to the candidate communication algorithm.

[0008] Optionally, after selecting a target communication algorithm from the candidate communication algorithms according to the execution duration of each candidate communication algorithm, the method further includes: applying the target communication algorithm in the computing cluster; monitoring the real-time input traffic size information and the real-time output traffic size information of the interfaces in the computing cluster; comparing the real-time input traffic size information of the interfaces with the input traffic size information corresponding to the interfaces in the test case, and determining that there is a potential fault in the topology structure corresponding to the interfaces when it is detected that the difference between the real-time input traffic size information and the input traffic size information is greater than a first preset threshold; comparing the real-time output traffic size information of the interfaces with the output traffic size information corresponding to the interfaces in the test case, and determining that there is a potential fault in the topology structure corresponding to the interfaces when it is detected that the difference between the real-time output traffic size information and the output traffic size information is greater than a second preset threshold; sending an alarm message when it is determined that there is a potential fault in the topology structure corresponding to the interfaces, where the alarm message is used to instruct to check the topology structure corresponding to the interfaces.

[0009] Optionally, the configuration information of the computing cluster further includes the identification information of the large model deployed in the computing cluster. Generating a traffic test case corresponding to the computing cluster according to the configuration information of the computing cluster includes: determining the historical traffic data corresponding to the identification information in the historical traffic database, where the historical traffic data includes parallel strategies, and the input traffic size information and the output traffic size information corresponding one-to-one to the parallel strategies, and the parallel strategies include at least one of the following: data parallelism, tensor parallelism, pipeline parallelism; determining the parallel strategy followed when communicating between each interface according to the configuration information of the computing cluster; generating a traffic test case according to the parallel strategy followed when communicating between each interface and the historical traffic data.

[0010] Optionally, before determining the historical traffic data corresponding to the identification information in the traffic database, the method further includes: obtaining first communication data of the large model deployed in the historical computing cluster in one iteration, where the first communication data includes the first input traffic size information, the first output traffic size information corresponding to one or more interfaces in the historical computing cluster, and the traffic transmission paths of one or more interfaces; determining, according to the first communication data and the topological structure of the historical computing cluster, the parallel strategy followed by the first input traffic size information and the first output traffic size information of one or more interfaces; counting the first input traffic size information and the first output traffic size information corresponding to each parallel strategy; storing the identification information of the large model, each parallel strategy, the first input traffic size information corresponding to the parallel strategy, and the first output traffic size information corresponding to the parallel strategy into the historical traffic database.

[0011] Optionally, before determining, according to the first communication data and the topological structure of the historical computing cluster, the parallel strategy followed by the first input traffic size information and the first output traffic size information of one or more interfaces, the method further includes: obtaining multiple rounds of second communication data of the large model in multiple iterations; determining the average value of the multiple rounds of second communication data as the third communication data; comparing the third communication data with the first communication data to determine the difference in the input-output traffic size corresponding to one or more interfaces; in the case where the difference in the input-output traffic size corresponding to the interfaces in the third communication data and the first communication data is greater than or equal to the third preset threshold, re-obtaining the first communication data and comparing the third communication data with the new first communication data until the difference in the input-output traffic size corresponding to one or more interfaces in the third communication data and the new first communication data is less than the third preset threshold.

[0012] Optionally, before executing the traffic test case in the computing cluster, the method further includes: determining the type of the graphics processor in the computing cluster; determining the corresponding mirror data according to the type of the graphics processor, where the mirror data is compatible with the graphics processor, and the mirror data includes a test program client for receiving the test case and executing the test case to send the mirror data to the computing cluster.

[0013] Optionally, determining the corresponding mirror data according to the type of the graphics processor includes: querying a compatibility list according to the type of the graphics processor, where the compatibility list records each mirror data identifier and the compatible type corresponding to the mirror data identifier; in the case where the compatible type of the mirror data includes the type of the graphics processor, determining the mirror data as the mirror data corresponding to the type of the graphics processor.

[0014] According to another aspect of the embodiments of the present application, there is also provided a communication algorithm selection device, including: a generation module, configured to generate a traffic test case corresponding to a computing cluster according to the configuration information of the computing cluster, where the configuration information includes the topology structure information of the computing cluster, and the traffic test case includes the input traffic size information and the output traffic size information corresponding to one or more interfaces in the topology structure; a configuration module, configured to configure candidate communication algorithms in the computing cluster, where the candidate communication algorithms are used to instruct one or more interfaces to communicate according to preset rules; an execution module, configured to execute the traffic test case in the computing cluster and obtain the execution duration of each candidate communication algorithm when executing the traffic test case; a selection module, configured to select a target communication algorithm from the candidate communication algorithms according to the execution duration of each candidate communication algorithm, where the execution duration corresponding to the target communication algorithm is less than or equal to the execution duration corresponding to the candidate communication algorithm.

[0015] According to another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium storing a program, where when the program runs, it controls the device where the non-volatile storage medium is located to execute the communication algorithm selection method.

[0016] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, where the processor is configured to run the program stored in the memory, and when the program runs, it executes the communication algorithm selection method.

[0017] According to another aspect of the embodiments of the present application, there is also provided a computer program product including a computer program, where when the computer program is executed by a processor, it implements the communication algorithm selection method.

[0018] In the embodiments of the present application, traffic test cases corresponding to a computing cluster are generated based on the configuration information of the computing cluster. The configuration information includes the topological structure information of the computing cluster, and the traffic test cases include the input traffic size information and the output traffic size information corresponding to one or more interfaces in the topological structure. A candidate communication algorithm is configured in the computing cluster, where the candidate communication algorithm is used to instruct one or more interfaces to communicate according to preset rules. The traffic test cases are executed in the computing cluster, and the execution duration of each candidate communication algorithm when executing the traffic test cases is obtained. The target communication algorithm is selected from the candidate communication algorithms based on the execution duration of each candidate communication algorithm. The execution duration corresponding to the target communication algorithm is less than or equal to the execution duration corresponding to the candidate communication algorithm. By generating real test cases corresponding to the deployment scenario, the purpose of quickly and accurately evaluating the performance of each communication algorithm is achieved, thereby realizing the technical effect of accurately selecting the optimal communication algorithm in the current deployment environment, and further solving the technical problem of difficult selection of communication algorithms due to the lack of authenticity in the existing collective communication tests. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0020] Figure 1 is a schematic structural diagram of a computer terminal provided according to an embodiment of the present application;

[0021] Figure 2 is a schematic flowchart of a method for selecting a communication algorithm provided according to an embodiment of the present application;

[0022] Figure 3 is a schematic overall flowchart of a method for selecting a communication algorithm provided according to an embodiment of the present application;

[0023] Figure 4 is a schematic structural diagram of a device for selecting a communication algorithm provided according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0025] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained as follows:

[0027] Large model training: Large model training refers to the training process of machine learning models with large-scale parameters and complex computing structures. These models are usually constructed by deep neural networks and have billions or even hundreds of billions of parameters. The design goal of large models is to improve the model's expressive ability and prediction performance, enabling it to handle more complex tasks and data. To achieve this goal, high-performance intelligent computing clusters are usually required.

[0028] Intelligent computing cluster: An intelligent computing cluster is a comprehensive computing network ecosystem that integrates computing power, computing quotient, algorithms, data, and applications. It represents an important development direction in the future technology field. The core of an intelligent computing cluster is to provide powerful computing power support and complete the delivery of intelligent computing power in a service-oriented manner, and is commonly used to support scenarios such as large model training and inference.

[0029] Collective communication: Collective communication is an important concept in parallel computing. It refers to communication operations carried out among a group of processes, and all processes participate in it. Such communication operations include a series of standard information exchange interfaces to solve the communication problems between different processes during parallel computing. The basic operations of collective communication can be combined into different communication templates, also known as communication primitives, such as P2P, AllReduce, AlltoAll, All-Gather, Reduce-Scatter, etc. The key to collective communication lies in communication efficiency and the best application of the network hardware connection topology structure. It is particularly important in distributed training because a large amount of data communication needs to be carried out between multiple hardware devices.

[0030] Collective communication library: Provides interfaces for implementing communication primitives, enabling developers to more easily perform efficient data exchange and collaborative work on multiple GPUs or multiple nodes.

[0031] RING Algorithm (Ring Algorithm): A collective communication algorithm that forms a ring among nodes for data or gradient exchange. In the ring algorithm, each node communicates only with its two adjacent nodes, and data or gradients are passed along the order of the ring.

[0032] TREE Algorithm (Tree Algorithm): A collective communication algorithm based on a tree structure, used for aggregating data or gradients in a parallel system. The tree algorithm usually centers around one or more root nodes, and data or gradients are aggregated from leaf nodes up to the root nodes level by level.

[0033] COLLNET Algorithm (Collective Network Algorithm): A collective communication algorithm that offloads the data reduction operation to the network switch. The switch directly participates in data processing, reducing the communication requirements between GPUs, thereby optimizing the use of network resources.

[0034] In the related art, existing collective communication testing methods often lack accuracy, generality, and flexibility, and cannot comprehensively evaluate the performance of collective communication in different scenarios. It is difficult to accurately select a communication algorithm suitable for the current intelligent computing cluster. Specifically, the current communication algorithm selection method has the following disadvantages:

[0035] 1. Lack of actual use cases: Currently, in the performance evaluation and selection of communication algorithms, collective communication tests are often conducted by sending traffic with multiple fixed information sizes. However, the traffic distribution of such test cases does not match the real traffic distribution in actual large model training, resulting in test results that cannot guide actual operations, nor can they represent the performance of communication algorithms in actual applications, and cannot guide the determination of the most suitable communication algorithm. At the same time, in different communication scenarios, such as intra-node communication and inter-node communication, different algorithms (such as Ring algorithm and Tree algorithm) need to be selected to optimize performance. However, this selection increases the complexity of test implementation, and the test results are also difficult to reflect the actual performance, exacerbating the difficulty of algorithm selection.

[0036] 2. Lack of cross-architecture support: Current test schemes often only target intelligent computing clusters of a certain type of chip. In reality, there are significant heterogeneities in the physical connection methods and topologies among different GPUs, which increases the complexity of collective communication testing.

[0037] To solve the above problems, relevant solutions are provided in the embodiments of this application, which are described in detail below.

[0038] According to an embodiment of the present application, a method embodiment of a communication algorithm selection method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0039] The method embodiment provided by the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal for implementing the communication algorithm selection method is shown. As Figure 1 shown, the computer terminal 10 may include one or more (shown as 102a, 102b,..., 102n in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.

[0040] It should be noted that the above one or more processors 102 and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit can be embodied in software, hardware, firmware or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10. As involved in the embodiment of the present application, the data processing circuit is used for processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0041] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the communication algorithm selection method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned communication algorithm selection method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0042] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0043] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 10.

[0044] Under the above operating environment, an embodiment of the present application provides a communication algorithm selection method, as Figure 2 shown, the method includes the following steps:

[0045] Step S202, generating a traffic test case corresponding to the computing cluster according to the configuration information of the computing cluster, where the configuration information includes the topology structure information of the computing cluster, and the traffic test case includes input traffic size information and output traffic size information corresponding to one or more interfaces in the topology structure.

[0046] In the technical solution provided in step S202, the configuration information of the computing cluster further includes the identification information of the large model deployed in the computing cluster. Generating a traffic test case corresponding to the computing cluster according to the configuration information of the computing cluster includes: determining historical traffic data corresponding to the identification information in the historical traffic database, where the historical traffic data includes parallel strategies, and input traffic size information and output traffic size information corresponding to the parallel strategies one by one. The parallel strategies include at least one of the following: data parallelism, tensor parallelism, and pipeline parallelism; determining the parallel strategy followed when communicating between interfaces according to the configuration information of the computing cluster; generating a traffic test case according to the parallel strategy followed when communicating between interfaces and the historical traffic data.

[0047] Optionally, determining the parallel strategy followed when communicating between interfaces according to the configuration information of the computing cluster includes determining parallel rules through the structure of the model to be deployed in the computing cluster (such as the number of attention heads, the number of layers, and the hidden dimension, and the size of the cluster), and following the rule that the number of cards = tp (tensor parallelism number) * dp (data parallelism number) * pp (pipeline parallelism number). For example, a parallel mode for 64 cards is: tp4, dp2, pp8. According to the parallel strategy followed when communicating between interfaces and the historical traffic data, determine the traffic size to be transmitted between interfaces, generate a series of communication steps in large model training, and generate corresponding communication loads according to the traffic (i.e., historical traffic data) in actual training for these communication steps.

[0048] Optionally, generating a traffic test case according to the parallel strategy followed when communicating between interfaces and the historical traffic data includes comparing the gap between the configuration information of the current computing cluster and the configuration information of the historical traffic data, and converting the historical traffic data according to a preset rule. For example, the configuration of the current computing cluster is tp4, dp2, pp8, while the configuration of the historical traffic data is tp2, dp2, pp8, and in one model iteration, the communication load related to tensor parallelism is x MB. When adjusting the parallel strategy and increasing tp from 4 to 8, it means that the data exchange originally carried out between 4 GPUs now needs to be carried out between 8 GPUs. Due to the increase in the number of GPUs, the overall communication load will increase significantly. Therefore, in order to simulate the communication requirements under the tp8 configuration, double the previous communication load related to tp4 according to the preset rule, that is, it becomes 2x MB, as the test case of the current computing cluster.

[0049] As an alternative implementation, before determining the historical traffic data corresponding to the identification information in the traffic database, the method further includes: obtaining first communication data of the large model deployed in the historical computing cluster during one iteration, where the first communication data includes the first input traffic size information, the first output traffic size information corresponding to one or more interfaces in the historical computing cluster, and the traffic transmission paths of one or more interfaces; determining, based on the first communication data and the topology of the historical computing cluster, the parallel strategy followed by the first input traffic size information and the first output traffic size information of one or more interfaces; counting the first input traffic size information and the first output traffic size information corresponding to each parallel strategy; storing the identification information of the large model, each parallel strategy, the first input traffic size information corresponding to the parallel strategy, and the first output traffic size information corresponding to the parallel strategy in the historical traffic database.

[0050] Optionally, obtaining the first communication data of the large model deployed in the historical computing cluster during one iteration includes real-time collecting the collective communication traffic data during the large model training process in the intelligent computing cluster, mainly collecting the traffic bandwidth of rdma switches, rdma network cards, pcie switches, XPU switches, or bus interfaces, etc., and the real input and output data of some of the iterations, and analyzing the collected real large model training. Since each iteration in the large model training process has repeatability and similarity, only extract the real collected data of some of the iterations for modeling; model according to the time nodes and input and output data of the collective communication data of a single iteration to form an accurate model.

[0051] Specifically, since the actual network traffic of each iteration is similar, mainly capture the data of one iteration (i.e., the first communication data), and obtain which data is currently transmitted from which card through which link to which card by capturing the bandwidth of each network card, pcie switch, XPU switch, or bus interface, etc. in one iteration, to form a sample model of the real network traffic, and cooperate with the buried points added in the collective communication library and the training framework to determine the communication data of different parallel strategies such as tp (tensor parallelism), dp (data parallelism), and pp (pipeline parallelism) corresponding to the traffic, so as to determine the traffic size corresponding to each parallel strategy for the large model deployed in the current computing cluster.

[0052] Optionally, before determining the parallel strategy followed by the first input traffic size information and the first output traffic size information of one or more interfaces based on the first communication data and the topology of the historical computing cluster, the method further includes: obtaining multi-round second communication data of the large model in multiple rounds of iteration; determining the average value of the multi-round second communication data as the third communication data; comparing the third communication data with the first communication data to determine the difference in the input and output traffic sizes corresponding to one or more interfaces; in the case where the difference in the input and output traffic sizes corresponding to the interfaces in the third communication data and the first communication data is greater than or equal to the third preset threshold, re-obtaining the first communication data and comparing the third communication data with the new first communication data until the difference in the input and output traffic sizes corresponding to one or more interfaces in the third communication data and the new first communication data is less than the third preset threshold.

[0053] Optionally, after obtaining the first communication data, model verification is performed based on the data of other iterations (i.e., multiple second communication data) to ensure that the first communication data is representative data when the model communication is stable. Specifically, calculate the average value of the input and output traffic sizes corresponding to each interface in the multiple second communication data as the third communication data, and then determine whether the difference between each input and output traffic size and the first communication is less than the third preset threshold. In the case of no, it indicates that the error is large, re-obtain the first communication data, and repeat the above steps until the error between the model and the real data is within a certain range (i.e., the difference in the input and output traffic sizes corresponding to one or more interfaces in the third communication data and the new first communication data is less than the third preset threshold), complete the construction of the traffic model, and store the identification information of the large model, each parallel strategy, the first input traffic size information corresponding to the parallel strategy, and the first output traffic size information corresponding to the parallel strategy in the historical traffic database.

[0054] Optionally, since the traffic characteristics of each large model are different during training, different large models are deployed multiple times in the computing cluster for training, the first communication data of multiple models is obtained, and a historical traffic database containing the first communication data of different large models is formed.

[0055] Step S204, configure candidate communication algorithms in the computing cluster, where the candidate communication algorithms are used to instruct one or more interfaces to communicate according to preset rules.

[0056] Step S206, execute the traffic test case in the computing cluster and obtain the execution duration of each candidate communication algorithm when executing the traffic test case.

[0057] In the technical solution provided in step S206, before executing the traffic test case in the computing cluster, the method further includes: determining the type of the graphics processor in the computing cluster; determining the corresponding mirror data according to the type of the graphics processor, where the mirror data is compatible with the graphics processor, and the mirror data includes a test program client for receiving the test case and executing the test case to send the mirror data to the computing cluster.

[0058] As an optional implementation manner, determining the corresponding mirror data according to the type of the graphics processor includes: querying a compatibility list according to the type of the graphics processor, where the compatibility list records each mirror data identifier and the compatible type corresponding to the mirror data identifier; in the case that the compatible type of the mirror data includes the type of the graphics processor, determining the mirror data as the mirror data corresponding to the type of the graphics processor.

[0059] Optionally, according to different chips adopted by different computing clusters, different adapted container images (i.e., mirror files) are used to execute the test, and a C / S architecture is adopted between the running test program and the test platform. The test program can accept the tasks of the test platform and execute the corresponding commands, thus avoiding binding to specific hardware.

[0060] Step S208, select a target communication algorithm from the candidate communication algorithms according to the execution duration of each candidate communication algorithm, where the execution duration corresponding to the target communication algorithm is less than or equal to the execution duration corresponding to the candidate communication algorithm.

[0061] Optionally, collect the collective communication data of the intelligent computing cluster set for executing the test case, and evaluate the performance of different algorithms (i.e., candidate communication algorithms) in the current scenario by sensing the bus bandwidth between the XPU cards and the PCIe and the bandwidth and topology of the RDMA network. According to the evaluation result (i.e., the execution duration), select the optimal (the one with the minimum execution duration) different collective communication algorithms such as tree or ring (i.e., the target communication algorithm) to maximize the collective communication performance of the cluster.

[0062] In the technical solution provided in step S208, after selecting the target communication algorithm from the candidate communication algorithms according to the execution duration of each candidate communication algorithm, the method further includes: applying the target communication algorithm in the computing cluster; monitoring the real-time input traffic size information and the real-time output traffic size information of the interfaces in the computing cluster; comparing the real-time input traffic size information of the interface with the input traffic size information corresponding to the interface in the test case, and determining that there is a potential fault in the topology structure corresponding to the interface when it is detected that the difference between the real-time input traffic size information and the input traffic size information is greater than the first preset threshold; comparing the real-time output traffic size information of the interface with the output traffic size information corresponding to the interface in the test case, and determining that there is a potential fault in the topology structure corresponding to the interface when it is detected that the difference between the real-time output traffic size information and the output traffic size information is greater than the second preset threshold; and sending an alarm message when it is determined that there is a potential fault in the topology structure corresponding to the interface, where the alarm message is used to indicate checking the topology structure corresponding to the interface.

[0063] Optionally, after selecting the target communication algorithm, it further includes real-time monitoring of the traffic bandwidth data of rdma switches, rdma network cards, pcieswitches, XPUswitches, or bus interfaces, etc., obtaining the collective communication traffic during the acquisition process, calculating the total collective communication bandwidth and the input and output traffic sizes of each interface, comparing with the test case, and locating the faulty hardware and sending an alarm message when the difference from the test case traffic size is too large.

[0064] An embodiment of the present application provides an organized flowchart of a communication algorithm selection method, as Figure 3 shown, the method includes the following steps:

[0065] Step S301, collect real model training collective communication data: For large model training traffic monitoring and collection, deploy monitoring tools on each node (Server1 - ServerN) of the intelligent computing cluster to real-time monitor the traffic bandwidth of key components such as rdma switches, rdma network cards, pcieswitches, XPUswitches, or bus interfaces. The collected data includes traffic bandwidth and real input and output data of some iterations, and these data will be used for the construction of the subsequent historical traffic database.

[0066] Step S302, establish a collective communication traffic model: Analyze the collected data. Since the iterations of large model training are repetitive and similar, select representative iteration data for modeling. Construct a collective communication data model for a single iteration, including time nodes and input and output data, to form an accurate model. Use other iteration data to verify the model and adjust the model parameters until the error between the model and the real data reaches the preset range.

[0067] Step S303: Traffic model library generation: Repeat the above modeling and verification steps to construct the collective communication traffic models of multiple large models (such as llama, glm, telechat, etc.). Store these models in the traffic model library to provide data support for the subsequent test execution module.

[0068] Step S304: Simulator generates test cases: Utilize the collective communication data of the large model set in the traffic database, and simulate the real traffic load through a load simulator. Form the collective communication test cases for different models to prepare for the subsequent test execution.

[0069] Step S305: Intelligent algorithm selection automatically senses the cluster network topology, including the bus bandwidth between xpu cards and PCIe, as well as the bandwidth and topology of the RDAM network. Evaluate the performance of different collective communication algorithms (such as tree or ring) in the current scenario. According to the evaluation results, select the optimal collective communication algorithm to maximize the collective communication performance of the cluster.

[0070] Step S306: Adaptive multi-architecture execution: According to different chip clusters, select the appropriate adapted container image to execute the test. Adopt the client / server (C / S) architecture, and the test program can receive the task execution commands of the test platform, avoiding being bound to specific hardware.

[0071] Step S307: Real-time monitoring and data collection: Real-time monitor the traffic bandwidth data of key components, including rdma switches, rdma network cards, pcie switches, XPU switches or bus interfaces, etc. Collect the collective communication traffic data and calculate the total collective communication bandwidth to provide a basis for performance evaluation and optimization.

[0072] Through the above steps, it is possible to implement a traffic database for large model collective communication, simulate real traffic loads, form collective communication test cases for different models with a load simulator, and then accurately test the performance of each communication algorithm in actual application scenarios, solving the problems existing in the current collective communication test, such as the lack of real use cases, the difficulty in selecting test algorithms, and the lack of multi-architecture support. By simulating the generation of test cases for real cluster collective communication, the topology-aware intelligent algorithm selection mechanism, and cross-architecture cluster support, the accuracy, efficiency, and universality of the test are improved. Comprehensively and flexibly evaluate the performance and stability of the collective communication of the intelligent computing cluster. Can more accurately reflect the real performance of the cluster and better guide the in-network services. Specifically, the method embodiment of the present application has the following advantages:

[0073] 1. Real traffic simulation: By collecting and establishing the collective communication data of actual large model training, constructing a traffic model library of multiple large models, different model test cases can be simulated and generated, enabling the test to reflect real services.

[0074] 2. Intelligent algorithm selection: Based on the current network topology and bandwidth of the cluster and the execution status of test cases, select the optimal collective communication algorithm of the tree or ring type, so that the test can reflect the real cluster performance.

[0075] 3. Adaptation to multiple architectures: The collective communication test can be adaptively executed for multiple different clusters, making the test flexible and general.

[0076] The method embodiments of this application can be used in scenarios such as acceptance or stress testing of various heterogeneous intelligent computing centers, collective communication performance testing of intelligent computing clusters, network fault location and performance analysis of intelligent computing clusters, etc. It can obtain the accurate and real network performance of the cluster, optimize the network performance, and improve the computing power utilization rate of the cluster; quickly locate network problems in the cluster, shorten the recovery time, and improve the availability of the cluster; through automatic testing with multi-architecture adaptation, shorten the testing time and reduce costs.

[0077] The embodiments of this application provide a communication algorithm selection device. Figure 4 It is a schematic structural diagram of the device, as Figure 4 shown. The device includes: a generation module 40, configured to generate a traffic test case corresponding to the computing cluster according to the configuration information of the computing cluster, where the configuration information includes the topology structure information of the computing cluster, and the traffic test case includes input traffic size information and output traffic size information corresponding to one or more interfaces in the topology structure; a configuration module 42, configured to configure candidate communication algorithms in the computing cluster, where the candidate communication algorithms are used to instruct one or more interfaces to communicate according to preset rules; an execution module 44, configured to execute the traffic test case in the computing cluster and obtain the execution duration of each candidate communication algorithm when executing the traffic test case; a selection module 46, configured to select a target communication algorithm from the candidate communication algorithms according to the execution duration of each candidate communication algorithm, where the execution duration corresponding to the target communication algorithm is less than or equal to the execution duration corresponding to the candidate communication algorithm.

[0078] In some embodiments of the present application, after the selection module 46 selects a target communication algorithm from the candidate communication algorithms according to the execution duration of each candidate communication algorithm, the following steps are further included: applying the target communication algorithm in the computing cluster; monitoring the real-time input traffic size information and the real-time output traffic size information of the interfaces in the computing cluster; comparing the real-time input traffic size information of the interfaces with the input traffic size information corresponding to the interfaces in the test case, and determining that there is a potential fault in the topology structure corresponding to the interfaces when it is detected that the difference between the real-time input traffic size information and the input traffic size information is greater than a first preset threshold; comparing the real-time output traffic size information of the interfaces with the output traffic size information corresponding to the interfaces in the test case, and determining that there is a potential fault in the topology structure corresponding to the interfaces when it is detected that the difference between the real-time output traffic size information and the output traffic size information is greater than a second preset threshold; and sending an alarm message when it is determined that there is a potential fault in the topology structure corresponding to the interfaces, where the alarm message is used to indicate checking the topology structure corresponding to the interfaces.

[0079] In some embodiments of the present application, the configuration information of the computing cluster further includes the identification information of the large model deployed in the computing cluster. The generating module 40 generates a traffic test case corresponding to the computing cluster according to the configuration information of the computing cluster, including: determining historical traffic data corresponding to the identification information in the historical traffic database, where the historical traffic data includes parallel strategies, and the input traffic size information and the output traffic size information corresponding one-to-one to the parallel strategies, and the parallel strategies include at least one of the following: data parallelism, tensor parallelism, and pipeline parallelism; determining the parallel strategy followed when communicating between each interface according to the configuration information of the computing cluster; and generating a traffic test case according to the parallel strategy followed when communicating between each interface and the historical traffic data.

[0080] In some embodiments of the present application, before the generating module 40 determines the historical traffic data corresponding to the identification information in the traffic library, the following steps are further included: obtaining the first communication data of the large model deployed in the historical computing cluster in one iteration, where the first communication data includes the first input traffic size information, the first output traffic size information corresponding to one or more interfaces in the historical computing cluster, and the traffic transmission paths of one or more interfaces; determining the parallel strategy followed by the first input traffic size information and the first output traffic size information of one or more interfaces according to the first communication data and the topology structure of the historical computing cluster; counting the first input traffic size information and the first output traffic size information corresponding to each parallel strategy; and storing the identification information of the large model, each parallel strategy, the first input traffic size information corresponding to the parallel strategy, and the first output traffic size information corresponding to the parallel strategy in the historical traffic database.

[0081] In some embodiments of the present application, before determining the parallel strategy followed by the first input traffic size information and the first output traffic size information of one or more interfaces according to the first communication data and the topology of the historical computing cluster, the following steps are further included: obtaining multi-round second communication data of the large model in multiple iterations; determining the average value of the multi-round second communication data as the third communication data; comparing the third communication data with the first communication data to determine the difference in the input and output traffic sizes corresponding to one or more interfaces; in the case where the difference in the input and output traffic sizes corresponding to the interfaces in the third communication data and the first communication data is greater than or equal to the third preset threshold, re-obtaining the first communication data and comparing the third communication data with the new first communication data until the difference in the input and output traffic sizes corresponding to one or more interfaces in the third communication data and the new first communication data is less than the third preset threshold.

[0082] In some embodiments of the present application, before the execution module 44 executes the traffic test case in the computing cluster, the following steps are further included: determining the type of the graphics processor in the computing cluster; determining the corresponding mirror data according to the type of the graphics processor, where the mirror data is compatible with the graphics processor, and the mirror data includes a test program client for receiving the test case and executing the test case to send the mirror data to the computing cluster.

[0083] In some embodiments of the present application, the execution module 44 determining the corresponding mirror data according to the type of the graphics processor includes: querying a compatibility list according to the type of the graphics processor, where the compatibility list records each mirror data identifier and the compatible type corresponding to the mirror data identifier; in the case where the compatible type of the mirror data includes the type of the graphics processor, determining the mirror data as the mirror data corresponding to the type of the graphics processor.

[0084] It should be noted that each module in the above communication algorithm selection device may be a program module (for example, a set of program instructions for implementing a specific function), or a hardware module. For the latter, it may be presented in the following forms, but is not limited thereto: the presentation form of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.

[0085] An embodiment of the present application provides a non-volatile storage medium, in which a program is stored. When the program runs, it controls the device where the non-volatile storage medium is located to execute the following communication algorithm selection method: generating a traffic test case corresponding to the computing cluster according to the configuration information of the computing cluster, where the configuration information includes the topology structure information of the computing cluster, and the traffic test case includes the input traffic size information and the output traffic size information corresponding to one or more interfaces in the topology structure; configuring candidate communication algorithms in the computing cluster, where the candidate communication algorithms are used to instruct one or more interfaces to communicate according to preset rules; executing the traffic test case in the computing cluster, and obtaining the execution time used by each candidate communication algorithm when executing the traffic test case; selecting a target communication algorithm from the candidate communication algorithms according to the execution time used by each candidate communication algorithm, where the execution time used by the target communication algorithm is less than or equal to the execution time used by the candidate communication algorithm.

[0086] An embodiment of the present application provides an electronic device, including: a memory and a processor, where the processor is used to run the program stored in the memory. When the program runs, it executes the following communication algorithm selection method: generating a traffic test case corresponding to the computing cluster according to the configuration information of the computing cluster, where the configuration information includes the topology structure information of the computing cluster, and the traffic test case includes the input traffic size information and the output traffic size information corresponding to one or more interfaces in the topology structure; configuring candidate communication algorithms in the computing cluster, where the candidate communication algorithms are used to instruct one or more interfaces to communicate according to preset rules; executing the traffic test case in the computing cluster, and obtaining the execution time used by each candidate communication algorithm when executing the traffic test case; selecting a target communication algorithm from the candidate communication algorithms according to the execution time used by each candidate communication algorithm, where the execution time used by the target communication algorithm is less than or equal to the execution time used by the candidate communication algorithm.

[0087] An embodiment of the present application provides a computer program product, including a computer program, which implements the following communication algorithm selection method when executed by a processor: generating a traffic test case corresponding to a computing cluster according to the configuration information of the computing cluster, where the configuration information includes the topology structure information of the computing cluster, and the traffic test case includes input traffic size information and output traffic size information corresponding to one or more interfaces in the topology structure; configuring candidate communication algorithms in the computing cluster, where the candidate communication algorithms are used to instruct one or more interfaces to communicate according to preset rules; executing the traffic test case in the computing cluster, and obtaining the execution time used by each candidate communication algorithm when executing the traffic test case; selecting a target communication algorithm from the candidate communication algorithms according to the execution time used by each candidate communication algorithm, where the execution time used by the target communication algorithm is less than or equal to the execution time used by the candidate communication algorithm.

[0088] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not elaborated in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0089] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0090] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0091] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0092] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.

[0093] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A communication algorithm selection method, characterized in that, Including: Generating a traffic test case corresponding to the computing cluster according to the configuration information of the computing cluster, where the configuration information includes the topology structure information of the computing cluster, and the traffic test case includes the input traffic size information and the output traffic size information corresponding to one or more interfaces in the topology structure; Configuring candidate communication algorithms in the computing cluster, where the candidate communication algorithms are used to instruct the one or more interfaces to communicate according to preset rules; Executing the traffic test case in the computing cluster, and obtaining the execution time used by each of the candidate communication algorithms when executing the traffic test case; Selecting a target communication algorithm from the candidate communication algorithms according to the execution time used by each of the candidate communication algorithms, where the execution time used by the target communication algorithm is less than or equal to the execution time used by the candidate communication algorithm; 2. The communication algorithm selection method according to claim 1, wherein After selecting a target communication algorithm from the candidate communication algorithms according to the execution time used by each of the candidate communication algorithms, the method further includes: Applying the target communication algorithm in the computing cluster; Monitoring the real-time input traffic size information and the real-time output traffic size information of the interfaces in the computing cluster; Comparing the real-time input traffic size information of the interface with the input traffic size information corresponding to the interface in the test case, and determining that there is a potential fault in the topology structure corresponding to the interface when it is detected that the difference between the real-time input traffic size information and the input traffic size information is greater than a first preset threshold; Comparing the real-time output traffic size information of the interface with the output traffic size information corresponding to the interface in the test case, and determining that there is a potential fault in the topology structure corresponding to the interface when it is detected that the difference between the real-time output traffic size information and the output traffic size information is greater than a second preset threshold; When it is determined that there is a potential fault in the topology structure corresponding to the interface, sending an alarm message, where the alarm message is used to instruct to check the topology structure corresponding to the interface; 3. The communication algorithm selection method according to claim 1, wherein The configuration information of the computing cluster further includes the identification information of the large model deployed in the computing cluster. Generating a traffic test case corresponding to the computing cluster according to the configuration information of the computing cluster includes: Determining historical traffic data corresponding to the identification information in the historical traffic database, where the historical traffic data includes parallel strategies, and the input traffic size information and the output traffic size information corresponding to the parallel strategies one by one, and the parallel strategies include at least one of the following: data parallelism, tensor parallelism, pipeline parallelism; Determining the parallel strategy followed when communicating between each interface according to the configuration information of the computing cluster; Generating a traffic test case according to the parallel strategy followed when communicating between each interface and the historical traffic data; 4. The communication algorithm selection method according to claim 3, characterized in that Before determining the historical traffic data corresponding to the identification information in the traffic library, the method further includes: Obtain the first communication data of the large model deployed in the historical computing cluster in one iteration, where the first communication data includes the first input traffic size information, the first output traffic size information corresponding to one or more interfaces in the historical computing cluster, and the traffic transmission paths of the one or more interfaces; Determine the parallel strategy followed by the first input traffic size information and the first output traffic size information of the one or more interfaces according to the first communication data and the topological structure of the historical computing cluster; Statistically analyze the first input traffic size information and the first output traffic size information corresponding to each parallel strategy; Store the identification information of the large model, each parallel strategy, the first input traffic size information corresponding to the parallel strategy, and the first output traffic size information corresponding to the parallel strategy into the historical traffic database.

5. The communication algorithm selection method according to claim 4, wherein Before determining the parallel strategy followed by the first input traffic size information and the first output traffic size information of the one or more interfaces according to the first communication data and the topological structure of the historical computing cluster, the method further includes: Obtain the multiple rounds of second communication data of the large model in multiple iterations; Determine the average value of the multiple rounds of second communication data as the third communication data; Compare the third communication data with the first communication data to determine the difference in the input and output traffic sizes corresponding to the one or more interfaces; In the case where the difference in the input and output traffic sizes corresponding to the interfaces in the third communication data and the first communication data is greater than or equal to the third preset threshold, re-obtain the first communication data and compare the third communication data with the new first communication data until the difference in the input and output traffic sizes corresponding to one or more interfaces in the third communication data and the new first communication data is less than the third preset threshold.

6. The communication algorithm selection method according to claim 1, characterized in that Before executing the traffic test case in the computing cluster, the method further includes: Determine the type of the graphics processor in the computing cluster; Determine the corresponding mirror data according to the type of the graphics processor, where the mirror data is compatible with the graphics processor, and the mirror data includes a test program client for receiving and executing the test case Send the mirror data to the computing cluster.

7. The communication algorithm selection method according to claim 6, wherein Determining the corresponding mirror data according to the type of the graphics processor includes: Query the compatibility list according to the type of the graphics processor, where the compatibility list records each mirror data identifier and the compatible type corresponding to the mirror data identifier; In the case where the compatible type of the mirror data includes the type of the graphics processor, determine the mirror data as the mirror data corresponding to the type of the graphics processor.

8. A communication algorithm selection device, characterized in that, Includes: A generation module for generating a traffic test case corresponding to the computing cluster according to the configuration information of the computing cluster, where the configuration information includes the topological structure information of the computing cluster, and the traffic test case includes the input traffic size information and the output traffic size information corresponding to one or more interfaces in the topological structure; A configuration module, configured to configure candidate communication algorithms in the computing cluster, where the candidate communication algorithms are used to instruct the one or more interfaces to communicate according to preset rules; An execution module, configured to execute the traffic test case in the computing cluster and obtain the execution duration of each candidate communication algorithm when executing the traffic test case; A selection module, configured to select a target communication algorithm from the candidate communication algorithms according to the execution duration of each candidate communication algorithm, where the execution duration corresponding to the target communication algorithm is less than or equal to the execution duration corresponding to the candidate communication algorithm.

9. A non-volatile storage medium, characterized in that, A program is stored in the non-volatile storage medium, and when the program runs, it controls the device where the non-volatile storage medium is located to execute the communication algorithm selection method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, Comprising: A memory and a processor, the processor is configured to run the program stored in the memory, and when the program runs, it executes the communication algorithm selection method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the communication algorithm selection method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Self-adaptive selection method and device of ensemble communication algorithm

    CN121727957A