Device communication method, device and medium
By analyzing topological connection information in a heterogeneous GPU environment and selecting the optimal communication strategy, the problem of low communication efficiency between heterogeneous GPUs is solved, and efficient resource utilization and robust execution of distributed computing tasks are achieved.
Patent Information
- Application Number
- CN202411732925.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-29
AI Technical Summary
In heterogeneous GPU environment, the communication efficiency of the existing distributed machine learning framework is low, resulting in performance bottlenecks and waste of resources, making it difficult to achieve efficient interconnection between heterogeneous GPUs.
By analyzing the topological connection information between multiple devices, multiple first communication strategies are determined, and the optimal second communication strategy is dynamically selected through communication performance evaluation to adapt to the characteristics and network conditions of heterogeneous devices.
It significantly improves communication efficiency and resource utilization in heterogeneous equipment environments, avoids efficiency losses caused by device performance differences or network bottlenecks, and ensures efficient execution of distributed computing tasks in complex heterogeneous systems.
Smart Images

Figure CN119211020B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning technology, and in particular to a device communication method, device and medium. Background Art
[0002] With the rapid development of artificial intelligence technology, especially the rise of large models, the demand for computing resources has exploded. The scale of these models, the scale of training data, and the number of graphics processing unit (GPU) resources required are all growing at an exponential rate. In some cases, thousands or even tens of thousands of GPUs are needed to meet training needs. However, in the current GPU cloud service and resource usage environment, these tens of thousands of GPUs may come from different manufacturers or belong to different product models of the same manufacturer, with different hardware architectures, that is, they are heterogeneous.
[0003] In this heterogeneous environment, existing distributed machine learning frameworks face huge challenges. Traditional distributed collaborative training methods often rely on homogeneous hardware and network environments, and their communication solutions are mainly optimized for GPUs with consistent performance and unified network topologies. When the performance, architecture or manufacturer of the devices are different, or the topology and bandwidth of the network are inconsistent, the efficiency of existing communication solutions will drop significantly, and may even fail to work properly. Especially when high-frequency communication such as gradient aggregation and model parameter synchronization is required across devices, this difference will lead to performance bottlenecks, which seriously restricts the scale and efficiency of distributed collaborative training. In addition, devices with different performance and configurations are difficult to work together, resulting in some resources in the cluster being used in isolation, and interconnection between heterogeneous GPUs is difficult to achieve, forming the so-called "computing power island". This not only wastes precious hardware resources, but also increases the cost and complexity of training. Summary of the invention
[0004] The present invention provides a device communication method, a device and a medium, which are used to solve the defect of low communication efficiency between heterogeneous devices in the related art.
[0005] The present invention provides a device communication method, comprising the following steps:
[0006] Determining a plurality of first communication strategies based on topological connection information between the plurality of devices;
[0007] Performing communication performance evaluation on various first communication strategies respectively to obtain communication performance of various first communication strategies, and selecting a second communication strategy from the plurality of first communication strategies based on the communication performance of various first communication strategies;
[0008] Based on the second communication strategy, the plurality of devices are controlled to communicate.
[0009] According to a device communication method provided by the present invention, the topology connection information includes a device connection relationship and a device communication bandwidth.
[0010] According to a device communication method provided by the present invention, the determining of multiple first communication strategies based on topological connection information between multiple devices includes:
[0011] Determine a bound network card for each of the devices based on the device connection relationship and the device communication bandwidth between the multiple devices, and the network cards in the device connection relationship;
[0012] Based on the bound network cards of the respective devices, a plurality of first communication strategies are determined.
[0013] According to a device communication method provided by the present invention, the first communication strategy includes a plurality of communication subgraphs executed sequentially, and the communication subgraphs include communication relationships between the plurality of devices.
[0014] According to a device communication method provided by the present invention, the communication performance evaluation of various first communication strategies is performed to obtain the communication performance of various first communication strategies, including:
[0015] For various first communication strategies, based on the device communication bandwidth corresponding to the communication relationship in each communication subgraph in the first communication strategy, the communication time of each communication subgraph is determined, and based on the communication time of each communication subgraph, the communication time of the first communication strategy is determined as the communication performance.
[0016] According to a device communication method provided by the present invention, when the multiple devices are used to execute multiple aggregate communication primitives, each of the aggregate communication primitives corresponds to multiple first communication strategies, and each of the aggregate communication primitives corresponds to one of the second communication strategies.
[0017] According to a device communication method provided by the present invention, there are at least two types of devices among the multiple devices.
[0018] The present invention also provides a device communication apparatus, comprising the following modules:
[0019] A first communication strategy determining unit, configured to determine a plurality of first communication strategies based on topological connection information between a plurality of devices;
[0020] a second communication strategy determining unit, configured to respectively evaluate the communication performance of various first communication strategies, obtain the communication performance of various first communication strategies, and select a second communication strategy from the plurality of first communication strategies based on the communication performance of various first communication strategies;
[0021] A communication unit is used to control the multiple devices to communicate based on the second communication strategy.
[0022] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-mentioned device communication methods is implemented.
[0023] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the device communication method described in any one of the above is implemented.
[0024] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the device communication method described above is implemented.
[0025] The device communication method, device and medium provided by the present invention generate multiple first communication strategies by analyzing the topological connection information between heterogeneous devices, and dynamically select the optimal second communication strategy through communication performance evaluation to adapt to the characteristics and network conditions of heterogeneous devices. During the communication execution process, the operation is strictly carried out in accordance with the selected second communication strategy, which not only effectively improves the communication efficiency, but also optimizes the coordinated allocation of resources between devices, avoiding efficiency losses caused by differences in device performance or network bottlenecks. This method significantly improves the communication adaptation capability and resource utilization efficiency in a heterogeneous device environment, and provides a strong guarantee for the efficient execution of distributed computing tasks in complex heterogeneous systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0027] Figure 1 It is a structural diagram of a computing system based on a heterogeneous GPU cluster in the related technology.
[0028] Figure 2 This is one of the flow charts of the device communication method provided by the present invention.
[0029] Figure 3 It is a schematic diagram of topological connection between multiple devices provided by the present invention.
[0030] Figure 4 It is a communication flow chart based on the Reduce-Scatter communication strategy provided by the present invention.
[0031] Figure 5 It is a communication flow chart based on the AllGather communication strategy provided by the present invention.
[0032] Figure 6 This is the second flow chart of the device communication method provided by the present invention.
[0033] Figure 7 It is a structural schematic diagram of the equipment communication device provided by the present invention.
[0034] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0036] As the parameter scale and training data volume of large language models increase, the demand for storage resources and hardware computing power has exceeded the upper limit that a single computing device can provide. In order to break through this limitation and improve the training throughput at the same time, a 3D parallel training strategy is often adopted, that is, distributed collaborative training of the model is implemented through data parallelism, tensor parallelism, and pipeline parallelism. Here, distributed collaborative training refers to the use of collaborative cooperation between multiple computing devices to jointly participate in the training process of the model to achieve model training acceleration and performance improvement. It should be noted that the computing device here is an artificial intelligence chip, such as a GPU, TPU (Tensor Processing Unit), NPU (Neural network Processing Unit), DPU (Deep learning Processing Unit), APU (Accelerated Processing Unit), and GPGPU (General-Purpose computing on Graphics Processing Unit), etc., and the embodiment of the present invention does not specifically limit this.
[0037] In distributed collaborative training, thousands or even tens of thousands of GPUs may be used to meet training needs. In the current GPU cloud service and resource usage environment, these tens of thousands of GPUs may come from different manufacturers or belong to different product models of the same manufacturer, with different hardware architectures, that is, they are heterogeneous. In the embodiment of the present invention, a cluster composed of GPUs with different hardware architectures is called a heterogeneous GPU cluster.
[0038] For example, Figure 1 It is a schematic diagram of the structure of a computing system based on a heterogeneous GPU cluster in the related technology. Figure 1 As shown, the system may include a host 110 and a device end 120, wherein the device end 120 may include multiple first-type devices, multiple second-type devices, and other types of devices, where the devices can all be regarded as GPUs or other types of artificial intelligence chips. The host 110 is connected to multiple devices on the device end 120 to control these devices to perform various computing tasks and jointly complete the training of the model. On the device end 120, multiple first-type devices are all devices of the same manufacturer and the same model, and these devices constitute a homogeneous device cluster, while the first-type devices and the second-type devices may be devices of different manufacturers or different models, and these devices constitute a heterogeneous device cluster. Model heterogeneous mixed training (also known as "heterogeneous training") is to use these different types of devices (including homogeneous and heterogeneous devices) to participate in the training task of the model during the model training process. These devices may have different computing capabilities, memory sizes, data formats, and other characteristics, but they can work together to complete the training of the model through the control and coordination of the host.
[0039] In this heterogeneous environment, existing distributed machine learning frameworks face severe challenges. Traditional distributed collaborative training methods rely on homogeneous hardware and a unified network environment, and their core communication algorithms (such as RingAllReduce and Tree AllReduce) are mostly designed and optimized for GPUs with consistent performance. When there are significant differences in performance, architecture, or network characteristics between devices, the efficiency of these communication schemes drops sharply, and may even cause the system to fail to operate normally. Especially in high-frequency communication tasks such as gradient aggregation and model parameter synchronization across devices, these differences can cause serious performance bottlenecks, greatly limiting the scalability and efficiency of distributed collaborative training. In addition, the lack of an effective coordination mechanism between devices with different performance and configurations leads to the inability to fully utilize resources in the computing cluster, forming a "computing island" phenomenon. The potential of high-performance devices is wasted, while low-performance devices may become a bottleneck for the overall training speed. This problem of low resource utilization not only increases the cost of training, but also makes the deployment and optimization of large-scale models more complicated.
[0040] In view of the above problems, an embodiment of the present invention provides a device communication method. Figure 2 It is one of the flow charts of the device communication method provided by the present invention, such as Figure 2 As shown, the method includes:
[0041] Step 210: Determine a plurality of first communication strategies based on topological connection information between a plurality of devices.
[0042] Specifically, this device communication method is applied to the communication between multiple devices. The devices here can be artificial intelligence chips, such as GPU, TPU, NPU, APU, and GPGPU, etc., which are not specifically limited in the embodiments of the present invention. In addition, the multiple devices here can come from different manufacturers or belong to different product models of the same manufacturer, that is, the multiple devices here can have different hardware architectures, and the multiple devices can be heterogeneous.
[0043] Topological connection information refers to the physical and logical connection characteristics between various computing devices in distributed computing, such as multiple GPUs. This information may include communication bandwidth, communication latency, network card type, network protocol support (such as Ethernet, InfiniBand, RDMA, etc.), and topological structure (such as star, ring, or tree). Topological connection information describes the communication capabilities and constraints between devices and is the basis for selecting efficient communication strategies.
[0044] After obtaining the topological connection information between multiple devices, multiple first communication strategies can be determined based on this. Here, the first communication strategies are all strategies for realizing communication between multiple devices under the constraints of the topological connection information between the above-mentioned multiple devices. It can be understood that under the topological connection information, the connection paths between different devices can be one or multiple, so the first communication strategies formed under the topological connection information can also be multiple, and the communication paths between multiple devices under different first communication strategies can be different. Here, each first communication strategy can be used as a strategy for communication between multiple devices, and each first communication strategy has its own advantages and disadvantages in the communication performance reflected when realizing communication between multiple devices.
[0045] Step 220: perform communication performance evaluation on various first communication strategies respectively to obtain the communication performance of various first communication strategies, and select a second communication strategy from the plurality of first communication strategies based on the communication performance of various first communication strategies.
[0046] Specifically, communication performance evaluation is a process of comprehensively analyzing the performance of different communication strategies in a specific network environment, aiming to provide data support for selecting the optimal communication strategy in a distributed collaborative training environment. The core of communication performance evaluation is to understand the advantages and disadvantages of each communication strategy in actual operation by measuring and analyzing a variety of communication performance indicators, and select the optimal strategy according to specific needs. These communication performance indicators may include but are not limited to: communication delay, bandwidth utilization, throughput, hardware resource consumption, scalability, etc. Delay is one of the important evaluation indicators. It indicates the time required for data to be sent and received, which is a direct reflection of communication efficiency. Bandwidth utilization reflects the degree of utilization of network resources during communication. The higher the utilization, the higher the efficiency of the communication strategy. Throughput is used to measure the ability to process data per unit time, and is an indicator that must be considered in large-scale data transmission tasks. Resource consumption focuses on the occupation of hardware resources (such as CPU, memory, network card, etc.) by communication, and scalability evaluates the performance adaptability of communication strategies when the scale of equipment changes.
[0047] The specific process of communication performance evaluation includes unified analysis of all first communication strategies based on the same indicators. Each communication strategy will be tested and compared for key indicators such as latency, bandwidth utilization, throughput, and resource consumption under the same network environment. For example, for the three first communication strategies, namely, first communication strategy 1, first communication strategy 2, and first communication strategy 3, their communication delays under the same data scale and network conditions will be measured respectively, and their bandwidth utilization efficiency and hardware resource occupancy rate will be uniformly evaluated. This evaluation method based on the same indicators ensures the fairness and accuracy of the comparison, and also facilitates finding the strategy with the best performance in a specific environment from multiple strategies. Through this comprehensive analysis, the advantages of various first communication strategies in different network environments can be discovered.
[0048] After completing the performance evaluation, the communication performance of all first communication strategies is compared, and the communication strategy with the best overall performance is selected as the second communication strategy for actual execution. For example, in tasks that require global synchronization, the strategy with the lowest latency can be selected; in scenarios where resources are tight, the strategy with the lowest resource consumption may be selected as the final solution. This decision-making method based on unified evaluation indicators ensures that the second communication strategy finally selected can maximize communication efficiency in the current network environment. This method greatly improves the efficiency of communication in heterogeneous distributed computing environments and ensures the smooth execution of complex computing tasks.
[0049] Step 230: Control the multiple devices to communicate based on the second communication strategy.
[0050] Specifically, after completing the communication performance evaluation of the first communication strategy and selecting the optimal second communication strategy, the communication behaviors among multiple devices will be coordinated according to the design logic of the second communication strategy. The second communication strategy is a comprehensive evaluation result combining topological connection information, device performance characteristics, and communication performance. Its core goal is to maximize communication efficiency and optimize resource utilization under the current network environment and task requirements.
[0051] The execution process of communication requires the second communication strategy to be concretized into actual operations between devices. For example, according to the data transmission path and order determined by the second communication strategy, the data sending and receiving operations are arranged in sequence to ensure that the communication process is completed efficiently and orderly. In this process, the role of each device in the communication, such as the data sender or receiver, will be clarified, and the load of the communication task will be reasonably distributed to avoid resource waste or efficiency reduction due to differences in device performance or network bottlenecks.
[0052] By implementing the second communication strategy, devices can achieve efficient collaboration, data can be transmitted in the shortest time, and bandwidth utilization and resource allocation efficiency can be maximized. This process not only improves communication efficiency, but also ensures the robustness and execution effect of distributed tasks in complex environments, thereby giving full play to the potential of computing resources.
[0053] The device communication method provided by the present invention generates multiple first communication strategies by analyzing the topological connection information between devices, and dynamically selects the optimal second communication strategy through communication performance evaluation to adapt to the characteristics and network conditions of heterogeneous devices. During the communication execution process, the operation is strictly carried out in accordance with the selected second communication strategy, which not only effectively improves the communication efficiency, but also optimizes the coordinated allocation of resources between devices, avoiding efficiency losses caused by differences in device performance or network bottlenecks. This method significantly improves the communication adaptation capability and resource utilization efficiency in a heterogeneous device environment, and provides a strong guarantee for the efficient execution of distributed computing tasks in complex heterogeneous systems.
[0054] Based on the above embodiment, the topology connection information includes device connection relationship and device communication bandwidth.
[0055] Device connection relationships refer to the physical or logical connection status between multiple devices in distributed computing. These connections can be achieved through network interfaces (such as network cards) or established through other hardware buses (such as PCIe or NVLink). In a multi-node distributed system, device connection relationships usually reflect the communication paths established between devices through network cards and network switches, such as the status of two GPUs connected to the same Ethernet switch or InfiniBand network through their network cards. In a single-node system, device connection relationships may be manifested as the connection between GPUs, CPUs, and storage devices through the PCIe bus on the motherboard or dedicated interconnects (such as NVLink). The characteristics of device connection relationships directly affect the complexity, latency, and potential communication bottlenecks of the data transmission path.
[0056] Device communication bandwidth refers to the amount of data that can be transmitted per unit time in the connection channel between two devices, usually measured in Mbps (million bits per second) or Gbps (billion bits per second). Communication bandwidth is a key indicator for evaluating the performance of connections between devices. High-bandwidth connections can support faster data synchronization and distribution, while low-bandwidth connections may limit data transmission efficiency and become a communication bottleneck.
[0057] By comprehensively analyzing the device connection relationship and device communication bandwidth, we can fully describe the communication capabilities and constraints between devices in distributed collaborative training. The device connection relationship determines the accessibility and path characteristics of communication between devices, while the device communication bandwidth measures the performance potential of these connections. Together, they provide a reliable data foundation for the design of communication strategies, enabling distributed collaborative training to efficiently utilize existing resources to execute tasks.
[0058] Based on the above embodiment, in step 220, the determining of multiple first communication strategies based on the topological connection information between the multiple devices includes:
[0059] Determine a bound network card for each of the devices based on the device connection relationship and the device communication bandwidth between the multiple devices, and the network cards in the device connection relationship;
[0060] Based on the bound network cards of the respective devices, a plurality of first communication strategies are determined.
[0061] Specifically, the device connection relationship not only includes the physical connection path between devices, but also reflects the specific way in which the device interacts with other devices through hardware interfaces such as network cards. In distributed computing, each device may be equipped with one or more network cards to support different types of communication protocols. By analyzing the device connection relationship and device communication bandwidth, the network card resources actually used by the device can be accurately identified, thus providing a basis for binding the device to the most suitable network card.
[0062] The bound network card of a device is a network card bound to the device that is selected based on the connection performance and communication requirements between devices. For example, on a device with multiple network cards, by analyzing the bandwidth, latency, and load of each network card, the network card with the best communication performance is selected as the bound network card to support subsequent communication operations. If some devices are connected through multiple paths, it is necessary to determine the optimal binding combination by evaluating the comprehensive performance of each network card in the path. The result of binding network cards ensures that each device can maximize the use of its network card resources during the communication process, while avoiding uneven use of network resources.
[0063] For example, Figure 3 Schematic diagram of topological connection between multiple devices provided by the present invention. Figure 3 As shown, multiple devices can form two systems, SYSTEM0 and SYSTEM1. SYSTEM0 includes 8 GPU devices, GPU0-GPU7, and SYSTEM1 also includes 8 GPU devices, GPU0-GPU7. For GPU0 in SYSTEM0, GPU0 can be connected to network cards NIC1 and NIC2 through PCIe switch SWITCH0, and connected to network card NIC0 through NODE0. Therefore, GPU0 can select a network card from NIC1, NIC2, and NIC0 as a bound network card based on the topological connection relationship.
[0064] After determining the bound network cards of each device, multiple first communication strategies can be determined in combination with the bound network cards of each device, the device connection relationship, the device communication bandwidth, etc.
[0065] The first communication strategy generated through this process can make full use of the device connection relationship and device communication bandwidth between devices, and effectively adapt to the diversity of heterogeneous devices. The diversity of the first communication strategy provides flexibility for subsequent performance evaluation, and can select the optimal solution under different network environments and task requirements, avoiding resource waste and performance bottlenecks caused by fixed communication modes.
[0066] Based on the above embodiment, the first communication strategy includes a plurality of communication subgraphs executed sequentially, and the communication subgraphs include the communication relationship between the plurality of devices.
[0067] Specifically, in heterogeneous systems, due to the differences in performance and connectivity between devices, it is difficult to directly adopt unified large-scale communication operations. Therefore, it is necessary to split the overall communication task into multiple smaller communication subgraphs and complete unified operations with a combination of these subgraphs.
[0068] The design of the splitting strategy for the communication subgraph needs to consider not only the algorithm requirements and the connection relationship between devices, but also the performance estimation and analysis of different splitting strategies. Algorithmic requirements determine the logical structure of the task, such as whether global synchronization or local aggregation is required in distributed training; the connection relationship between devices affects the specific structure of the subgraph, including the bandwidth, latency, and load balancing between devices. However, different splitting strategies may lead to significant differences in performance. Therefore, when creating a subgraph, it is necessary to perform a performance evaluation on each possible splitting scheme to ensure that the selected splitting strategy can maximize communication efficiency. Performance evaluation needs to take into account factors such as the hardware characteristics of the device, network conditions, and the amount of task data.
[0069] After determining the splitting strategy with the best performance, the communication subgraph is created to clarify the set of devices involved in the communication, the communication relationship between the devices, and the operations that need to be performed at each stage. The communication subgraphs are executed in sequence according to the established order. Each subgraph completes an independent communication task and collaborates with other subgraphs to complete the global task. For example, when calling AllReduce on multiple devices, you can first select a locally synchronized subgraph, complete gradient aggregation within the device group, and then complete the global aggregation operation through a cross-group synchronized subgraph.
[0070] For example, Figure 4 : is a communication flow chart based on the Reduce-Scatter communication strategy provided by the present invention. Figure 4 As shown, in the embodiment of the present invention, Reduce-Scatter is split into three sequentially executed communication subgraphs, namely, Figure 4 The first step communication subgraph, the second step communication subgraph and the third step communication subgraph are shown in .
[0071] exist Figure 4 In the first device type, there are two GPUs, G10 and G11, and the second device type includes six GPUs, G20, G21, G22, G23, G24, and G25. In this communication strategy, the Recursive Halving algorithm is used to complete the communication. In the first communication subgraph, the device groups with a relative distance of 4 communicate (for example, G10 and G22, G11 and G23), and each pair of GPUs exchanges 1 / 2 of their data, which are then aggregated to the corresponding device. In the second communication subgraph, the GPU groups with a relative distance of 2 communicate (for example, G10 and G20, G11 and G21), and each pair of GPUs exchanges 1 / 4 of their data based on the previous round of data and aggregates. In the third communication subgraph, the GPU groups with a relative distance of 1 communicate (for example, G10 and G11, G20 and G21), and each pair of GPUs exchanges 1 / 8 of their data and aggregates. By sequentially executing the three communication subgraphs, the Reduce-Scatter communication between the above two types of GPUs can be realized.
[0072] Through the optimization of this splitting strategy, the communication subgraph can fully adapt to the device characteristics and network conditions in the heterogeneous environment, which not only improves the communication efficiency but also avoids the waste of resources. The overall task is split into multiple performance-optimized subgraphs and executed sequentially, ensuring the efficient completion of complex tasks in a variety of scenarios, providing a flexible and efficient solution for distributed communication tasks.
[0073] Based on the above embodiment, the performing communication performance evaluation on various first communication strategies to obtain the communication performance of various first communication strategies includes:
[0074] For various first communication strategies, based on the device communication bandwidth corresponding to the communication relationship in each communication subgraph in the first communication strategy, the communication time of each communication subgraph is determined, and based on the communication time of each communication subgraph, the communication time of the first communication strategy is determined as the communication performance.
[0075] Specifically, the core of communication performance is communication time, which is directly affected by the communication bandwidth of the device. The communication bandwidth defines the amount of data that can be transmitted per unit time. The higher the bandwidth, the higher the transmission efficiency and the shorter the communication time. For example, for a given amount of data S , communication time T It can be expressed as T = S / B ,in B is the communication bandwidth. The communication subgraph is the specific execution unit of the communication strategy. The communication time of each subgraph reflects the efficiency of data exchange between devices. Therefore, the cumulative communication time of all subgraphs can represent the communication performance of the entire first communication strategy. The purpose of communication performance evaluation is to find the communication strategy with the best performance in the current environment by analyzing the time consumption of each subgraph in different strategies.
[0076] For example, Figure 4 The execution steps of the entire Reduce-Scatter are log2 P (in P is the number of GPUs), that is, the amount of data for each communication is reduced by half. Each GPU needs to establish log2 P The total transmission time is:
[0077]
[0078] That is, the communication performance of the Reduce-Scatter communication strategy is The communication time is consumed.
[0079] For example, Figure 5This is a communication flow chart based on the AllGather communication strategy provided by the present invention. Figure 5 As shown, similar to Figure 4 , distributed collaborative training includes two device types. The first device type includes two GPUs, G10 and G11, and the second device type includes six GPUs, G20, G21, G22, G23, G24, and G25. In this communication strategy, the Recursive doubling algorithm is used to complete the communication. In the first step of the communication subgraph, the device group with a relative distance of 1 communicates (for example, G10 and G11, G20 and G21), and each pair of GPUs exchanges 1 / 8 of their data, which are then aggregated to the corresponding device. In the second step of the communication subgraph, the GPU group with a relative distance of 2 communicates (for example, G10 and G20, G11 and G21), and each pair of GPUs exchanges 1 / 4 of their data based on the previous round of data and aggregates. In the third step of the communication subgraph, the GPU group with a relative distance of 4 communicates (for example, G10 and G22, G11 and G23), and each pair of GPUs exchanges 1 / 2 of their data and aggregates. The execution steps of the entire AllGather communication strategy are log2 P (in P = number of GPUs), that is, the amount of data is increased by half for each communication. Each GPU needs to establish log2 P The total transmission time is:
[0080]
[0081] That is, the communication performance of the AllGather communication strategy is The communication time is consumed.
[0082] Based on the above embodiment, when the multiple devices are used to execute multiple aggregate communication primitives, each of the aggregate communication primitives corresponds to multiple first communication strategies, and each of the aggregate communication primitives corresponds to one of the second communication strategies.
[0083] Aggregate communication primitives are basic operations or protocols used for communication and data aggregation in distributed computing, such as AllReduce, Reduce-Scatter, and AllGather. Among them, AllReduce is used to aggregate data between all devices, such as gradient aggregation in distributed collaborative training; Reduce-Scatter is an efficient data distribution and partial aggregation strategy, suitable for scenarios where each device only needs to pay attention to a part of the data; AllGather is used to share the data of each device to all other devices, suitable for scenarios that require global synchronization of data.
[0084] Due to the diversity of device hardware performance, device connection conditions and specific task requirements, different aggregate communication primitives need to use different communication strategies in the same distributed system to adapt to specific network environments and device configurations. The multiple first communication strategies of each aggregate communication primitive are multiple groups of alternatives designed for its specific functions, and the second communication strategy finally adopted is the best solution selected from these alternatives through performance evaluation. For example, when executing the Reduce-Scatter aggregate communication primitive, multiple first communication strategies may be included, such as recursive splitting strategy, hierarchical communication strategy and direct distribution strategy. The recursive splitting strategy achieves efficient synchronization by gradually reducing the amount of data, which is suitable for a uniform network environment; the hierarchical communication strategy first performs intra-group synchronization and then inter-group communication, which performs better in heterogeneous networks; the direct distribution strategy is suitable for small-scale, low-latency systems. By evaluating the performance of these strategies, such as analyzing performance evaluation indicators such as communication time consumption and bandwidth utilization, the best second communication strategy can be selected. Similarly, when executing the AllGather communication primitive, multiple first communication strategies may also be included, such as recursive doubling strategy, group broadcast strategy and ring transmission strategy. The recursive doubling strategy achieves global synchronization by gradually expanding the communication range, which is suitable for networks with uniform bandwidth; the group broadcast strategy optimizes bandwidth utilization by combining intra-group and inter-group broadcasts; the ring transmission strategy achieves synchronization by gradually transmitting data, which is suitable for multi-node scenarios. The second communication strategy selected through performance evaluation can adapt to the current environment and maximize communication efficiency. Under this mechanism, the multiple first communication strategies corresponding to each aggregate communication primitive provide a flexible selection space, and the communication performance evaluation ensures that each aggregate communication primitive can adopt the optimal execution plan. This method improves the adaptability of distributed collaborative training in heterogeneous environments, significantly reduces communication overhead, and ensures the efficient completion of complex tasks.
[0085] Based on the above embodiment, there are at least two types of devices in the plurality of devices.
[0086] Specifically, two types of devices refer to devices that have significant differences in certain key attributes or functions, mainly including differences between hardware types and different models or configurations of the same hardware category. For example, different types of devices may be GPUs produced by different manufacturers, or different models of GPUs produced by the same manufacturer. This feature of at least two types of devices is one of the core characteristics of heterogeneous environments, which puts higher requirements on the selection of communication strategies and also provides more possibilities for task optimization.
[0087] Based on the above embodiments, Figure 6 This is the second flow chart of the device communication method provided by the present invention. Figure 6 As shown, the method of heterogeneous device aggregation communication is mainly divided into three stages.
[0088] The first is the topology awareness stage. By detecting the device connection relationship and bus bandwidth between devices, the device topology information is established. The device topology information includes the device's communication bandwidth, latency, and bound network card, etc., which provides a basis for the subsequent communication strategy design. By analyzing this information, the performance differences between devices and the complexity of the network can be effectively identified to ensure that the communication strategy can adapt to the heterogeneous environment.
[0089] Next is the subgraph splitting stage, where multiple communication strategies are generated based on the topology information. For example, for different aggregate communication primitives such as AllReduce, ReduceScatte, and AllGather, different first communication strategies can be generated for each aggregate communication primitive. For example, for AllReduce, AllReduce communication strategies one, two, and three can be generated. AllReduce communication strategies one, two, and three refer to different first communication strategies generated for AllReduce. Optimize the design for different task requirements. The communication strategy will undergo performance evaluation, calculate indicators such as communication time consumption and bandwidth utilization, and select the best performing solution, that is, obtain the second communication strategy corresponding to each aggregate communication primitive. Subsequently, the second communication strategy under the global communication task can be split into several subgraphs to represent it. Each subgraph corresponds to the communication relationship and task logic of a specific device. For example, AllReduce uses a recursive halving and doubling algorithm to achieve task decomposition. Subgraph splitting not only considers the topological relationship between devices, but also combines the heterogeneous characteristics of devices to optimize communication efficiency.
[0090] Finally, in the communication group coordination phase, a communication group is created based on the generated subgraph to clarify the communication path between devices. The communication task executes the communication strategy according to the optimized communication group. For example, Reduce-Scatter gradually reduces the amount of data by recursive halving, and AllGather completes data synchronization by recursive doubling. The communication group optimizes the device load distribution and bandwidth usage to avoid communication bottlenecks caused by differences in device performance.
[0091] The overall process improves the communication efficiency between heterogeneous devices through topology awareness, subgraph splitting and communication group collaboration, solves the problem of traditional communication methods' dependence on homogeneous devices, and can efficiently complete distributed training tasks and optimize resource utilization.
[0092] The device communication apparatus provided by the present invention is described below. The device communication apparatus described below and the device communication method described above can be referenced to each other. Figure 7 It is a schematic diagram of the structure of the device communication device provided by the present invention, such as Figure 7 As shown, the device comprises:
[0093] The first communication strategy determining unit 710 is configured to determine a plurality of first communication strategies based on topological connection information between a plurality of devices.
[0094] The second communication strategy determination unit 720 is used to evaluate the communication performance of various first communication strategies respectively, obtain the communication performance of various first communication strategies, and select a second communication strategy from the multiple first communication strategies based on the communication performance of various first communication strategies.
[0095] The communication unit 730 is configured to control the plurality of devices to communicate based on the second communication strategy.
[0096] The device communication device provided by the present invention generates multiple first communication strategies by analyzing the topological connection information between heterogeneous devices, and dynamically selects the optimal second communication strategy through communication performance evaluation to adapt to the characteristics and network conditions of heterogeneous devices. During the communication execution process, the operation is strictly carried out in accordance with the selected second communication strategy, which not only effectively improves the communication efficiency, but also optimizes the coordinated allocation of resources between devices, avoiding efficiency losses caused by differences in device performance or network bottlenecks. This method significantly improves the communication adaptation capability and resource utilization efficiency in a heterogeneous device environment, and provides a strong guarantee for the efficient execution of distributed computing tasks in complex heterogeneous systems.
[0097] Based on any of the foregoing embodiments, the first communication strategy determination unit is specifically configured to:
[0098] The topology connection information includes device connection relationship and device communication bandwidth.
[0099] There are at least two types of devices among the plurality of devices.
[0100] Determine a bound network card for each of the devices based on the device connection relationship and the device communication bandwidth between the multiple devices, and the network cards in the device connection relationship;
[0101] Based on the bound network cards of the respective devices, a plurality of first communication strategies are determined.
[0102] The first communication strategy includes a plurality of communication subgraphs executed sequentially, and the communication subgraphs include communication relationships between the plurality of devices.
[0103] Based on any of the foregoing embodiments, the second communication strategy determination unit is specifically configured to:
[0104] For various first communication strategies, based on the device communication bandwidth corresponding to the communication relationship in each communication subgraph in the first communication strategy, the communication time of each communication subgraph is determined, and based on the communication time of each communication subgraph, the communication time of the first communication strategy is determined as the communication performance.
[0105] In the case where the plurality of devices are used to execute a plurality of aggregate communication primitives, each of the aggregate communication primitives corresponds to a plurality of the first communication strategies, and each of the aggregate communication primitives corresponds to a second communication strategy.
[0106] Figure 8 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the device communication method, which includes:
[0107] Determining a plurality of first communication strategies based on topological connection information between the plurality of devices;
[0108] Performing communication performance evaluation on various first communication strategies respectively to obtain communication performance of various first communication strategies, and selecting a second communication strategy from the plurality of first communication strategies based on the communication performance of various first communication strategies;
[0109] Based on the second communication strategy, the plurality of devices are controlled to communicate.
[0110] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the relevant technology or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0111] On the other hand, the present invention further provides a computer program product, the computer program product comprising a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, when the computer program is executed by a processor, the computer can execute the device communication method provided by the above methods, the method comprising:
[0112] Determining a plurality of first communication strategies based on topological connection information between the plurality of devices;
[0113] Performing communication performance evaluation on various first communication strategies respectively to obtain communication performance of various first communication strategies, and selecting a second communication strategy from the plurality of first communication strategies based on the communication performance of various first communication strategies;
[0114] Based on the second communication strategy, the plurality of devices are controlled to communicate.
[0115] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the device communication method provided by the above methods is implemented, and the method includes:
[0116] Determining a plurality of first communication strategies based on topological connection information between the plurality of devices;
[0117] Performing communication performance evaluation on various first communication strategies respectively to obtain communication performance of various first communication strategies, and selecting a second communication strategy from the plurality of first communication strategies based on the communication performance of various first communication strategies;
[0118] Based on the second communication strategy, the plurality of devices are controlled to communicate.
[0119] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0120] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A device communication method, characterized in that: include: Determine multiple first communication strategies based on topological connection information between multiple devices, where the device is an artificial intelligence chip, and the multiple devices are from different manufacturers or belong to different product models of the same manufacturer; Performing communication performance evaluation on various first communication strategies respectively to obtain communication performance of various first communication strategies, and selecting a second communication strategy from the multiple first communication strategies based on the communication performance of various first communication strategies, wherein the indicators of the communication performance evaluation include communication delay and hardware resource occupancy rate; Based on the second communication strategy, the plurality of devices are controlled to communicate.
2. The device communication method according to claim 1, characterized in that: The topology connection information includes device connection relationship and device communication bandwidth.
3. The device communication method according to claim 2, characterized in that: The determining of a plurality of first communication strategies based on the topological connection information between the plurality of devices includes: Determine a bound network card for each of the devices based on the device connection relationship and the device communication bandwidth between the multiple devices, and the network cards in the device connection relationship; Based on the bound network cards of the respective devices, a plurality of first communication strategies are determined.
4. The device communication method according to claim 1, characterized in that: The first communication strategy includes a plurality of communication subgraphs executed sequentially, and the communication subgraphs include communication relationships between the plurality of devices.
5. The device communication method according to claim 4, characterized in that: The performing communication performance evaluation on various first communication strategies to obtain the communication performance of various first communication strategies includes: For various first communication strategies, based on the device communication bandwidth corresponding to the communication relationship in each communication subgraph in the first communication strategy, the communication time of each communication subgraph is determined, and based on the communication time of each communication subgraph, the communication time of the first communication strategy is determined as the communication performance.
6. The device communication method according to any one of claims 1 to 5, characterized in that: In the case where the plurality of devices are used to execute a plurality of aggregate communication primitives, each of the aggregate communication primitives corresponds to a plurality of the first communication strategies, and each of the aggregate communication primitives corresponds to a second communication strategy.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the device communication method according to any one of claims 1 to 6 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the device communication method according to any one of claims 1 to 6 is implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the device communication method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Processing performance optimization method and device based on heterogeneous system
CN112256623A
Multi-GPU set communication path selection method based on ring algorithm
CN118282923A