Integrated communication method, device, storage medium, and computer program product

By constructing a logical topology in the computing system and prioritizing the use of different ring topology connections for data transmission, the problem of bandwidth resource waste in aggregated communication is solved, and the overall communication performance is improved.

WO2026098187A1PCT designated stage Publication Date: 2026-05-15CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
Filing Date
2025-10-16
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing ensemble communication methods are significantly affected by inter-node and intra-node bandwidth resources in distributed training, leading to a decline in overall communication performance and a waste of bandwidth resources.

Method used

The logical topology of the computing system is constructed, including a first ring topology between different computing nodes and a second ring topology within the same computing node. Data transmission is carried out by preferentially using different ring topology connections through two different transmission paths, dynamically adjusting the network load between machines and making full use of bandwidth resources.

Benefits of technology

It improves overall communication performance, reduces the waste of bandwidth resources, and maximizes the utilization of bandwidth resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025128124_15052026_PF_FP_ABST
    Figure CN2025128124_15052026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an integrated communication method, a device, a storage medium, and a computer program product. When a computing system comprises N computing nodes and the computing nodes comprise M computing units, a logical topology of the computing system is constructed, the logical topology comprising a first ring topology between N computing units pertaining to different computing nodes and a second ring topology between M computing units pertaining to a same computing node. On the basis of two different ring topologies, two different transmission paths can be constructed, one of the two different transmission paths preferentially using connection of the first ring topology for data transmission, and the other transmission path preferentially using connection of the second ring topology for data transmission. The number of data channels adopting two different transmission paths is controlled, and the corresponding transmission paths are used to perform data transmission on the data channels, so as to fully utilize bandwidth resources within the computing nodes and bandwidth resources between the computing nodes and achieve maximum utilization of the bandwidth resources.
Need to check novelty before this filing date? Find Prior Art

Description

Products that combine communication methods, devices, storage media, and computer programs

[0001] This disclosure claims priority to Chinese Patent Application No. 202411603876.5, filed on November 11, 2024, entitled "Collection Communication Method, Device, Storage Medium and Computer Program Product", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of communication technology, and in particular to a product that integrates communication methods, devices, storage media, and computer programs. Background Technology

[0003] As large language models (LLMs) continue to grow in size, training them requires increasingly more training data and computational resources. This training data and resources are often distributed across multiple computing nodes for distributed training. To improve the efficiency of distributed training, aggregated communication is employed. Aggregated communication is a communication mode that enables efficient data exchange between multiple computing nodes in a distributed computing environment. For example, parameter synchronization and gradient aggregation in distributed model training can both utilize aggregated communication.

[0004] Currently, aggregated communication commonly employs a ring topology, where multiple computing nodes form a ring communication path connected end-to-end via network interface cards (NICs). Data is transmitted along this ring path between computing nodes and between multiple computing units within a single node, enabling rapid data synchronization. However, the overall communication performance of this aggregated communication method is significantly affected by both inter-node and intra-node bandwidth resources. If either bandwidth resource becomes a bottleneck, it will lead to a decrease in overall communication performance and a waste of bandwidth resources. Summary of the Invention

[0005] This disclosure provides a combination of communication methods, devices, storage media, and computer program products to improve overall communication performance and reduce the waste of bandwidth resources.

[0006] This disclosure provides a method for aggregated communication applied to a computing system. The computing system includes N computing nodes, and each computing node includes M computing units, where M and N are both positive integers. The method includes: determining a logical topology built upon the physical topology of the computing system, the logical topology including a first ring topology between N computing units belonging to different computing nodes and a second ring topology between M computing units belonging to the same computing node; creating K1 data channels suitable for a first transmission path and K2 data channels suitable for a second transmission path on the logical topology based on the bandwidth resources between computing nodes and the bandwidth resources within computing nodes, where K1 and K2 are both positive integers; wherein the first transmission path is a path that preferentially uses connections in the second ring topology to transmit data between the M*N computing units, and the second transmission path is a path that preferentially uses connections in the first ring topology to transmit data between the M*N computing units; transmitting data between the M*N computing units using the first transmission path in the K1 data channels; and transmitting data between the M*N computing units using the second transmission path in the K2 data channels.

[0007] This disclosure also provides an electronic device, including: a memory and a processor; the memory stores a computer program; the processor is coupled to the memory and is used to execute the computer program to implement the steps in the collective communication method.

[0008] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the processor to implement the steps in the collection communication method.

[0009] This disclosure also provides a computer program product, including a computer program / instruction that, when executed by a processor, enables the processor to implement the steps in the collection communication method.

[0010] In this embodiment, for a computing system comprising N computing nodes, each computing node comprising M computing units, a logical topology is constructed. This logical topology includes a first ring topology between the N computing units belonging to different computing nodes and a second ring topology between the M computing units belonging to the same computing node. Based on these two different ring topologies, two different transmission paths can be constructed. One transmission path prioritizes the connection using the first ring topology for data transmission, while the other prioritizes the connection using the second ring topology. The number of data channels employing the two different transmission paths is controlled, and data transmission is performed on the corresponding transmission paths. This fully utilizes the bandwidth resources within and between computing nodes, maximizing bandwidth utilization, improving overall communication performance, and reducing bandwidth waste. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this disclosure, illustrate exemplary embodiments of the present disclosure and are used to explain the disclosure, but do not constitute an undue limitation of the disclosure. In the drawings:

[0012] Figure 1 is a schematic diagram of an exemplary ring topology;

[0013] Figure 2 is a flowchart of a collection communication method provided in an embodiment of this disclosure;

[0014] Figure 3 is a schematic diagram of the logical topology of an exemplary computing system;

[0015] Figure 4 is an exemplary path diagram for data transmission based on connections in a second ring topology in a first transmission path;

[0016] Figure 5 is an exemplary path diagram of data transmission based on connections in a first ring topology in a first transmission path;

[0017] Figure 6 is an exemplary path diagram for data transmission based on connections in the first ring topology in a second transmission path;

[0018] Figure 7 is an exemplary path diagram for data transmission based on connections in a second ring topology in a second transmission path;

[0019] Figure 8 is a schematic diagram of the structure of a collective communication device provided in an embodiment of this disclosure;

[0020] Figure 9 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0022] In the embodiments of this disclosure, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the access relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone, where A and B can be singular or plural. In the textual description of this disclosure, the character " / " generally indicates that the preceding and following associated objects have an "or" relationship. Furthermore, in the embodiments of this disclosure, "first," "second," "third," etc., are only used to distinguish the content of different objects and have no other special meaning.

[0023] As large language models (LLMs) continue to grow in size, training them requires increasingly more training data and computational resources. This training data and resources are often distributed across multiple computing nodes for distributed training. To improve the efficiency of distributed training, aggregated communication is employed. Aggregated communication is a communication mode that enables efficient data exchange between multiple computing nodes in a distributed computing environment. For example, parameter synchronization and gradient aggregation in distributed model training can both utilize aggregated communication.

[0024] Currently, aggregated communication commonly employs a ring topology, where multiple computing nodes form a ring communication path connected end-to-end via network interface cards (NICs). Data is transmitted along this ring path between computing nodes and between multiple computing units within a single node, enabling rapid data synchronization. However, the overall communication performance of this aggregated communication method is significantly affected by both inter-node and intra-node bandwidth resources. If either bandwidth resource becomes a bottleneck, it will lead to a decrease in overall communication performance and a waste of bandwidth resources.

[0025] For ease of understanding, we'll use a physical machine as the computing node and a GPU (Graphics Processing Unit) as the computing unit on the physical machine as an example. Aggregate communication commonly uses the Ring algorithm. The Ring algorithm arranges all participating GPUs in a ring, and each GPU communicates only with the GPU preceding and following it in the ring. Once the ring topology is established, the communication path between GPUs is fixed.

[0026] Figure 1 is a schematic diagram of an exemplary ring topology. Referring to Figure 1, physical machine 1 has eight GPUs: GPU1, GPU2, GPU3, GPU4, GPU5, GPU6, GPU7, and GPU8. Physical machine 2 has eight GPUs: GPU9, GPU10, GPU11, GPU12, GPU13, GPU14, GPU15, and GPU16. Physical machines 1 and 2 communicate via network interface cards (NICs 1 and 2), which can be considered an inter-machine connection. Adjacent GPUs on the same physical machine communicate using high-speed interconnect technology, which can be considered an intra-machine connection. Through intra-machine and inter-machine connections, the 16 GPUs (GPU1, GPU2, GPU3, GPU4, GPU5, GPU6, GPU7, GPU8, GPU9, GPU10, GPU11, GPU12, GPU13, GPU14, GPU15, and GPU16) are sequentially connected, forming the ring topology shown by the dashed line in Figure 1.

[0027] The drawback of the ring network algorithm is that the bandwidth resource usage ratio between the inter-machine network and the internal network is fixed, and it cannot dynamically adjust the load on the inter-machine network, as detailed below:

[0028] 1. When the network card bandwidth between physical machines becomes a bottleneck, data is still transmitted along a fixed ring communication path, and the data transmission speed between physical machines decreases. Even if the internal network has sufficient bandwidth resources, it cannot be fully utilized.

[0029] 2. When the internal communication bandwidth of the physical machine becomes a bottleneck, data is still transmitted along a fixed ring communication path, and the data transmission speed within the physical machine decreases. Even if the inter-machine network has sufficient bandwidth resources, they cannot be fully utilized.

[0030] To address this, embodiments of this disclosure provide a combined communication method, device, storage medium, and computer program product. In this embodiment, for a computing system comprising N computing nodes, each computing node comprising M computing units, a logical topology is constructed. This logical topology includes a first ring topology between the N computing units belonging to different computing nodes and a second ring topology between the M computing units belonging to the same computing node. Based on these two different ring topologies, two different transmission paths can be constructed. One transmission path prioritizes the connection using the first ring topology for data transmission, while the other prioritizes the connection using the second ring topology. The number of data channels employing the two different transmission paths is controlled, and data transmission is performed on the corresponding transmission paths. This fully utilizes the bandwidth resources within and between computing nodes, maximizing bandwidth resource utilization, improving overall communication performance, and reducing bandwidth waste.

[0031] The technical solutions of this disclosure and how they solve the aforementioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The technical solutions provided by each embodiment of this disclosure are described in detail below with reference to the accompanying drawings.

[0032] Figure 2 is a flowchart of a collection communication method provided in an embodiment of this disclosure. Referring to Figure 2, the method may include the following steps:

[0033] 201. Determine the logical topology built on the physical topology of the computing system. The logical topology includes a first ring topology between N computing units belonging to different computing nodes and a second ring topology between M computing units belonging to the same computing node.

[0034] 202. Based on the bandwidth resources between computing nodes and the bandwidth resources within computing nodes, create K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path on the logical topology. K1 and K2 are both positive integers. The first transmission path is the path that prioritizes the use of the connection in the second ring topology to transmit data between M*N computing units, and the second transmission path is the path that prioritizes the use of the connection in the first ring topology to transmit data between M*N computing units.

[0035] 203. Data transmission is performed between M*N computing units using the first transmission path in K1 data channels; data transmission is performed between M*N computing units using the second transmission path in K2 data channels.

[0036] The aggregate communication method provided in this disclosure can be applied to a computing system, which includes N computing nodes, where N is a positive integer. These computing nodes include, but are not limited to, various computing resources such as physical machines, virtual machines, and cloud servers. Each computing node includes M computing units, where M is a positive integer. These computing units include, but are not limited to, GPUs, DPUs (Data Processing Units), NPUs (Neural Processing Units), or TPUs (Tensor Processing Units).

[0037] In practical applications, there are no restrictions on the executing entity of the collective communication method. The executing entity may include, but is not limited to, the CPU (Central Processing Unit) in any computing node, or other devices outside of the N computing nodes.

[0038] In this embodiment, the physical topology of the computing system reflects the system information of the computing system. The system information of the computing system includes, but is not limited to, the number of computing nodes included in the computing system, the number of computing units included in the computing nodes, the communication connection relationship between computing nodes, the relative positional relationship between computing nodes, the relative positional relationship between computing units, the communication connection relationship between computing units, and so on.

[0039] In this embodiment, a logical topology can be established in real time or in advance on the physical topology of the computing system. The logical topology includes a first ring topology between N computing units belonging to different computing nodes and a second ring topology between M computing units belonging to the same computing node.

[0040] Understandably, based on the physical topology of the computing system, we can identify N computing units belonging to different computing nodes and M computing units belonging to the same computing node. We can arrange the N computing units belonging to different computing nodes in a ring to form a first ring topology; and arrange the M computing units belonging to the same computing node in a ring to form a second ring topology, thereby establishing the logical topology of the computing system.

[0041] Specifically, the first ring topology is a ring topology spanning across computing nodes. It is established by arranging the N computing units connected via network interface cards (NICs) in a ring among N computing nodes. The second ring topology is a ring topology within a computing node. It is established by arranging the M computing units connected via the internal network of each computing node in a ring among M computing units. For the case of N computing nodes, each containing M computing units, M first ring topologies and N second ring topologies can be established.

[0042] In practical applications, N*M computing units can be arranged in a star topology, a ring topology, or a serial topology, etc., but are not limited to these. Connecting the M computing units in each computing node, which are connected end-to-end through the internal network of the computing node, in a first ring topology can be established; connecting the N computing units in different computing nodes, which are connected end-to-end through network cards, in a second ring topology can be established.

[0043] For example, in order to accurately and efficiently establish the logical topology, the implementation method of the logical topology built on the physical topology of the computing system is as follows: N*M computing units are arranged in an array, the array arrangement includes N rows and M columns, the N rows correspond to N computing nodes, and the M columns correspond to M computing units in the same computing node; based on the physical topology of the computing system, the M computing units in each row are connected end to end in sequence to obtain N second ring topologies; based on the physical topology of the computing system, the N computing units in each column are connected end to end in sequence to obtain M first ring topologies; wherein, the N second ring topologies and the M first ring topologies form the logical topology.

[0044] Taking the logical topology shown in Figure 3 as an example, the logical topology presents a two-dimensional network topology. All computing units participating in communication are arranged in an array. Computing units belonging to the same physical machine in the X direction (i.e., the row direction) form a second ring topology through intra-machine connections, while computing units belonging to different physical machines in the Y direction (i.e., the column direction) form a first ring topology through inter-machine connections.

[0045] Referring to Figure 3, there are four physical machines: physical machine 0, physical machine 1, physical machine 2, and physical machine 3. Each physical machine has eight GPUs: G0, G1, G2, G3, G4, G5, G6, and G7. The eight GPUs inside physical machine 0 form a second ring topology through internal connections. The eight GPUs inside physical machine 1 form a second ring topology through internal connections. The eight GPUs inside physical machine 3 form a second ring topology through internal connections. The eight GPUs inside physical machine 4 form a second ring topology through internal connections.

[0046] Referring to Figure 3, the four GPUs (G0 of physics machine 0, G0 of physics machine 1, G0 of physics machine 2, and G0 of physics machine 3) form a first ring topology through inter-machine connections; the four GPUs (G1 of physics machine 0, G1 of physics machine 1, G1 of physics machine 2, and G1 of physics machine 3) form a first ring topology through inter-machine connections; the four GPUs (G2 of physics machine 0, G2 of physics machine 1, G2 of physics machine 2, and G2 of physics machine 3) form a first ring topology through inter-machine connections; and the four GPUs (G3 of physics machine 0, G3 of physics machine 1, G3 of physics machine 2, and G3 of physics machine 3) form a first ring topology through inter-machine connections. Topology: The four GPUs of physics machine 0 (G4), physics machine 1 (G4), physics machine 2 (G4), and physics machine 3 (G4) form a first ring topology through inter-machine connections; the four GPUs of physics machine 0 (G5), physics machine 1 (G5), physics machine 2 (G5), and physics machine 3 (G5) form a first ring topology through inter-machine connections; the four GPUs of physics machine 0 (G6), physics machine 1 (G6), physics machine 2 (G6), and physics machine 3 (G6) form a first ring topology through inter-machine connections; the four GPUs of physics machine 0 (G7), physics machine 1 (G7), physics machine 2 (G7), and physics machine 3 (G7) form a first ring topology through inter-machine connections.

[0047] Based on the creation of the logical topology, K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path can be created on the logical topology according to the bandwidth resources between computing nodes and the bandwidth resources within computing nodes.

[0048] In this embodiment, a data channel refers to a logical channel defined from a data transmission perspective, responsible for carrying data transmitted between M*N computing units. Each data channel has its own channel identifier. This logical channel relies on the transmission path between computing units for data transmission. Different data channels can exist in parallel, and different data channels can use the same transmission path or different transmission paths.

[0049] In this embodiment, the transmission path is the actual physical or logical connection path between computing units, used for the actual data transmission. The transmission path consists of a series of links connecting the computing units, each link may have specific attributes such as bandwidth and latency. The planning of the transmission path may need to consider constraints such as bandwidth and latency to ensure efficient and reliable data transmission.

[0050] In practical applications, the first and second transmission paths can be pre-established, or they can be created in real time based on the path definition information; there are no restrictions on this.

[0051] Optionally, in order to better improve the overall communication performance, a first transmission path and a second transmission path can be determined based on the obtained predefined first path description information and second path description information, and the logical topology, the first path description information and the second path description information.

[0052] In practical applications, users can input predefined first path description information and second path description information on the interactive interface, or obtain predefined first path description information and second path description information from the configuration file, without any restrictions.

[0053] The first path description information describes relevant information about the first transmission path, including but not limited to: the transmission order between specified computing units, the minimum bandwidth requirement for each path segment, and the maximum allowable delay for each path segment. Path planning is performed based on the connection relationships between computing units across different computing nodes in the logical topology, the connection relationships within each computing node, and the first path description information to obtain the first transmission path. Specifically, with the goal of satisfying the transmission order between specified computing units, the minimum bandwidth requirement for each path segment, and the maximum allowable delay for each path segment, transmission links are first established between M computing units on the same computing node. Then, transmission links are established between N computing units belonging to different computing nodes, thus forming the first transmission path. The second path description information describes relevant information about the second transmission path, including but not limited to: the transmission order between specified computing units, the minimum bandwidth requirement for each path segment, and the maximum allowable delay for each path segment. Path planning is performed based on the connection relationships between computing units across different computing nodes in the logical topology, the connection relationships within each computing node, and the second path description information to obtain the second transmission path.

[0054] In this embodiment, the first transmission path is a path that preferentially uses connections in the second ring topology to transmit data between M*N computing units. That is, the first transmission path sequentially uses connections in the second ring topology and connections in the first ring topology to transmit data between M*N computing units. Optionally, when N*M computing units are arranged in an array, the first transmission path includes: all connections in the first second ring topology except for the head-to-tail connections, and all connections in each first ring topology except for the head-to-tail connections; wherein, the head-to-tail connection refers to the connection between the first computing unit and the tail computing unit in the first second ring topology or each first ring topology.

[0055] Taking the data transmission of G0 on physical machine 0 using the first transmission path as an example, firstly, referring to Figure 4, the data of G0 on physical machine 0 is transmitted to other GPUs such as G1, G2, G3, G4, G5, G6, and G7 on physical machine 0 through the internal connections of the physical machine. Next, referring to Figure 5, the data on G0, G1, G2, G3, G4, G5, G6, and G7 on physical machine 0 is transmitted to other GPUs through the inter-machine connections between physical machines. Therefore, the data transmitted on physical machine 0 using the first transmission path includes: the segment from physical machine 0 to physical machine 0 and ending at physical machine 0 and ending at physical machine 0 and ending at physical machine 0 and ending at physical machine 3 ...0 and ending at physical machine 3 and ending at physical machine 0 and ending at physical machine 0 and ending at physical machine 3 and ending at physical machine 0 and ending at physical machine 0 and ending at physical machine 3 and ending at physical machine 0 and ending at physical machine 0 and ending at physical machine 3 and ending at physical machine 0 and ending at physical machine 0 and ending at physical machine 3 and ending at physical machine 0 and ending at physical machine 0 and ending at physical machine 0 and ending at physical machine 3 and ending at physical machine 0 and ending at physical machine 0 and ending at physical machine 0 and ending at physical machine

[0056] In this embodiment, the second transmission path is a path that preferentially uses connections in the first ring topology to transmit data between M*N computing units. That is, the second transmission path sequentially uses connections in the first ring topology and connections in the second ring topology to transmit data between M*N computing units. Optionally, when N*M computing units are arranged in an array, the second transmission path includes: all connections in the first ring topology except for the head-to-tail connections and all connections in each second ring topology except for the head-to-tail connections; wherein, the head-to-tail connection refers to the connection between the first computing unit and the tail computing unit in the first ring topology or each second ring topology.

[0057] Taking the data transmission of G0 on physical machine 0 using the second transmission path as an example, firstly, referring to Figure 6, the data of G0 on physical machine 0 is transmitted to G0 on physical machine 1, G0 on physical machine 2, and G0 on physical machine 3 through inter-machine connections; then, referring to Figure 7, the data of G0 on physical machine 0 is transmitted to other GPUs such as G1, G2, G3, G4, G5, G6, and G7 on physical machine 0 through internal connections within the physical machine; the data of G0 on physical machine 1 is transmitted through... Data from G0 on physical machine 1 is transmitted via internal connections to other GPUs such as G1, G2, G3, G4, G5, G6, and G7 on physical machine 1; data from G0 on physical machine 2 is transmitted via internal connections to other GPUs such as G1, G2, G3, G4, G5, G6, and G7 on physical machine 2; data from G0 on physical machine 3 is transmitted via internal connections to other GPUs such as G1, G2, G3, G4, G5, G6, and G7 on physical machine 3. Therefore, the second transmission path used to transmit data on physical machine 0 includes: the segment from physical machine 0's G0 to physical machine 3's G0 (i.e., column direction), the segment from physical machine 0's G0 to physical machine 0's G7 (i.e., row direction), the segment from physical machine 1's G0 to physical machine 1's G7 (i.e., row direction), the segment from physical machine 2's G0 to physical machine 2's G7 (i.e., row direction), and the segment from physical machine 3's G0 to physical machine 3's G7 (i.e., row direction).

[0058] In this embodiment, in order to make full use of the bandwidth resources inside the computing nodes and the bandwidth resources between computing nodes, and to maximize the utilization of bandwidth resources, K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path can be created on the logical topology according to the bandwidth resources between computing nodes and the bandwidth resources inside the computing nodes.

[0059] Specifically, both using the first transmission path and using the second transmission path for data transmission consume bandwidth resources between computing nodes and within computing nodes. By controlling the number of data channels suitable for the first transmission path and the number of data channels suitable for the second transmission path, the utilization of bandwidth resources can be maximized.

[0060] For example, based on the bandwidth resources between computing nodes and the bandwidth resources within computing nodes, the implementation of creating K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path on the logical topology is as follows: Create C data channels on the logical topology, where C is a positive integer; with the goal of adapting the inter-node bandwidth and intra-node bandwidth consumed by the C data channels to the bandwidth resources between computing nodes and the bandwidth resources within computing nodes, respectively, determine the number of data channels K1 suitable for the first transmission path and the number of data channels K2 suitable for the second data transmission path, where K1 + K2 = C; divide the C data channels into K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path.

[0061] In this embodiment, the specific implementation method for adapting the inter-node bandwidth and intra-node bandwidth consumed by the C data channels to the bandwidth resources between computing nodes and within computing nodes is not limited. For example, the inter-node bandwidth consumed by the C data channels can be proportional to the bandwidth resources between computing nodes, and the intra-node bandwidth consumed by the C data channels can be proportional to the bandwidth resources within computing nodes, so as to achieve adaptation of the inter-node bandwidth and intra-node bandwidth consumed by the C data channels to the bandwidth resources between computing nodes and within computing nodes, respectively. Alternatively, the inter-node bandwidth and intra-node bandwidth consumed by the C data channels can be added together to obtain the total bandwidth resources consumed by the C data channels; the bandwidth resources between computing nodes and the bandwidth resources within computing nodes can be added together to obtain the standard total bandwidth resources; and the total bandwidth resources consumed by the C data channels can be proportional to the standard total bandwidth resources, so as to achieve adaptation of the inter-node bandwidth and intra-node bandwidth consumed by the C data channels to the bandwidth resources between computing nodes and within computing nodes, respectively.

[0062] Furthermore, to maximize bandwidth resource utilization, with the goal of adapting the inter-node bandwidth and intra-node bandwidth consumed by C data channels to the bandwidth resources between computing nodes and within computing nodes, respectively, the following method is used to determine the number of data channels K1 suitable for the first transmission path and the number of data channels K2 suitable for the second data transmission path: taking the number of data channels K1 as the unknown variable, calculate the inter-node bandwidth and intra-node bandwidth consumed by C data channels based on the first number of connections, the second number of connections, the third number of connections, and the fourth number of connections; taking the minimum absolute value of the difference between the first bandwidth ratio and the second bandwidth ratio as the optimization objective, solve for the number of data channels K1, and use C-K1 as the number of data channels K2;

[0063] Wherein, the first number of connections and the second number of connections refer to the number of connections belonging to the first ring topology and the number of connections belonging to the second ring topology in the first transmission path; the third number of connections and the fourth number of connections refer to the number of connections belonging to the first ring topology and the number of connections belonging to the second ring topology in the second transmission path; the first bandwidth ratio is the ratio between the inter-node bandwidth and the intra-node bandwidth consumed by C data channels, and the second bandwidth ratio is the ratio between the bandwidth resources between computing nodes and the bandwidth resources within the computing nodes.

[0064] Assume there are N computing nodes, each consisting of M computing units and C data channels. K1 is the number of data channels suitable for the first transmission path, and K2 is the number of data channels suitable for the second transmission path. Let BW denote the bandwidth resource between computing nodes corresponding to a single network interface card (NIC) connection (i.e., the NIC bandwidth allowed for a single NIC connection). nic The bandwidth resources within a compute node that allow connections between single compute units within the compute node (which can be referred to as the intra-machine bandwidth allowed for a single intra-machine connection) are denoted as BW. nvl BW nic BW nvl The quantity is known.

[0065] For each data channel, if the data channel uses the first transmission path for data transmission, then during the data transmission process, data transmission is first carried out through the network connection inside the computing node, and then through the network card connection. In this way, (N-1)×M network card connections are required, and M-1 network connections inside the computing node are required. The (N-1)×M network card connections are considered as the number of connections belonging to the first ring topology in the first transmission path, and the M-1 network connections inside the computing node are considered as the number of connections belonging to the second ring topology in the first transmission path.

[0066] For each data channel, if the data channel uses the second transmission path for data transmission, then during the data transmission process, data transmission is first performed through the network card connection, and then through the network connection inside the computing node. In this way, N-1 network card connections are required, and (M-1)×N internal connections are required. The N-1 network card connections are considered as the number of connections belonging to the first ring topology in the second transmission path, and the (M-1)×N internal connections are considered as the number of connections belonging to the second ring topology in the second transmission path.

[0067] For C data channels, with K1 suitable for the first transmission path and K2 suitable for the second data transmission path, the C data channels require K1×(N-1)×M+K2×(N-1) network card connections and K1×(M-1)+K2×(M-1)×N internal connections. The K1×(N-1)×M+K2×(N-1) network card connections can be considered as the inter-node bandwidth consumed by the C data channels, and the K1×(M-1)+K2×(M-1)×N internal connections can be considered as the intra-node bandwidth consumed by the C data channels.

[0068] According to the following formula (1), K1 is taken as the unknown quantity and K1 is solved according to formula (1).

[0069] In practical applications, so as to make Minimize (Min) as the optimization objective, and solve for K1, K2 = C - K1.

[0070] in, This represents the first bandwidth ratio. This represents the second bandwidth ratio.

[0071] In this embodiment, when dividing the C data channels into K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path, K1 data channels can be randomly selected from the C data channels as data channels suitable for the first transmission path, and the remaining K2 data channels can be used as data channels suitable for the second transmission path; or, based on the size of the identifiers of the C data channels, K1 data channels with the smallest or largest identifiers can be selected as data channels suitable for the first transmission path, and the remaining K2 data channels can be used as data channels suitable for the second transmission path, but this is not a limitation.

[0072] Based on the determination of K1 data channels and K2 data channels, data transmission can be performed between M*N computing units using the first transmission path in the K1 data channel; and data transmission can be performed between M*N computing units using the second transmission path in the K2 data channel.

[0073] As an example, the implementation of data transmission between M*N computing units using the first transmission path in K1 data channels is as follows: For any one of the K1 data channels, control data is transmitted from the first computing unit in the first second ring topology to the subsequent computing units in the first second ring topology and to the tail computing unit in the first second ring topology; and control data is transmitted from the first computing unit in each first ring topology to the subsequent computing units in each first ring topology and to the tail computing unit in each first ring topology.

[0074] In practical applications, each first ring topology can transmit data in parallel or serially, without restriction. Parallel data transmission can be understood as synchronous transmission, while serial data transmission can be understood as asynchronous transmission.

[0075] As an example, data transmission is performed between M*N computing units using a second transmission path in K2 data channels, including: for any one of the K2 data channels, controlling data to be transmitted from the first computing unit in the first first ring topology to the subsequent computing units in the first first ring topology and to the tail computing unit in the first second ring topology; and controlling data to be transmitted from the first computing unit in each second ring topology to the subsequent computing units in each second ring topology and to the tail computing unit in each second ring topology.

[0076] In practical applications, each second ring topology can transmit data in parallel or serially without restriction.

[0077] The technical solution provided in this disclosure addresses a computing system comprising N computing nodes, each node containing M computing units. It constructs a logical topology for this computing system, including a first ring topology connecting the N computing units belonging to different computing nodes and a second ring topology connecting the M computing units belonging to the same computing node. Based on these two different ring topologies, two different transmission paths can be constructed. One transmission path prioritizes the connection using the first ring topology for data transmission, while the other prioritizes the connection using the second ring topology. The number of data channels employing these two different transmission paths is controlled, and data transmission is performed on these channels using the corresponding transmission paths. This fully utilizes the bandwidth resources within and between computing nodes, maximizing bandwidth utilization, improving overall communication performance, and reducing bandwidth waste.

[0078] It is understood that each computing unit in the logical topology provided in this embodiment participates in two ring topologies simultaneously: one ring topology within a computing node and one ring topology between computing nodes. With this design, each computing unit has two predecessor computing units and two successor computing units. One pair of predecessor and successor computing units uses an internal network connection within the computing node, while the other pair uses a network interface card (NIC) connection.

[0079] In practical applications, based on logical topology, the optimal data transmission channel is dynamically selected according to the internal bandwidth (bandwidth resources within the computing node) and inter-machine bandwidth (bandwidth resources between computing nodes) in the real environment, so as to maximize the use of bandwidth resources and avoid resource waste and communication bottlenecks.

[0080] Specifically, when the bandwidth of the network card between machines is limited, more internal connections can be used for data transmission to make full use of the high-speed communication bandwidth within the machine and reduce the load on the network card.

[0081] When internal bandwidth resources are scarce, more use can be made of inter-machine connections for data transfer, making full use of network card bandwidth resources and avoiding internal network congestion.

[0082] Figure 8 is a schematic diagram of a collective communication device provided in an embodiment of this disclosure. Applied to a computing system, the computing system includes N computing nodes, and each computing node includes M computing units, where M and N are both positive integers. Referring to Figure 8, the device may include:

[0083] The topology determination module 81 is used to determine a logical topology built upon the physical topology of the computing system. This logical topology can be built in real-time or pre-built. Specifically, the topology determination module 81 can be used to: build the logical topology upon the physical topology of the computing system. The logical topology includes a first ring topology between N computing units belonging to different computing nodes and a second ring topology between M computing units belonging to the same computing node.

[0084] The channel creation module 82 is used to create K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path on the logical topology, based on the bandwidth resources between computing nodes and the bandwidth resources within computing nodes. K1 and K2 are both positive integers. The first transmission path is the path that prioritizes the use of connections in the second ring topology for data transmission between M*N computing units, and the second transmission path is the path that prioritizes the use of connections in the first ring topology for data transmission between M*N computing units.

[0085] The data transmission module 83 is used to transmit data between M*N computing units using a first transmission path in K1 data channels; and to transmit data between M*N computing units using a second transmission path in K2 data channels.

[0086] Optionally, the topology determination module 81 is specifically used for: arranging N*M computing units in an array, the array arrangement including N rows and M columns, where the N rows correspond to N computing nodes and the M columns correspond to M computing units within the same computing node; based on the physical topology of the computing system, connecting the M computing units in each row sequentially to obtain N second ring topologies; based on the physical topology of the computing system, connecting the N computing units in each column sequentially to obtain M first ring topologies; wherein the N second ring topologies and the M first ring topologies form a logical topology.

[0087] Optionally, the channel creation module 82 is specifically used for: creating C data channels on the logical topology, where C is a positive integer; determining the number of data channels K1 suitable for the first transmission path and the number of data channels K2 suitable for the second data transmission path, with the goal of adapting the inter-node bandwidth and intra-node bandwidth consumed by the C data channels to the bandwidth resources between computing nodes and the bandwidth resources within computing nodes, respectively, where K1 + K2 = C; and dividing the C data channels into K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path.

[0088] Optionally, when the channel creation module 82 determines the number of data channels K1 suitable for the first transmission path and the number of data channels K2 suitable for the second data transmission path, with the goal of adapting the inter-node bandwidth and intra-node bandwidth consumed by C data channels to the bandwidth resources between computing nodes and within computing nodes, respectively, the following is specifically used: taking the number of data channels K1 as the unknown quantity, and based on the first connection number, the second connection number, the third connection number, and the fourth connection number, calculate the inter-node bandwidth and intra-node bandwidth consumed by C data channels; the first connection number and the second connection number refer to the bandwidth resources between computing nodes and within computing nodes in the first transmission path. The number of connections in the first ring topology and the number of connections in the second ring topology; the third number of connections and the fourth number of connections refer to the number of connections in the second transmission path belonging to the first ring topology and the number of connections belonging to the second ring topology; with the optimization objective of minimizing the absolute value of the difference between the first bandwidth ratio and the second bandwidth ratio, the number of data channels K1 is solved, and C-K1 is taken as the number of data channels K2; the first bandwidth ratio is the ratio between the inter-node bandwidth and the intra-node bandwidth consumed by C data channels, and the second bandwidth ratio is the ratio between the bandwidth resources between computing nodes and the bandwidth resources within computing nodes.

[0089] Optionally, when the channel creation module 82 divides the C data channels into K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path, it specifically performs the following: randomly selects K1 data channels from the C data channels as data channels suitable for the first transmission path, and uses the remaining K2 data channels as data channels suitable for the second transmission path; or selects the K1 data channels with the smallest or largest identifiers according to the size of the identifiers of the C data channels as data channels suitable for the first transmission path, and uses the remaining K2 data channels as data channels suitable for the second transmission path.

[0090] Optionally, the channel creation module 82 is further configured to: obtain predefined first path description information and second path description information; and determine the first transmission path and the second transmission path based on the logical topology, the first path description information, and the second path description information.

[0091] Optionally, when N*M computing units are arrayed, the first transmission path includes: other connections in the first second ring topology besides the head-to-tail connection and other connections in each first ring topology besides the head-to-tail connection; wherein, the head-to-tail connection refers to the connection between the first computing unit and the tail computing unit in the first second ring topology or each first ring topology.

[0092] Optionally, when N*M computing units are arrayed, the second transmission path includes: all connections in the first ring topology except for the head-to-tail connection and all connections in each second ring topology except for the head-to-tail connection; wherein, the head-to-tail connection refers to the connection between the first computing unit and the tail computing unit in the first ring topology or each second ring topology.

[0093] Optionally, the data transmission module 83 is specifically used for: controlling the transmission of data from the first computing unit in the first second ring topology to the subsequent computing units in the first second ring topology and to the tail computing unit in the first second ring topology for any one of the K1 data channels; and controlling the transmission of data from the first computing unit in each first ring topology to the subsequent computing units in each first ring topology and to the tail computing unit in each first ring topology.

[0094] Optionally, the data transmission module 83 is specifically used for: controlling the transmission of data from the first computing unit in the first first ring topology to the subsequent computing units in the first first ring topology and to the tail computing unit in the first first ring topology for any one of the K2 data channels; and controlling the transmission of data from the first computing unit in each second ring topology to the subsequent computing units in each second ring topology and to the tail computing unit in each second ring topology.

[0095] Optionally, the computing node is a physical machine, and the computing unit is a GPU, DPU, NPU, or TPU.

[0096] The device shown in Figure 8 can execute the method shown in the embodiment of Figure 2, and its implementation principle will not be repeated here. The specific ways in which each module and unit of the device shown in Figure 8 in the above embodiments performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0097] The technical solution provided in this disclosure addresses a computing system comprising N computing nodes, each node containing M computing units. It constructs a logical topology for this computing system, including a first ring topology connecting the N computing units belonging to different computing nodes and a second ring topology connecting the M computing units belonging to the same computing node. Based on these two different ring topologies, two different transmission paths can be constructed. One transmission path prioritizes the connection using the first ring topology for data transmission, while the other prioritizes the connection using the second ring topology. The number of data channels employing these two different transmission paths is controlled, and data transmission is performed on these channels using the corresponding transmission paths. This fully utilizes the bandwidth resources within and between computing nodes, maximizing bandwidth utilization, improving overall communication performance, and reducing bandwidth waste.

[0098] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 201 to 203 can be device A; or the execution subject of steps 201 and 202 can be device A, and the execution subject of step 203 can be device B; and so on.

[0099] Furthermore, in some of the processes described in the above embodiments and accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 401, 402, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0100] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0101] Figure 9 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. As shown in Figure 9, the electronic device includes: a memory 91 and a processor 92;

[0102] Memory 91 is used to store computer programs and can be configured to store various other data to support operation on the computing platform. Examples of this data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc.

[0103] The memory 91 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0104] Processor 92, coupled to memory 91, is used to execute computer programs in memory 91 for:

[0105] Determine the logical topology built on the physical topology of the computing system, wherein the computing system includes N computing nodes, each computing node includes M computing units, where M and N are both positive integers, and the logical topology includes a first ring topology between N computing units belonging to different computing nodes and a second ring topology between M computing units belonging to the same computing node.

[0106] Based on the bandwidth resources between computing nodes and within computing nodes, K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path are created on the logical topology, where K1 and K2 are both positive integers; wherein, the first transmission path is a path that preferentially uses the connection in the second ring topology to transmit data between M*N computing units, and the second transmission path is a path that preferentially uses the connection in the first ring topology to transmit data between M*N computing units.

[0107] In the K1 data channels, the first transmission path is used to transmit data between M*N computing units; in the K2 data channels, the second transmission path is used to transmit data between M*N computing units.

[0108] In one optional embodiment, the processor 92 establishes a logical topology on top of the physical topology of the computing system, including: arranging N*M computing units in an array, the array arrangement including N rows and M columns, where the N rows correspond to N computing nodes and the M columns correspond to M computing units in the same computing node; based on the physical topology of the computing system, connecting the M computing units in each row sequentially to obtain N second ring topologies; and based on the physical topology of the computing system, connecting the N computing units in each column sequentially to obtain M first ring topologies; wherein the N second ring topologies and the M first ring topologies form the logical topology.

[0109] In an optional embodiment, the processor 92 creates K1 data channels suitable for a first transmission path and K2 data channels suitable for a second transmission path on the logical topology based on the bandwidth resources between computing nodes and the bandwidth resources within computing nodes. This includes: creating C data channels on the logical topology, where C is a positive integer; determining the number of data channels K1 suitable for the first transmission path and the number of data channels K2 suitable for the second transmission path, with the goal of adapting the inter-node bandwidth and intra-node bandwidth consumed by the C data channels to the bandwidth resources between computing nodes and the bandwidth resources within computing nodes, respectively, where K1 + K2 = C; and dividing the C data channels into K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path.

[0110] In an optional embodiment, the processor 92 aims to adapt the inter-node bandwidth and intra-node bandwidth consumed by the C data channels to the bandwidth resources between computing nodes and within computing nodes, respectively. It determines the number of data channels K1 suitable for the first transmission path and the number of data channels K2 suitable for the second data transmission path, including: using the number of data channels K1 as the variable to be determined, calculating the inter-node bandwidth and intra-node bandwidth consumed by the C data channels based on the first number of connections, the second number of connections, the third number of connections, and the fourth number of connections; the first number of connections and the second number of connections refer to the bandwidth within the first ring of the first transmission path. The number of connections in the first ring topology and the number of connections in the second ring topology; the third number of connections and the fourth number of connections refer to the number of connections in the second transmission path belonging to the first ring topology and the number of connections belonging to the second ring topology; with the optimization objective of minimizing the absolute value of the difference between the first bandwidth ratio and the second bandwidth ratio, the number of data channels K1 is solved, and C-K1 is taken as the number of data channels K2; the first bandwidth ratio is the ratio between the inter-node bandwidth and the intra-node bandwidth consumed by C data channels, and the second bandwidth ratio is the ratio between the bandwidth resources between computing nodes and the bandwidth resources within computing nodes.

[0111] In an optional embodiment, the processor 92 divides the C data channels into K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path, including: randomly selecting K1 data channels from the C data channels as data channels suitable for the first transmission path, and using the remaining K2 data channels as data channels suitable for the second transmission path; or, selecting the K1 data channels with the smallest or largest identifiers according to the size of the identifiers of the C data channels as data channels suitable for the first transmission path, and using the remaining K2 data channels as data channels suitable for the second transmission path.

[0112] In an optional embodiment, the processor 92 is further configured to: acquire predefined first path description information and second path description information; and determine the first transmission path and the second transmission path based on the logical topology, the first path description information and the second path description information.

[0113] In an optional embodiment, when N*M computing units are arrayed, the first transmission path includes: other connections in the first second ring topology besides the head-to-tail connection and other connections in each first ring topology besides the head-to-tail connection; wherein, the head-to-tail connection refers to the connection between the first computing unit and the tail computing unit in the first second ring topology or each first ring topology.

[0114] Accordingly, when N*M computing units are arrayed, the second transmission path includes: all connections in the first ring topology except for the head-to-tail connection and all connections in each second ring topology except for the head-to-tail connection; wherein, the head-to-tail connection refers to the connection between the first computing unit and the tail computing unit in the first ring topology or each second ring topology.

[0115] In an optional embodiment, the processor 92 performs data transmission among M*N computing units using the first transmission path in the K1 data channels, including: for any one of the K1 data channels, controlling the data to be transmitted from the first computing unit in the first second ring topology to the subsequent computing units in the first second ring topology and to the tail computing unit in the first second ring topology; and controlling the data to be transmitted from the first computing unit in each first ring topology to the subsequent computing units in each first ring topology and to the tail computing unit in each first ring topology.

[0116] In an optional embodiment, the processor 92 uses the second transmission path to transmit data between M*N computing units in the K2 data channels, including: for any one of the K2 data channels, controlling the data to be transmitted from the first computing unit in the first first ring topology to the subsequent computing units in the first first ring topology and to the tail computing unit in the first first ring topology; and controlling the data to be transmitted from the first computing unit in each second ring topology to the subsequent computing units in each second ring topology and to the tail computing unit in each second ring topology.

[0117] Optionally, the electronic device is implemented as a computing node in a computing system, and the electronic device also includes: M computing units, where M is a positive integer.

[0118] Optionally, as shown in Figure 9, the electronic device may also include other components such as a communication component 93, a display 94, a power supply component 95, and an audio component 96. Figure 9 only schematically shows some components and does not imply that the electronic device only includes the components shown in Figure 9. Furthermore, the components within the dashed boxes in Figure 9 are optional, not mandatory, and their specific inclusion depends on the product form of the electronic device. The electronic device of this embodiment can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT (Internet of Things) device, or a server-side device such as a conventional server, cloud server, or server array. If the electronic device of this embodiment is a terminal device such as a desktop computer, laptop computer, or smartphone, it may include the components within the dashed boxes in Figure 9; if the electronic device of this embodiment is a server-side device such as a conventional server, cloud server, or server array, it may not include the components within the dashed boxes in Figure 9.

[0119] For a detailed description of the implementation process of each action by the processor, please refer to the relevant descriptions in the foregoing method embodiments or device embodiments, which will not be repeated here.

[0120] Accordingly, this disclosure also provides a computer-readable storage medium storing a computer program, which, when executed, can perform the steps that can be executed by an electronic device in the above method embodiments.

[0121] Accordingly, this disclosure also provides a computer program product, including a computer program / instructions, which, when executed by a processor, enable the processor to perform the steps that can be executed by an electronic device in the above method embodiments.

[0122] The aforementioned communication components are configured to facilitate wired or wireless communication between the device containing the communication components and other devices. The device containing the communication components can access wireless networks based on communication standards, such as WiFi (Wireless Fidelity), 2G (2nd Generation), 3G (3rd Generation), 4G (4th Generation) / LTE (long Term Evolution), 5G (5th Generation), or combinations thereof. In one exemplary embodiment, the communication components receive broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication components also include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth, and other technologies.

[0123] The aforementioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.

[0124] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.

[0125] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0126] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams.

[0128] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0129] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0130] In a typical configuration, a computing device includes one or more processors (Central Processing Units, CPUs), input / output interfaces, network interfaces, and memory.

[0131] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0132] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0133] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0134] The above are merely embodiments of this disclosure and are not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.

Claims

1. A collective communication method, characterized in that, Applied to a computing system, the computing system comprising N computing nodes, each computing node comprising M computing units, where M and N are both positive integers, the method includes: Determine the logical topology built on the physical topology of the computing system, the logical topology including a first ring topology between N computing units belonging to different computing nodes and a second ring topology between M computing units belonging to the same computing node; Based on the bandwidth resources between computing nodes and within computing nodes, K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path are created on the logical topology, where K1 and K2 are both positive integers; wherein, the first transmission path is a path that preferentially uses the connection in the second ring topology to transmit data between M*N computing units, and the second transmission path is a path that preferentially uses the connection in the first ring topology to transmit data between M*N computing units. In the K1 data channels, the first transmission path is used to transmit data between M*N computing units; in the K2 data channels, the second transmission path is used to transmit data between M*N computing units.

2. The method according to claim 1, characterized in that, Determine the logical topology built upon the physical topology of the computing system, including: N*M computing units are arranged in an array, the array arrangement including N rows and M columns, where the N rows correspond to N computing nodes and the M columns correspond to M computing units in the same computing node; Based on the physical topology of the computing system, the M computing units in each row are connected end to end to obtain N second ring topologies. Based on the physical topology of the computing system, the N computing units in each column are connected end to end to obtain M first ring topologies. The N second ring topologies and the M first ring topologies form the logical topology.

3. The method according to claim 1, characterized in that, Based on the bandwidth resources between computing nodes and within computing nodes, K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path are created on the logical topology, including: Create C data channels on the logical topology, where C is a positive integer; With the goal of adapting the inter-node bandwidth and intra-node bandwidth consumed by C data channels to the bandwidth resources between computing nodes and within computing nodes, respectively, the number of data channels K1 suitable for the first transmission path and the number of data channels K2 suitable for the second data transmission path are determined, where K1 + K2 = C; The C data channels are divided into K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path.

4. The method according to claim 3, characterized in that, With the goal of adapting the inter-node bandwidth and intra-node bandwidth consumed by C data channels to the bandwidth resources between computing nodes and within computing nodes, respectively, the number of data channels K1 suitable for the first transmission path and the number of data channels K2 suitable for the second data transmission path are determined, including: With the number of data channels K1 as the unknown quantity, calculate the inter-node bandwidth and intra-node bandwidth consumed by C data channels based on the first connection number, the second connection number, the third connection number, and the fourth connection number. The first number of connections and the second number of connections refer to the number of connections belonging to the first ring topology and the number of connections belonging to the second ring topology in the first transmission path; The third number of connections and the fourth number of connections refer to the number of connections belonging to the first ring topology and the number of connections belonging to the second ring topology in the second transmission path, respectively. The optimization objective is to minimize the absolute value of the difference between the first bandwidth ratio and the second bandwidth ratio. The number of data channels K1 is then calculated, and C-K1 is taken as the number of data channels K2. The first bandwidth ratio is the ratio between the inter-node bandwidth and the intra-node bandwidth consumed by C data channels, and the second bandwidth ratio is the ratio between the bandwidth resources between computing nodes and the bandwidth resources within the computing nodes.

5. The method according to claim 3, characterized in that, Dividing the C data channels into K1 data channels suitable for the first transmission path and K2 data channels suitable for the second transmission path includes: K1 data channels are randomly selected from C data channels as data channels suitable for the first transmission path, and the remaining K2 data channels are selected as data channels suitable for the second transmission path. or Based on the size of the identifiers of the C data channels, select the K1 data channels with the smallest or largest identifiers as the data channels suitable for the first transmission path, and use the remaining K2 data channels as the data channels suitable for the second transmission path.

6. The method according to claim 1, characterized in that, Also includes: Obtain predefined first path description information and second path description information; The first transmission path and the second transmission path are determined based on the logical topology, the first path description information, and the second path description information.

7. The method according to any one of claims 1-6, characterized in that, When N*M computing units are arrayed, the first transmission path includes: all connections in the first second ring topology except for the head-to-tail connection and all connections in each first ring topology except for the head-to-tail connection; The term "head-to-tail connection" refers to the connection between the first second ring topology or the first computing unit and the tail computing unit in each first ring topology.

8. The method according to any one of claims 1-6, characterized in that, When N*M computing units are arrayed, the second transmission path includes: all connections in the first ring topology except for the head-to-tail connection and all connections in each second ring topology except for the head-to-tail connection; The term "head-to-tail connection" refers to the connection between the first computing unit and the tail computing unit in the first first ring topology or each second ring topology.

9. The method according to claim 7, characterized in that, In the K1 data channels, data transmission is performed between M*N computing units using the first transmission path, including: For any one of the K1 data channels, control data is transmitted from the first computing unit in the first second ring topology to the subsequent computing units in the first second ring topology and finally to the tail computing unit in the first second ring topology; and The data is controlled to be transmitted from the first computing unit in each first ring topology to the subsequent computing units in each first ring topology and to the tail computing unit in each first ring topology.

10. The method according to claim 8, characterized in that, In the K2 data channels, data is transmitted between M*N computing units using the second transmission path, including: For any one of the K2 data channels, control data is transmitted from the first computing unit in the first ring topology to the subsequent computing units in the first ring topology and finally to the tail computing unit in the first ring topology; and The data is controlled to be transmitted from the first computing unit in each second ring topology to the subsequent computing units in each second ring topology and to the tail computing unit in each second ring topology.

11. The method according to claim 10, characterized in that, At least some of the N computing nodes further include a CPU, and the execution subject of the collective communication method is the CPU on any of the computing nodes including the CPU; or, the execution subject of the collective communication method is another device other than the N computing nodes.

12. An electronic device, characterized in that, include: Memory and processor; The memory stores a computer program; the processor is coupled to the memory and is used to execute the computer program to implement the steps of the method according to any one of claims 1-11.

13. The electronic device according to claim 12, characterized in that, The electronic device is implemented as a computing node in the computing system, and the electronic device further includes M computing units, where M is a positive integer.

14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it causes the processor to perform the steps of the method according to any one of claims 1-11.

15. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, causes the processor to perform the steps of the method according to any one of claims 1-11.