Computing system and computing method

By adopting a ring communication topology and direct connection of shared memory devices in a distributed computing system, the communication path between computing devices is simplified, the latency and congestion problems caused by switches are solved, and efficient data transmission is achieved.

CN120980077AActive Publication Date: 2025-11-18INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511504463.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2025-11-18
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

In distributed computing systems, data transmission speeds are low, especially in high-frequency communication scenarios. Switch architectures lead to increased network latency and congestion, making it difficult to meet the high-efficiency transmission requirements of tasks such as model training.

Method used

A ring communication topology and shared memory devices are adopted. The computing devices and shared memory devices are directly connected through direct device links and direct memory links, which simplifies the communication path, enables hardware-level memory access, and avoids switch forwarding delay.

Benefits of technology

It significantly improves the data transmission speed of distributed computing systems, solves the transmission delay and congestion problems caused by switches, and enhances the efficiency of computing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980077A_ABST
    Figure CN120980077A_ABST
Patent Text Reader

Abstract

The invention provides a computing system and a computing method, which can be applied to the field of distributed technologies. The computing system comprises a plurality of computing devices, any computing device is directly connected to the other two computing devices in the plurality of computing devices through a host direct connection link so as to form a ring communication topology based on the plurality of computing devices, the adjacent computing devices in the ring communication topology carry out data transmission through the host direct connection links between the adjacent computing devices; the plurality of shared memory devices are directly connected with one another through the memory direct-connection links, and the plurality of shared memory devices are directly connected to the corresponding computing devices in the plurality of computing devices through the device direct-connection links respectively.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of distributed technology, in particular to a computing system and a computing method. BACKGROUND

[0002] In the face of huge computing demand, the computing power and storage resources of a single device are difficult to bear, and distributed computing systems emerge as the times require. However, in the related art, the network architecture of the distributed computing system has low data transmission speed, which limits the efficiency of model training. SUMMARY

[0003] In view of the above problems, the present application provides a computing system and a computing method.

[0004] According to an aspect of the present application, a computing system is provided, comprising: a plurality of computing devices, wherein any computing device is directly connected to another two computing devices in the plurality of computing devices through a host direct link, to constitute a ring communication topology based on the plurality of computing devices, wherein adjacent computing devices in the ring communication topology perform data transmission through the host direct link between each other; a plurality of shared memory devices directly connected to each other through a memory direct link, the plurality of shared memory devices are respectively directly connected to corresponding computing devices in the plurality of computing devices through a device direct link, wherein the plurality of computing devices comprise a first computing device and a second computing device which are spaced apart in the ring communication topology, the plurality of shared memory devices comprise a first shared memory device directly connected to the first computing device and a second shared memory device directly connected to the second computing device, the first computing device is configured to directly write first data for the second computing device to the first shared memory device through the device direct link, the first shared memory device is configured to directly write the first data to the second shared memory device through the memory direct link, and the second computing device is configured to directly read the first data from the second shared memory device through the device direct link.

[0005] According to another aspect of this application, a computing method is provided, applied to a computing system. The computing method includes: transmitting data between multiple computing devices in the computing system, and performing a data processing task based on the transmitted data to obtain a data processing result. Each computing device is directly connected to two other computing devices among the multiple computing devices via a host direct link to form a ring communication topology based on the multiple computing devices. Adjacent computing devices in the ring communication topology transmit data through host direct links. The computing system also includes multiple shared memory devices directly connected to each other via memory direct links. The multiple shared memory devices are respectively directly connected to corresponding computing devices among the multiple computing devices via device direct links. The multiple computing devices include a first computing device and a second computing device spaced apart in the ring communication topology. The multiple shared memory devices include a first shared memory device directly connected to the first computing device and a second shared memory device directly connected to the second computing device, such that the first computing device directly writes first data for the second computing device to the first shared memory device via the device direct link, the first shared memory device directly writes the first data to the second shared memory device via the memory direct link, and the second computing device directly reads the first data from the second shared memory device via the device direct link.

[0006] According to embodiments of this application, for a first computing device and a second computing device spaced apart in a ring communication topology, when the first computing device needs to write first data to the second computing device, the first computing device can directly write the first data to the first shared memory device via a direct device link. Then, the first shared memory device can directly write the first data to the second shared memory device via a direct memory link, allowing the second computing device to read the first data from the second shared memory device via the direct device link. Thus, by using shared memory for data transmission between spaced computing devices in a ring communication topology, this application simplifies the communication path between spaced computing devices in a ring communication topology to a hardware-level access path, achieving communication between spaced computing devices in a ring communication topology through memory transfer. This at least partially solves the transmission latency caused by using switches for data transmission, improving the data transmission speed of the distributed computing system. Attached Figure Description

[0007] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, which will be explained in conjunction with the drawings.

[0008] Figure 1 A schematic diagram of a computing system according to an embodiment of this application is shown.

[0009] Figure 2 A schematic diagram of a computing system according to another embodiment of this application is shown.

[0010] Figure 3 A schematic diagram of a computing system according to another embodiment of this application is shown.

[0011] Figure 4A A schematic diagram of a computing system according to another embodiment of this application is shown.

[0012] Figure 4B A schematic diagram of a computing system according to another embodiment of this application is shown.

[0013] Figure 4C A schematic diagram of a computing system according to another embodiment of this application is shown.

[0014] Figure 5 A schematic diagram of a computing network based on a computing system according to an embodiment of this application is shown.

[0015] Figure 6 A flowchart of a calculation method according to an embodiment of this application is shown. Detailed Implementation

[0016] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0018] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0019] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0020] It should be noted that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. The terms "installed," "connected," and "linked" should be interpreted broadly, for example, they can be fixed connections, detachable connections, or integral connections; they can be mechanical connections or electrical connections; they can be direct connections or indirect connections through an intermediate medium; they can be internal connections between two elements. The terms "parallel," "perpendicular," and "equal" include the described situation and situations similar to the described situation, the range of which is within an acceptable deviation range, wherein the acceptable deviation range is determined by those skilled in the art taking into account the measurement under discussion and the error associated with the measurement of a particular quantity (i.e., the limitations of the measurement system). For example, "parallel" includes absolute parallelism and approximate parallelism, where an acceptable deviation range for approximate parallelism can be, for example, within 5°; "perpendicular" includes absolute perpendicularity and approximate perpendicularity, where an acceptable deviation range for approximate perpendicularity can also be, for example, within 5°. "Equal" includes absolute equality and approximate equality, where an acceptable deviation range for approximate equality can be, for example, a difference between the two equal items being less than or equal to 5% of either one. Those skilled in the art will understand the specific meaning of the above terms in this application based on the specific circumstances.

[0021] Faced with massive computing demands, the computing power and storage resources of a single device are insufficient, leading to the development of distributed computing systems. However, a key drawback of distributed computing systems is their low data transmission speed, specifically the low data transmission speed between nodes within the system.

[0022] On the one hand, in distributed computing systems, the subtasks of the same data processing task have low independence from each other. For example, during iterative training, nodes in a distributed computing system need to communicate periodically to synchronize gradients among themselves. Artificial Intelligence (AI) tasks have extremely low independence, requiring periodic communication during iterative training to achieve crucial gradient synchronization.

[0023] On the other hand, the operation of artificial intelligence services heavily relies on units such as Graphics Processing Units (GPUs) and Neural Processing Units (NPUs), with a high degree of concurrency in direct communication between these units. Therefore, it is necessary to improve the computing power of distributed computing systems by innovating the technologies related to GPUs and NPUs.

[0024] Building upon this foundation, significant technological innovations have been implemented around units such as graphics processing units and neural network processing units, resulting in a leapfrog improvement in the computing power of distributed computing systems. However, while the computing power bottleneck has been gradually overcome, the strong synchronicity of AI business communications means that the final performance of any communication transaction depends on the slowest communication link, making the network links for data transmission and interaction a new performance constraint. Under these circumstances, the computing network topology of some distributed computing systems is no longer sufficient to meet the requirements of efficient data transmission and low-latency interaction in distributed machine learning.

[0025] In some solutions, distributed computing systems use switches to achieve large-scale expansion of computing clusters, such as the Fat-tree and Spine-leaf architectures of the CLOS architecture, where hosts are interconnected via switches. However, this architecture has significant bottlenecks in high-frequency communication scenarios. Taking model training tasks as an example, the low-latency and high-bandwidth communication requirements of model training tasks are fundamentally contradictory to the routing mechanism of switches: each data transmission needs to be forwarded through multiple levels of switches, and the number of switch hops directly leads to a linear increase in network latency. The number of switch hops refers to the number of switches through which data passes during transmission. In many-to-one communication modes, such as distributed computing systems used for training model parameters, congestion problems such as incasts are easily triggered when a large number of concurrent requests are concentrated on a single receiving node. For example, when multiple computing nodes send data to the same computing node through switch ports, the data from multiple computing nodes must first converge to the target port buffer of the switch before being forwarded to the receiving computing node. If the concurrent traffic exceeds the capacity of the target port buffer, it can lead to packet loss or retransmission, resulting in a spike in latency.

[0026] In some model training algorithms, such as Allreduce, model parameter synchronization can be achieved through multiple rounds of data aggregation and broadcasting. Each round of communication relies on the efficient forwarding of switches. As the scale of the computing cluster increases, bandwidth contention on switch ports intensifies, further exacerbating head-of-line blocking and ACK (Acknowledgement) backlog problems caused by caching mechanisms, ultimately resulting in a significant increase in synchronization latency during training. Furthermore, hardware limitations such as the depth of the switch's buffer area and the mismatch between switch port rates become key bottlenecks restricting performance improvements in distributed computing systems.

[0027] For example, Ethernet communication requires traversing the entire network protocol stack, resulting in a latency of 2-4 microseconds. This network protocol stack can include, for example, the application layer, the Transmission Control Protocol / Internet Protocol (TCP / IP) layer, the network interface card (NIC) driver, and the physical NIC. Alternatively, Remote Direct Memory Access (RDMA) communication can be implemented by bypassing the kernel protocol stack and utilizing a host channel adapter. With InfiniBand technology, the latency can be reduced to approximately 1 microsecond, and with Remote Direct Memory Access over Converged Ethernet Version 2 (RoCEv2) technology, the latency can be reduced to approximately 1-2 microseconds.

[0028] In view of this, this application provides a computing system in which computing devices can directly read and write to shared memory devices for data transmission, thereby simplifying the communication path to hardware-level memory access and compressing latency to 300-400 nanoseconds. In some solutions, the communication latency is high because the data transmission path requires a large number of switch hops. In contrast, the computing system of this application can significantly reduce communication latency. In this application, the number of switch hops for data transmission between at least some or even all computing devices is 0. The computing system of this application will be described below with reference to the accompanying drawings.

[0029] Figure 1 A schematic diagram of a computing system according to an embodiment of this application is shown.

[0030] like Figure 1As shown, the computing system in this embodiment may include a computing device and a shared memory device. However, it should be understood that this application is not limited thereto. In other embodiments of this application, the computing system may also include other devices, such as control devices, etc., and this application does not limit this.

[0031] Computing devices can be used to perform computational tasks to obtain computational data. Further, in this embodiment, there can be multiple computing devices. In the multiple computing devices, any one device is directly connected to two other computing devices to form a ring communication topology based on the multiple computing devices, wherein adjacent computing devices in the ring communication topology transmit data through direct connections to each other.

[0032] Furthermore, the computing device may include one or more computing nodes. The data transmitted by multiple computing devices may be computational data calculated by the computing nodes, but this application is not limited to this; the multiple computing devices may also transmit other data.

[0033] In other embodiments of this application, data from multiple computing devices can also be transmitted via other devices. For example, in embodiments of this application, computational data from multiple computing devices can be transmitted via a shared memory device.

[0034] A shared memory device can be used as common memory for multiple computing devices. Specifically, any computing device can perform data write operations or data read operations on the shared memory device. However, the embodiments of this application are not limited to this. In other embodiments of this application, the multiple computing devices may include a first computing device and a second computing device. The first computing device and the second computing device may refer to two computing devices spaced apart in a ring communication topology. Based on this, when the first computing device writes data to the shared memory device, the second computing device can read the data written by the first computing device from the shared memory device. In this way, data transmission between the first computing device and the second computing device can be realized.

[0035] For example, shared memory devices can be implemented based on Compute Express Link (CXL) multi-head devices. Here, "head" can refer to a port. Compared to the nanosecond-level latency of switched solutions, CXL multi-head devices can compress communication latency, resulting in an order-of-magnitude improvement in data transmission performance. Specifically, shared memory devices enable hardware-level memory access. Transmitting data from computing devices through shared memory devices avoids the entire network forwarding process on the switch link, thus eliminating the speed mismatch problem associated with switch transmission ports. For instance, a switch, existing independently of the computing device, can lead to speed mismatch issues. However, as a memory extension for the computing device, the CXL multi-head device can abstract the computing device and the corresponding CXL multi-head device into a unified hardware device connected via a bus, resulting in a more accurate speed match. This avoids the data transmission congestion problems caused by the aforementioned port speed mismatch issues associated with switches.

[0036] Furthermore, in this embodiment, there can be multiple shared memory devices. Based on this, the shared memory devices can be connected to corresponding computing devices. In addition, multiple shared memory devices can be connected to each other. For example, multiple shared memory devices can be directly connected to each other via a memory direct link, and multiple shared memory devices can also be directly connected to corresponding computing devices among multiple computing devices via device direct links. Thus, multiple computing devices and multiple shared memory devices constitute the shared memory devices in this application. Figure 1 The ring communication topology is shown.

[0037] Based on this, a first computing device among multiple computing devices can be connected to a first shared memory device among multiple shared memory devices. In this way, the first computing device can read and write data to the first shared memory device. Similarly, a second computing device among multiple computing devices can be connected to a second shared memory device among multiple shared memory devices. In this way, the second computing device can read and write data to the second shared memory device. However, the embodiments of this application are not limited to this. In the embodiments of this application, the first shared memory device can also read and write data to the second shared memory device to realize data transmission between the first computing device and the second computing device. Similarly, the second shared memory device can also read and write data to the first shared memory device, which will not be elaborated here. For example, the shared memory device can be deployed with a corresponding controller, which can respond to control instructions to control the shared memory device to output the stored data through a port to realize data transmission, as long as the concept of this application can be realized, it will not be elaborated here. For example, a first computing device can directly write first data for a second computing device to a first shared memory device via a direct device link. The first shared memory device can then directly write the first data to a second shared memory device via a direct memory link. The second computing device can then directly read the first data from the second shared memory device via a direct device link. The first data can be data that needs to be sent from the first computing device to the second computing device.

[0038] Based on this, in this application, for a first computing device and a second computing device spaced apart in a ring communication topology, when the first computing device needs to transmit first data to the second computing device, the first computing device can directly write the first data to the first shared memory device via a direct device link. Then, the first shared memory device can directly write the first data to the second shared memory device via a direct memory link, allowing the second computing device to read the first data from the second shared memory device via the direct device link. Thus, by using shared memory devices for data transmission between spaced computing devices in a ring communication topology, this application simplifies the communication path between spaced computing devices in a ring communication topology to a hardware-level access path, achieving communication between spaced computing devices in a ring communication topology through memory transfer. This at least partially solves the transmission delay caused by using switches for data transmission, improving the data transmission speed of the distributed computing system.

[0039] Figure 2 A schematic diagram of a computing system according to another embodiment of this application is shown. It should be noted that, for ease of identification, in... Figure 2 The text uses different shades of gray to show lines that are obscured, and the same applies to the following text, so it will not be repeated here.

[0040] like Figure 2 As shown in the embodiments of this application, any computing device may include multiple computing nodes. For example, the computing device may be a host device such as a server, and the computing node may be a graphics processing unit. Multiple computing nodes may be electrically connected, for example, via a bus. However, it should be understood that this application is not limited thereto. In other embodiments of this application, the computing node may also be other units capable of performing data processing tasks, such as a neural network processing unit.

[0041] Multiple computing devices can be interconnected to form a ring communication topology for data transmission between them. Furthermore, in this ring communication topology, adjacent computing devices can communicate via network interface cards (NICs). Each computing device can have at least two high-performance NICs, the specific number depending on requirements, NIC usage, and NIC virtualization. In this ring communication topology, nodes are computing devices. Edges can be communication links formed by NICs.

[0042] There can be multiple shared memory devices. These devices can communicate via a network interface card (NIC) or an electrical connection via a bus. Since data packet encapsulation and parsing are required during communication via the NIC, this application preferably uses a bus to connect the multiple shared memory devices to ensure communication speed.

[0043] Any shared memory device can connect to multiple computing devices. For example, corresponding computing devices and shared memory devices can be electrically connected via a bus.

[0044] Furthermore, the storage space of any shared memory device may include storage areas. For example, a storage area may include a storage area for storing data of computing devices connected to any shared memory device. However, it should be understood that the embodiments of this application are not limited thereto. In another embodiment of this application, the storage space of the shared memory device may also include a storage area for storing data of other shared memory devices and / or other computing devices besides the computing device corresponding to the shared memory device. Specifically, the storage area of ​​any shared memory device may also include a storage area for storing data of computing devices in the ring communication topology that are not connected to the shared memory device but are connected to other shared memory devices. In yet another embodiment of this application, the storage space of the shared memory device connected to any computing device includes multiple storage areas. The multiple storage areas are respectively configured to store data of multiple computing nodes of any computing device.

[0045] In the embodiments of this application, a computing device or other shared memory device connected to any shared memory device can read and write to the storage area of ​​any shared memory device according to the address of the storage area in any shared memory device. For example, a computing node connected to any shared memory device can read and write to the storage area of ​​any shared memory device according to the address of the storage area in any shared memory device. It should be understood that in a switch solution, when multiple computing devices send data to the same target node through a switch port, all data must first be aggregated into the target port buffer of the switch before being forwarded to the receiver. If the concurrent traffic exceeds the port buffer capacity, it will cause data packet loss and retransmission, thereby causing a surge in latency.

[0046] Specifically, if the high-performance network interface card (NIC) port speed of the computing device (e.g., 100Gbps) is higher than the port speed of the connected switch (e.g., 40Gbps), the amount of data sent by the NIC will exceed the receiving capacity of the switch port, and the excess data will be temporarily stored in the switch's buffer queue. When the buffer queue is full, problems such as "packet loss retransmission" or "head-of-queue blocking" will be triggered, resulting in a sharp increase in data transmission latency. However, in this application, the computing device, as the sender, directly writes the data into an independent block of the shared memory device. The receiver does not need to wait for the data to be forwarded to its local machine through a single switch port, but directly reads the data from the independent blocks of the shared memory device pool. This method is implemented through memory partitioning, rather than port-level data aggregation. Thus, this application can be applied to high-concurrency scenarios and solves the above-mentioned problems.

[0047] Based on this, in one embodiment of this application, the first shared memory device includes multiple first storage regions, each of which is associated with a plurality of shared memory devices. For example, the storage addresses of the multiple first storage regions can be associated with the multiple shared memory devices. Thus, in response to a first computing device writing first data to a target storage region associated with a second shared memory device within the multiple first storage regions, the first shared memory device can directly write the first data to the second shared memory device. Furthermore, the first shared memory device can directly write data stored in any of the first storage regions to another shared memory device associated with that first storage region, achieving direct memory write-to-memory data transfer and improving data transfer speed.

[0048] In another embodiment of this application, the first shared memory may include a first port and a second port. The first port is connected to a first computing device, and the second port is connected to a second shared memory device. The first port and the second port are associated with a target first storage region among multiple first storage regions. For example, the port identifiers of the first port and the second port may be associated with the storage address of the target first storage region and stored in the shared memory device. Thus, in response to the first computing device writing first data to the target storage region via the first port, the first shared memory device may directly write the first data to the second shared memory device via the second port. Based on this, when data is written to a storage region via a certain port, the first shared memory device may directly output the data of that storage region via another port associated with that storage region, thereby directly writing the data to another shared memory device, realizing direct memory write-to-memory transfer of data and improving data transfer speed. However, the embodiments of this application are not limited to this. The aforementioned first storage region may also include multiple sub-regions, and different sub-regions may be associated with different shared memory devices, so that when data is stored in a sub-region, the data stored in the sub-region may be directly written to the shared memory device associated with that sub-region.

[0049] Furthermore, in another embodiment of this application, at least two computing devices connected to any shared memory device are spaced apart from each other in the ring communication topology. For example, the computing devices connected to any shared memory device may be substantially uniformly spaced apart in the ring communication topology, but in some embodiments, the computing devices connected to the shared memory device may not be uniformly spaced in the ring communication topology; this application does not limit this. Further, any shared memory device includes a second storage area associated with at least two computing devices connected to the shared memory device. The association method here is similar to that described above and will not be repeated here. Thus, when one of the at least two computing devices writes second data for the other computing device into the second storage area, the other computing device reads the second data from the second storage area. The second data may be data that needs to be sent by one of the at least two computing devices to the other computing device, etc.

[0050] Thus, a shared memory device can achieve memory transfer between at least two connected computing devices, thereby improving data transfer efficiency. However, the embodiments of this application are not limited to this; the aforementioned storage area may also include multiple sub-regions, and each sub-region may be associated with at least two computing devices, so that when one of the at least two computing devices writes data to that sub-region, the shared memory device can directly write that data to other computing devices associated with that sub-region. The computing devices associated with different sub-regions are different, and will not be elaborated here.

[0051] In this embodiment, the shared memory devices and computing devices can be connected according to the principle of uniform division. For example, any shared memory device has a predetermined number of ports, which correspond to the "heads" of the CXL multi-head device described above. Each port is used to connect to the corresponding computing device, and the number of shared memory devices is equal to the total number of computing devices divided by the predetermined number and rounded up. For example, in one embodiment of this application, if the number of ports of the shared memory devices is 3 and the total number of computing devices is 9, then the number of shared memory devices can be 3. However, it should be understood that this embodiment is not limited to this. In another embodiment of this application, if the number of ports of the shared memory devices is 3 and the total number of computing devices is 8, then the number of shared memory devices can be 3. It should be understood that when 3 is divided by 8, the remainder is 2, that is, there are 2 remaining computing devices. Therefore, a shared memory device with 3 ports should be added to connect the remaining 2 computing devices. In this way, the number of shared memory devices can still be 3, where 3 is the rounded-up value.

[0052] Specifically, multiple shared memory devices each have a port for connecting to a corresponding computing device, and each shared memory device has the same number of such ports, so that the multiple shared memory devices can connect to the same number of computing devices via the same number of ports. For example, in Figure 2 In this configuration, any shared memory device has multiple (e.g., three) ports, which are connected to multiple computing devices (e.g., three computing devices). The purpose is to ensure that any shared memory device is logically connected to the same number of computing devices. This simplifies the management of the ring communication topology. Furthermore, multiple shared memory devices may also have ports for connecting to each other, which will not be elaborated upon here.

[0053] Furthermore, in data processing tasks, communication links between multiple computing devices can be constructed based on multiple shared memory devices. Since each shared memory device has a corresponding port to receive or send data from computing devices or other shared memory devices, data received via any shared memory device's port can be directly stored in the storage area corresponding to that port, or data in the corresponding storage area can be directly output based on read requests received via any shared memory device's port. This at least partially avoids the incast problem caused by using switches for data transmission, improves the redundancy of communication links between multiple computing devices, and increases the data transmission bandwidth of multiple computing devices. Simultaneously, the multiple ports of each shared memory device can perform data transmission in parallel. For example, a first computing device and a third computing device can write data to a first shared memory device in parallel. The first shared memory device can write data from the first computing device and data from the third computing device to a second shared memory device. The second computing device can read data from the first computing device and data from the third computing device from the second shared memory device. Based on this, shared memory devices can transmit data with two or more computing devices, thus improving the utilization of multiple shared memory devices during high-bandwidth transmission. This improves the execution efficiency of data processing tasks. Based on this, this application removes the centralized switch used for data transmission and adopts a direct link of a distributed computing system and a shared memory mechanism of CXL multi-head devices in parallel to solve the congestion problem of switch architecture in many-to-one communication scenarios in some solutions.

[0054] Furthermore, regarding the "uniform partitioning" principle described above, when the number of computing devices is insufficient, this can be addressed by adding virtual nodes. The following will combine... Figure 3 Please provide an explanation.

[0055] Figure 3 A schematic diagram of a computing system according to another embodiment of this application is shown.

[0056] like Figure 3 As shown in the embodiments of this application, multiple computing devices can be communicatively connected to each other to form a ring communication topology for data transmission between the multiple computing devices. Furthermore, in this ring communication topology, adjacent computing devices can also be communicatively connected via network interface cards (NICs). Each computing device may include multiple connected computing nodes.

[0057] There can be multiple shared memory devices. Any shared memory device can connect to multiple computing devices. For example, corresponding computing devices and shared memory devices can be electrically connected via a bus.

[0058] In this embodiment, the shared memory device and the computing device are still connected according to the principle of uniform division. However, in this embodiment, the number of ports of the shared memory device is 3, and the total number of computing devices is 8, so the number of shared memory devices can be 3. It should be understood that when 3 is divided by 8, the remainder is 2, that is, there are 2 remaining computing devices. Therefore, a shared memory device with 3 ports should be added to connect the remaining 2 computing devices.

[0059] At this point, for the third shared memory device, one port is still not connected to a computing device. In this case, the third port of the shared memory device can be configured to connect to a virtual node. This virtual node is not the actual computing device described earlier. That is, even when the third shared memory device is configured to connect to a virtual node, the aforementioned port of the third shared memory device is still not actually connected to a computing device. Figure 3 In the illustrated ring communication topology, the computing devices located on either side of the virtual node (or virtual device), i.e., the two computing devices adjacent to the virtual device, can directly communicate with each other. In one configuration of the virtual node in this application, logical configuration can be performed via software to change the writing of data that should be written to the virtual node to other computing devices, or to change the reading of data that should be read from the virtual node to other computing devices. However, it should be understood that the embodiments of this application are not limited to this; virtual nodes can also be added in other ways, which will not be elaborated here. This simplifies the process of constructing the ring communication topology.

[0060] Furthermore, in this embodiment, at least one of the following communication connections may be included: a communication connection between at least two shared memory devices via a network interface card (NIC) or a bus; a communication connection between corresponding computing devices and shared memory devices via a bus; a communication connection between computing devices via a NIC; and a communication connection between computing nodes within the same computing device via a bus. The bus may support the CXL protocol.

[0061] Building upon this, multiple shared memory devices are connected to each other via direct memory links, which support memory interconnect protocols. Multiple computing devices are connected to each other via direct host links. Direct host links support protocols different from memory interconnect protocols. For example, direct host links can work collaboratively using multiple layers of the TCP / IP protocol stack. Any computing device and its corresponding shared memory device are connected via direct device links. Direct device links support memory interconnect protocols.

[0062] Thus, in the event that communication between at least two shared memory devices is interrupted, these two interrupted shared memory devices can communicate via an alternative link, as described below. Figure 4A Please provide an explanation.

[0063] Figure 4A A schematic diagram of a computing system according to another embodiment of this application is shown. It should be noted that in this application, intersecting short lines (e.g., "+", but not limited to, "×") represent faulty lines, and arrows indicate the direction of data transmission; the same applies below.

[0064] like Figure 4A As shown, in this embodiment, when the direct memory link between the first shared memory device and the second shared memory device fails, the first computing device and the second computing device are configured to transmit data between each other via a direct device link between the first computing device and the first shared memory device, another direct memory link between the first shared memory device and the third shared memory device, yet another direct memory link between the third shared memory device and the second shared memory device, and yet another direct device link between the second shared memory device and the second computing device.

[0065] For example, the computing node of the first computing device can send a read instruction to the first shared memory device. Under the control of the read instruction, the first shared memory device can send an instruction to the third shared memory device, so that the third shared memory device reads the data indicated by the instruction from the second shared memory device, and then writes the data to the first shared memory device.

[0066] For example, upon receiving a read instruction, the first shared memory device can determine whether the direct memory link between itself and the second shared memory device is faulty based on the read address indicated by the read instruction. If the direct memory link between the first and second shared memory devices is determined to be faulty, it can then send an instruction carrying the memory address of the second shared memory device to the third shared memory device. This allows the third shared memory device to read data from the second shared memory device via the direct memory link between itself and the second shared memory device. The data writing process is similar, requiring only the computing node of the first computing device to send a write instruction, and will not be elaborated upon here.

[0067] Conversely, the computing node of the first computing device can read from and write to the first shared memory device via another direct memory link between the third and second shared memory devices, and another direct memory link between the first and third shared memory devices. The specific process is similar to that described above and will not be repeated here.

[0068] Furthermore, the embodiments of this application are not limited to this. In another embodiment of this application, the first computing device may also perform data read and write to the second shared memory device via a host direct link between the first computing device and the third computing device, and a device direct link between the third computing device and the third shared memory device.

[0069] Furthermore, the embodiments of this application are not limited to this. In other embodiments of this application, when communication between corresponding computing devices and shared memory devices is disconnected, the corresponding computing devices and shared memory devices that have disconnected communication can communicate via another link. The following is in conjunction with... Figure 4B Please provide an explanation.

[0070] Figure 4B A schematic diagram of a computing system according to another embodiment of this application is shown.

[0071] like Figure 4B As shown, in this embodiment, when the direct device link between the first computing device and the first shared memory device fails, the computing node of the first computing device performs data read and write operations with the first shared memory device via the host direct link between the first computing device and the third computing device, the direct device link between the third computing device and the corresponding third shared memory device, and the memory direct link between the third shared memory device and the first shared memory device.

[0072] Specifically, the first computing device can send a read command to the third computing device via a host direct link between the first and third computing devices. This allows the third computing device to send a read command carrying the memory address of the first shared memory device to the third shared memory device via a device direct link between the third computing device and the corresponding third shared memory device. This enables the third shared memory device to read data from the first shared memory device via a memory direct link between the third and first shared memory devices. The data writing process from the first computing device to the first shared memory device is similar and will not be described in detail here.

[0073] Furthermore, the embodiments of this application are not limited to this. In other embodiments of this application, when communication between computing devices is interrupted, the disconnected computing devices can communicate via another link. The following is in conjunction with... Figure 4C Please provide an explanation.

[0074] Figure 4C A schematic diagram of a computing system according to another embodiment of this application is shown.

[0075] like Figure 4CAs shown, in this embodiment, when the host direct link between the third computing device and the fourth computing device, which are adjacent to each other in the ring communication topology, fails, the third computing device and the fourth computing device transmit data to each other via the device direct link between the third computing device and the corresponding third shared memory device, the memory direct link between the third shared memory device and the corresponding fourth shared memory device of the fourth computing device, and another device direct link between the fourth shared memory device and the fourth computing device.

[0076] For example, a third computing device can send a data transfer command to a third shared memory device via a direct device link between the third computing device and a corresponding third shared memory device. This enables the third shared memory device to send a data transfer command to a fourth shared memory device, and then the fourth shared memory device can transmit the data carried by the data transfer command to the fourth computing device. The reverse is also true, and will not be elaborated further here.

[0077] Furthermore, in this embodiment, any computing device can be equipped with multiple high-performance network interface cards (NICs). For example, there can be four NICs, but it is not limited to this. The NICs are also duplex and support virtualization. Based on this, multiple direct host links can be configured between adjacent computing devices. The third and fourth computing devices are configured to switch to other communication links among the multiple direct host links in the event of a failure of the current communication link. The multiple direct host links can include a primary link and backup links, where the primary link can be the link with the highest communication priority and used for daily communication. In the event of a failure of the primary link used for communication among the multiple direct host links, adjacent computing devices can automatically switch to another unfailed backup link for communication.

[0078] It should be noted that the link switching method described above in this application is only an example. In fact, communication can also be based on other links described above in this application, which will not be listed here. It should also be noted that for communication between shared memory devices, and for communication between computing devices and corresponding shared memory devices, bus links are preferred to minimize latency. Furthermore, this application supports parallel read / write operations on shared memory devices by multiple computing devices.

[0079] Based on the above, the ring communication topology of this application has rich fault tolerance mechanisms, which can ensure that in the event of a faulty communication link, other communication links will automatically take over the tasks originally performed by the faulty link. This at least partially guarantees that the transfer of memory data will not be interrupted, and at least partially guarantees communication continuity, thereby improving the fault tolerance and reliability of the system. Compared with some solutions that use multi-level switches for data forwarding, the solution of this application can reduce the overall communication latency of the computing system by 1 to 2 orders of magnitude.

[0080] Furthermore, this application can also achieve elastic expansion of the transmission links of the ring communication topology by increasing the number of shared memory devices, thereby improving link fault tolerance. Specifically, the ring communication topology in this application has high structural flexibility and high scalability. Thus, by rationally planning the direct connection links between shared memory devices and computing devices, the scale of the computing cluster can be easily expanded without significantly increasing system complexity and cost.

[0081] The architecture of the computing system of this application has been described above. The following will focus on the specific data processing tasks performed by the computing system. For example, data processing tasks may include model training tasks. For instance, model training tasks can be implemented based on the 2D-Torus AllReduce algorithm. The computing system is configured to execute model training tasks, and each computing node of any computing device obtains its own gradient data in the model training task. The gradient data represents the model parameters used to optimize the model. The following will combine... Figure 5 Please provide an explanation.

[0082] Figure 5 A schematic diagram of a computing network based on a computing system according to an embodiment of this application is shown.

[0083] exist Figure 5 The diagram illustrates a computing network based on compute nodes and shared memory devices. For example, the computing system may include nine compute devices, each of which may include four compute nodes. Based on this, refer to... Figure 5 As can be seen, each row corresponds to a computing device, and each column corresponds to a group of computing nodes. In the embodiments of this application, computing nodes that are spaced apart from each other can communicate via shared memory devices, and computing nodes that are adjacent to each other can communicate directly.

[0084] Specifically, in this embodiment of the application, gradient data on multiple computing nodes of any computing device can be locally reduced to obtain local gradient data for each of the multiple computing devices. This process corresponds to... Figure 5 The diagram shows the lateral computation process of the computational network.

[0085] Then, global reduction can be performed on the local gradient data of multiple computing devices, resulting in global gradient data on each device. The first and second computing devices are spaced apart in a ring communication topology and transmit their local gradient data to each other via a first shared memory device and a second shared memory device. For example, the aforementioned first data may include the local gradient data of the first computing device. This process corresponds to... Figure 5 The diagram shows the vertical computation process of the computational network.

[0086] Specifically, the gradient data on any computing node is divided into a predetermined number of gradient blocks. For example, the data dimension ranges corresponding to the predetermined number of gradient blocks are different from each other. Multiple computing nodes within the same computing device can correspond to different data dimension ranges. Based on this, the data dimension range of the local computing data reduced during topology reduction is the same. It should be noted that the predetermined number here corresponds to the number of computing nodes, which is different from the predetermined number of ports of the shared memory device described above. For ease of distinction, they can be referred to as the first predetermined number or the second predetermined number, respectively.

[0087] Furthermore, the local specification may include: passing gradient blocks between multiple computing nodes within any computing device, such that each node has a local gradient block of the corresponding dimension, which can be used to compute the corresponding local gradient data. Other similar considerations are not elaborated further. For example, the same computing device may include computing node 11, computing node 12, computing node 13, and computing node 14. Specifically, computing node 11, computing node 12, computing node 13, and computing node 14 can be... Figure 5 The four computing nodes are located in the same row of the computing network.

[0088] Based on this, compute node 11 can send the first-dimensional gradient block to compute node 12; compute node 12 can perform calculations on the received first-dimensional gradient block and the locally stored first-dimensional gradient block (according to a pre-defined calculation method, not limited here, and the same applies below) to obtain the calculated first-dimensional gradient block, and then send the calculated first-dimensional gradient block to compute node 13; compute node 13 can perform calculations on the received first-dimensional gradient block and the locally stored first-dimensional gradient block to obtain the calculated first-dimensional gradient block, and then send the calculated first-dimensional gradient block to compute node 14; compute node 14 can perform calculations on the received first-dimensional gradient block and the locally stored first-dimensional gradient block to obtain the calculated first-dimensional gradient block, and then send the calculated first-dimensional gradient block to compute node 11. Thus, compute node 11 can use the calculated first-dimensional gradient block received from compute node 14 as its locally reduced local gradient block. The gradient blocks for the corresponding dimensions of other computing nodes can be calculated in a similar manner, which will not be elaborated here. In one embodiment, after splitting the gradients on N computing nodes within the computing device into N smaller blocks and iteratively reducing them, after iteration, the gradient block of any computing node will have a complete gradient of the same dimension. The gradient block of any computing node's dimension will contain the sum of all gradients corresponding to that block across N computing nodes. N is a positive integer greater than 1, corresponding to 4 in the above embodiment.

[0089] Furthermore, the global specification can include: passing gradient blocks of the same dimension between corresponding nodes with gradient blocks of the same dimension in the respective computing nodes of multiple computing devices, so that any computing node in any computing device has a global gradient block of the corresponding dimension. For example, the nine computing devices can respectively include computing node 11, computing node 21, computing node 31, computing node 41, computing node 51, computing node 61, computing node 71, computing node 81, and computing node 91 (hereinafter simply referred to as "computing node 11 to computing node 91"), and will not be elaborated further. Computing nodes 11 to 91 can be in Figure 5 Nine computation nodes are located in the same column of the computation network. Furthermore, computation nodes 11 to 91 can have local gradient blocks of the same dimension, specifically, local gradient blocks of the first dimension. For example, computation node 11 can send its local gradient block to other computation nodes, and in a similar manner to the above, the gradient block is passed and computed among computation nodes 11 to 91 to obtain a global gradient block of the first dimension. The global gradient block of the first dimension can then be returned from computation node 91 to computation node 11. This iterative process completes the reduction of global gradient blocks across various data dimensions.

[0090] Based on this, for corresponding nodes with gradient blocks of the same dimension that are spaced apart within the computing nodes of multiple computing devices, the gradient blocks of the same dimension can be transmitted via the corresponding shared memory device. For corresponding nodes with gradient blocks of the same dimension that are adjacent to each other within the computing nodes of multiple computing devices, the gradient blocks of the same dimension can be transmitted directly, so that any computing node within any computing device has a global gradient block of the corresponding dimension. In this way, this application performs a vertical global reduction on the data on the computing nodes of any computing device within the cluster, so that the gradient block of the dimension where any computing node is located includes the sum of all gradients corresponding to that block in the vertical computing node. Furthermore, by using different communication modes (which can be called hybrid modes) based on the spacing relationship between nodes as described above, the latency can be compressed to 300~400 nanoseconds.

[0091] After obtaining the global gradient blocks, the model training task can further include globally aggregating the global gradient blocks of corresponding dimensions from multiple computing nodes on any computing device to obtain the final gradient information. Specifically, multiple computing nodes on the same computing device can broadcast their corresponding dimension global gradient blocks within the device, so that the same computing node has global gradient blocks covering multiple data dimensions. This allows for the aggregation of global gradient blocks, ensuring that each computing node possesses the final gradient information. Thus, this application achieves the global aggregation of gradients from the gradient block of any computing node's dimension onto the corresponding dimension gradient block within the computing device. Based on this, the model training task can be completed by iterating over the above process.

[0092] Building upon this, for the aforementioned processes, especially the high-concurrency and high-bandwidth global reduction and global convergence processes, since the computing nodes of the separated computing devices can communicate via corresponding shared memory devices, and the computing nodes of adjacent computing devices can communicate directly, the problems caused by data transmission through switches can be at least partially avoided. This optimizes the transmission path of the distributed algorithm, reduces data synchronization latency and switch hop count in the aforementioned computing process, improves communication efficiency, and accelerates algorithm convergence. Furthermore, this application overcomes the congestion bottleneck of switch architectures; compared to switch architectures, the ring communication topology of this application exhibits a significantly improved communication gain.

[0093] Building upon this foundation, shared memory devices are applied to rack-level server solutions, making them ideal for small to medium-sized high-performance clusters, such as memory-intensive workloads in single-rack or inter-rack artificial intelligence and cloud computing, maximizing the transfer speed between computing devices. Shared memory devices are used to directly interconnect two isolated computing devices, creating an efficient and low-cost memory sharing solution. As a data transfer path for cross-node data exchange, shared memory devices support parallel read / write operations across multiple computing devices, improving model training efficiency.

[0094] For small- to medium-sized AI training clusters, each rack contains eight servers, each equipped with eight image processing units (also known as computing cards or accelerator cards). These clusters need to support distributed training of large language models, requiring low-latency gradient synchronization (such as the reduction and convergence operations mentioned above) and high-bandwidth data interaction, while avoiding congestion issues caused by switches. Based on this, the computing system of this application distributes concurrent requests across direct links and shared memory devices, thereby meeting the above requirements and at least partially solving the incast congestion problem caused by switches. In the computing system of this application, multiple computing devices can adopt a direct-connection design. Each computing device is equipped with two high-performance network cards, forming a closed-loop topology through direct network card connections. That is, any computing device is directly connected to two adjacent computing devices, achieving "direct-connection ring" communication between computing devices without switches.

[0095] For servers spanning multiple racks (applicable to scenarios such as hybrid AI training and cloud computing workloads), it is necessary to support dynamic expansion of the cluster size and have link failure tolerance capabilities, meaning that the failure of a single link does not affect overall communication. The computing system in this application also meets these requirements.

[0096] Furthermore, distributed database clusters for high-frequency trading systems require millisecond-level data synchronization (such as order information and account balances), demanding high-concurrency read / write operations and low-latency data consistency guarantees. The computing system in this application also meets these requirements. For example, six computing devices can be directly connected in a ring via their respective deployed single high-performance network cards, enabling low-latency direct communication between nodes (e.g., synchronizing order data between adjacent computing devices). Additionally, two 4-port CXL multi-head devices can be used as a global shared memory pool. The multi-head devices are connected to the computing devices via the CXL bus, supporting parallel read / write of core data such as database logs and indexes. Based on this, real-time transaction data generated by order interactions between adjacent computing devices is directly transmitted through the computing device direct-connection ring, with latency controlled to the microsecond level. Moreover, during cross-node transaction submission, global data consistency verification can be completed through shared memory implemented using CXL multi-head devices. After data is written to the shared memory device, all computing devices can directly read it, reducing TCP / IP protocol stack latency; for example, this application compresses the latency from 2-4 microseconds to within 300 nanoseconds. Based on this, this application can reduce the commit latency of distributed transactions, meet the real-time requirements of high-frequency transactions, and avoid the caching bottleneck of the switch, thus avoiding the head-of-line blocking problem caused by concurrent write requests from multiple nodes.

[0097] Furthermore, in some application switch solutions, the total communication delay between nodes consists of multi-level switch forwarding delay, protocol stack delay, etc., as shown in the following formula (1):

[0098] T_switching scheme = N_hop count × t_single-hop switching + t_protocol stack + t_congestion wait (1)

[0099] Wherein, T_switch scheme represents the total communication latency between nodes. N_hop count represents the number of switch stages through which data transmission passes, such as 2-3 hops in some architectures. t_single-hop switching represents the forwarding latency of a single switch. t_protocol stack represents the processing latency of the network protocol stack. t_congestion wait represents the queuing latency caused by incast problems, port contention, buffer depth, etc., and it increases significantly with the increase of cluster size.

[0100] Based on this, some solutions also involve a congestion model, as shown in the following formula (2):

[0101] P_congestion = f(M, B, C) (2)

[0102] Where P_congestion represents the probability of the switch becoming congested. M represents the number of concurrent requests to the switch, such as many-to-one requests in a parameter server architecture. B represents the port bandwidth of the switch. C represents the buffer depth of the switch. f indicates that M, B, and C are related to P_congestion.

[0103] Based on this, this application, in contrast to the above-mentioned scheme, states:

[0104] Regarding the "N_hop count × t_single-hop exchange" part, this application can reduce the number of hops in the data transmission process to 1-2 hops. Furthermore, the data in this application passes through multi-head devices or directly connected devices that support high concurrency during transmission. Moreover, the latency of the ring-shaped direct communication between computing devices is reduced, with no switch required and only physical link latency, which is close to the hardware-level latency. The physical link includes the link connected to the shared memory device.

[0105] For the "t_protocol stack" part, this application reduces transmission latency by utilizing the transmission latency of the shared memory device, eliminating the need for protocol stack forwarding.

[0106] Regarding the "t_congestion wait" part, even if the port bandwidth of the shared device in this application is the same as the port bandwidth of the aforementioned switch, the computing system of this application will reduce the probability of port contention in scenarios with concurrent requests. In addition, since the switch has a limited cache depth, while the shared memory device has cost and storage advantages, the impact of cache depth is basically non-existent. Therefore, the probability of congestion wait and the waiting time can be further reduced.

[0107] The improvement effect can be represented by the following formula (3):

[0108] T_Improvement = T_Switch Scheme - T_This Application (3)

[0109] Wherein, T_improvement indicates the reduction in latency compared to the aforementioned switching scheme. T_switching scheme can be found in the preceding description. T_this application indicates the data transmission latency of this application.

[0110] As cluster size increases and communication frequency rises, the problems with switch solutions and latency become more prominent, while the gains from improved solutions will increase significantly. Specifically, in distributed training, taking the algorithm mentioned above as an example, the total synchronization time is a key performance indicator. Based on this, assuming the algorithm has K iterations, the overall improvement is K×T_improvement.

[0111] Based on this, in the model training task, this application can use the Compute ExpressLink (CXL) technology to divide the model or data into multiple nodes, so as to quickly realize cross-node data copying with the help of high-speed memory copying, thereby quickly copying parameters and other data from one node to other nodes and reducing the waiting time during the training process.

[0112] Furthermore, this application fully leverages the inherent communication latency compression advantages of CXL technology, employs the CXL multi-device memory transfer mechanism to construct a distributed computing network solution, and specifically designs a ring communication topology adapted to this mechanism. This topology, serving as the core transmission link for data exchange between different nodes, significantly improves the overall data transmission efficiency through the direct memory transfer mode implemented by the CXL multi-device.

[0113] The core advantage of this architecture lies in its complete elimination of the additional latency inherent in traditional switch-based network architectures, such as queuing and store-and-forward operations. On one hand, hosts form a direct communication path through a ring-shaped direct connection, eliminating the need for data transmission through multiple levels of switches and ensuring the shortest path at the physical link level. On the other hand, CXL multi-head devices achieve cross-node data interaction through a shared memory pool, simplifying the communication path to hardware-level memory access and further reducing transmission latency. This dual design together meets the stringent real-time requirements of distributed computing scenarios, providing a more efficient and low-latency underlying network solution, especially for high-frequency communication scenarios such as large-scale distributed machine learning and cloud computing.

[0114] Thus, based on the CXL multi-head device topology scheme, this application can realize a switchless computing network architecture. Firstly, it significantly reduces communication latency and improves data transmission efficiency. By reducing switch forwarding latency and combining the advantages of direct connection and architectural design, it can optimize synchronous operation performance in algorithm applications (such as AllReduce). Secondly, it overcomes the hardware limitations of traditional switches, solving congestion and bottleneck problems in traditional architectures while maintaining good network topology scalability. Furthermore, the link redundancy design and flexible topology adaptation can be extended to improve system fault tolerance and stability, adapting to high-performance computing scenarios such as AI training in data centers, thus broadening its application scope.

[0115] Based on this, this application reduces overall system latency through the memory transfer mode of CXL multi-head devices. It uses the shared memory pool of multi-head devices as the core path for cross-node data exchange, providing larger cache capacity and supporting parallel read / write operations on shared memory by multiple hosts. This achieves "hardware-level memory access" data transmission, replacing the traditional forwarding mechanism of switches that relies on network protocol stacks and limited cache. Furthermore, based on the communication mode and fault-tolerant optimization mechanism of this application, the ring-shaped direct communication between hosts (reducing intermediate media latency) ensures low-latency direct communication between adjacent nodes, while the memory transfer communication mediated by multi-head devices (avoiding switch latency and congestion, reducing protocol stack latency) provides cross-device communication efficiency and link redundancy. Moreover, the fully connected design of the shared memory devices provides abundant links, ensuring strong fault tolerance and high bandwidth utilization, supporting path switching in case of failure; it is compatible with distributed algorithms such as 2D-Torus AllReduce, and improves the communication efficiency of operations such as global reduction and aggregation through topology optimization.

[0116] Therefore, the computing system of this application can be used in any application that requires memory expansion and sharing, including scenarios involving artificial intelligence, cloud computing, distributed databases, scientific computing, big data, and other scenarios that require distributed high-performance data computing and transmission.

[0117] In addition, this application also provides a calculation method based on the above-mentioned calculation system.

[0118] Figure 6 A flowchart of a calculation method according to an embodiment of this application is shown.

[0119] like Figure 6 As shown, the calculation method of this embodiment may include operations S610~S620.

[0120] The S610 is used to transfer data between multiple computing devices.

[0121] When operating the S620, data processing tasks are performed based on the transmitted data to obtain the data processing results.

[0122] In this embodiment, any computing device is directly connected to two other computing devices among multiple computing devices to form a ring communication topology based on multiple computing devices. Adjacent computing devices in the ring communication topology transmit data through direct connections. The computing system also includes multiple shared memory devices directly connected to each other via direct memory links. These shared memory devices are respectively directly connected to corresponding computing devices among the multiple computing devices via direct device links. The multiple computing devices include a first computing device and a second computing device spaced apart in the ring communication topology. The multiple shared memory devices include a first shared memory device connected to the first computing device and a second shared memory device connected to the second computing device. This allows the first computing device to directly write first data for the second computing device to the first shared memory device via the direct device link, the first shared memory device to directly write the first data to the second shared memory device via the direct memory link, and the second computing device to directly read the first data from the second shared memory device via the direct device link. However, it should be understood that the computing system of this application can also implement any of the methods described above, which will not be elaborated upon here.

[0123] For example, data processing tasks include model training tasks; data processing results include gradient data; gradient data represents the model parameters used to optimize the model; multiple computing devices transmit data via multiple shared memory devices and perform data processing tasks based on the transmitted data to obtain data processing results, including: any computing node of any computing device performs a model training task to obtain the corresponding gradient data.

[0124] For example, any computing node of any computing device performs a model training task to obtain corresponding gradient data, including: performing local reduction on the gradient data on multiple computing nodes of the computing device within the computing device to obtain local gradient data for each of the multiple computing devices; and performing global reduction on the local gradient data of each of the multiple computing devices to obtain global gradient data for each of the multiple computing devices, wherein the first computing device and the second computing device among the multiple computing devices are spaced apart from each other in a ring communication topology, and transmit their respective local gradient data to each other through a first shared memory device and a second shared memory device.

[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0126] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

[0127] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.

Claims

1. A computing system, characterized in that, include: Multiple computing devices, wherein any one of the computing devices is directly connected to two other computing devices among the multiple computing devices via a host direct link, to form a ring communication topology based on the multiple computing devices, wherein adjacent computing devices in the ring communication topology transmit data through a host direct link between them; Multiple shared memory devices are directly connected to each other via direct memory links, and each of these shared memory devices is directly connected to a corresponding computing device among the multiple computing devices via a direct device link. The plurality of computing devices includes a first computing device and a second computing device spaced apart in the ring communication topology, and the plurality of shared memory devices includes a first shared memory device directly connected to the first computing device and a second shared memory device directly connected to the second computing device. The first computing device is configured to directly write first data for the second computing device to the first shared memory device via a device direct link, the first shared memory device is configured to directly write the first data to the second shared memory device via a memory direct link, and the second computing device is configured to directly read the first data from the second shared memory device via a device direct link.

2. The computing system according to claim 1, characterized in that, The first shared memory device includes a plurality of first storage areas, and the plurality of first storage areas are respectively associated with the plurality of shared memory devices; The first shared memory device is further configured to directly write the first data into the second shared memory device in response to the first computing device writing the first data into a target first storage area associated with the second shared memory device among the plurality of first storage areas.

3. The computing system according to claim 2, characterized in that, The first shared memory device includes a first port and a second port; the first port is directly connected to the first computing device, the second port is directly connected to the second shared memory device, and the target first storage area is associated with the first port and the second port; The first shared memory device is further configured to write the first data directly to the second shared memory device via the second port in response to the first computing device writing the first data to the target first storage area via the first port.

4. The computing system according to any one of claims 1 to 3, characterized in that, In the ring communication topology, at least two computing devices directly connected to any shared memory device are spaced apart from each other. The shared memory device includes a second storage area associated with at least two computing devices directly connected to the shared memory device, such that if one of the at least two computing devices writes second data for the other computing device to the second storage area, the other computing device reads the second data from the second storage area.

5. The computing system according to any one of claims 1 to 3, characterized in that, The plurality of computing devices also includes a third computing device connected to the first shared memory device; The first computing device and the third computing device are configured to write data to the first shared memory device in parallel; The first shared memory device is further configured to write data from the first computing device and data from the third computing device into the second shared memory device; The second computing device is configured to read data from the first computing device and data from the third computing device from the second shared memory device.

6. The computing system according to any one of claims 1 to 3, characterized in that, The computing system is configured to perform a model training task. Each computing device includes multiple computing nodes. Each computing node of any computing device obtains corresponding gradient data in the model training task. The gradient data represents the model parameters used to optimize the model.

7. The computing system according to claim 6, characterized in that, Within any computing device, the gradient data on multiple computing nodes of the computing device are locally reduced to obtain the local gradient data of each of the multiple computing devices. Global reduction is performed on the local gradient data of each of the plurality of computing devices among the plurality of computing devices, and global gradient data is obtained on the plurality of computing devices respectively, wherein the first data includes the local gradient data of the first computing device.

8. The computing system according to claim 7, characterized in that, The gradient data on any computing node is divided into a predetermined number of gradient blocks, and the local specification includes: transferring the gradient blocks between multiple computing nodes within any computing device, so that any computing node has local gradient blocks of a corresponding dimension. The global reduction includes: passing the gradient blocks of the same dimension between corresponding computing nodes that have gradient blocks of the same dimension in the computing nodes of the multiple computing devices, so that any computing node in any computing device has a global gradient block of the corresponding dimension; the model training task also includes: globally converging the global gradient blocks of the corresponding dimensions of the multiple computing nodes of the computing device in any computing device to obtain the final gradient data.

9. The computing system according to any one of claims 1 to 3, characterized in that, In the event of a failure of the direct memory link between the first shared memory device and the second shared memory device, the first shared memory device writes the first data to the second shared memory device via another direct memory link between the first shared memory device and the third shared memory device, and yet another direct memory link between the third shared memory device and the second shared memory device.

10. The computing system according to any one of claims 1 to 3, characterized in that, In the event of a failure of the direct device link between the first computing device and the first shared memory device, the first computing device performs data read and write operations on the first shared memory device via the host direct link between the first computing device and the third computing device, the direct device link between the third computing device and the corresponding third shared memory device, and the memory direct link between the third shared memory device and the first shared memory device.

11. The computing system according to any one of claims 1 to 3, characterized in that, In the event of a failure of the host direct link between the third and fourth computing devices, which are adjacent to each other in the ring communication topology, the third and fourth computing devices are configured to transmit data to each other via a device direct link between the third computing device and a corresponding third shared memory device, a memory direct link between the third shared memory device and a corresponding fourth shared memory device of the fourth computing device, and another device direct link between the fourth shared memory device and the fourth computing device.

12. The computing system according to any one of claims 1 to 3, characterized in that, In the ring communication topology, there are multiple host direct links between the third and fourth computing devices that are adjacent to each other, and the third and fourth computing devices are configured to switch to another host direct link among the multiple host direct links for communication in the event of a failure of the current host direct link.

13. The computing system according to any one of claims 1 to 3, characterized in that, The memory direct link includes a connection via a network card or a bus; The device direct connection link includes a connection via a bus; The host direct connection link includes a connection via a network interface card; Computing nodes within the same computing device are connected via a bus.

14. A calculation method, characterized in that, The calculation method, applied to a computing system, includes: Data is transmitted between multiple computing devices in the computing system, and data processing tasks are performed based on the transmitted data to obtain data processing results. In this configuration, any computing device is directly connected to two other computing devices among the plurality of computing devices via a host direct link to form a ring communication topology based on the plurality of computing devices. In this ring communication topology, adjacent computing devices are directly connected to each other via host direct links for data transmission. The computing system also includes multiple shared memory devices that are directly connected to each other via direct memory links. Each of the multiple shared memory devices is directly connected to a corresponding computing device among the multiple computing devices via a direct device link. The plurality of computing devices include a first computing device and a second computing device spaced apart in the ring communication topology. The plurality of shared memory devices include a first shared memory device connected to the first computing device and a second shared memory device connected to the second computing device, such that the first computing device directly writes first data for the second computing device to the first shared memory device via a direct device link, the first shared memory device directly writes the first data to the second shared memory device via a direct memory link, and the second computing device directly reads the first data from the second shared memory device via a direct device link.

15. The calculation method according to claim 14, characterized in that, The data processing task includes a model training task; the data processing result includes gradient data; the gradient data represents the model parameters used to optimize the model; Data is transferred between multiple computing devices in the computing system, and data processing tasks are performed based on the transferred data to obtain data processing results, including: Any computing node of any of the aforementioned computing devices executes the model training task to obtain the corresponding gradient data.

Citation Information

Patent Citations

  • Inter-domain link fast failure recovery method based on software defined network

    CN103428031A

  • Model training method, computing device and system

    CN118278540A

  • Memory access method, computing system and electronic equipment

    CN118331922A

  • Intelligent visual management method and system for enterprise big data

    CN120144416A

  • Memory performance detection method, device and equipment and memory medium

    CN120540926A

Cited By

  • Communication method and device of computing cluster, equipment, storage medium and program product

    CN122132190A