A hardware synchronization interface for a supercomputer cluster

CN122654061APending Publication Date: 2026-08-28陈立波
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610395014.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-29
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

本发明的目的在于克服现有技术的缺陷,提供一种智算中心集群硬件同步接口,解决现有集群同步依赖网络协议、软件解析延迟高、同步精度低、易受网络波动影响、可靠性不足的技术问题,实现无网络协议依赖的、纳秒级高精度的硬件级全域集群同步

Benefits of technology

1. 实现了无网络协议、无软件依赖的硬件级全域同步,同步信号为原生硬件脉冲,无任何网络协议封装、软件解析环节,大幅降低了协议封装、软件处理带来的延迟,同步信号传输延迟低至1纳秒以内,较现有精密时间协议方案提升了6个数量级以上,可满足人工智能大模型训练纳秒级同步精度的需求;

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The application discloses a kind of wisdom calculation center cluster hardware synchronous interface, belong to wisdom calculation center cluster management and control technical field, solve the existing cluster synchronization dependence network protocol, software analysis delay is high, low synchronous accuracy, easily influenced by network fluctuation Problem.The interface is exclusive physical interface, independent of all network synchronization protocol, specially designed for wisdom calculation center multi-node cluster;Core includes global synchronization hardware pulse generation circuit, point-to-point cascaded hardware transmission port, synchronization hardware calibration module, cluster node hardware docking unit;Synchronization signal is original hardware pulse, without any protocol package, software analysis link;Through point-to-point hardware cascade, realize global delay-free transmission, synchronization calibration is automatically completed by fixed hardware logic, without software calibration link.The application realizes the hardware level global cluster synchronization without network protocol dependence, and the synchronization accuracy can be controlled within 1 nanosecond, is not easily influenced by network fluctuation, system load, can greatly improve the consistency and efficiency of wisdom calculation center cluster parallel computing, suitable for wisdom calculation center, supercomputing center multi-node cluster Global synchronization management and control scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent computing center cluster management and control technology, specifically involving a hardware synchronization interface for intelligent computing center clusters that is independent of network protocols and provides hardware-level full-domain synchronization. Background Technology

[0002] The training of large-scale AI models and high-performance parallel computing scenarios in intelligent computing centers require high-precision global synchronization across hundreds or even thousands of computing nodes within the cluster, including clock synchronization, task synchronization, and state synchronization. The synchronization accuracy directly determines the efficiency and consistency of parallel computing results. Current technologies for intelligent computing center cluster synchronization rely on network protocols and software architectures, including network time protocols, precise time protocols, and remote direct data access network synchronization. These technologies suffer from the following unresolved technical limitations: 1. High synchronization latency due to reliance on network protocols and software parsing: Synchronization signals need to go through multiple stages such as network protocol encapsulation, software parsing, routing and forwarding, and kernel processing, resulting in synchronization latency as high as milliseconds, which cannot meet the nanosecond-level synchronization accuracy requirements for training large artificial intelligence models. 2. Insufficient synchronization accuracy and susceptibility to external factors: Synchronization accuracy depends on software clock resolution and network link quality, and is easily affected by network fluctuations, bandwidth congestion, system load, and process scheduling. Synchronization deviation can reach tens of microseconds, making it impossible to achieve high-precision and consistent synchronization across the entire cluster, resulting in deviations in parallel computing tasks and a decrease in training efficiency. 3. Complex architecture and insufficient reliability: Existing synchronization solutions rely on a large number of devices such as network switches, routers, and servers, resulting in a complex architecture. Failure of any node or link will affect the global synchronization process and may even cause the cluster synchronization to crash, making it impossible to achieve long-term stable operation. 4. No true hardware-level global synchronization solution: All existing publicly available solutions cannot be separated from the participation of network protocols, software, and firmware. They cannot achieve pure hardware-based synchronization signal generation, transmission, calibration, and execution throughout the entire process, and cannot fundamentally solve the problems of high latency, low accuracy, and poor reliability of existing solutions. Summary of the Invention

[0003] Technical problems to be solved The purpose of this invention is to overcome the shortcomings of the prior art and provide a hardware synchronization interface for intelligent computing center clusters. This interface solves the technical problems of existing cluster synchronization, such as reliance on network protocols, high software parsing latency, low synchronization accuracy, susceptibility to network fluctuations, and insufficient reliability. It enables high-precision, nanosecond-level hardware-level global cluster synchronization without network protocol dependence. Technical solution

[0004] To achieve the above objectives, the present invention adopts the following technical solution: A hardware synchronization interface for a smart computing center cluster includes a global synchronization hardware pulse generator circuit, a point-to-point cascaded hardware transmission port, a synchronization hardware calibration module, and a cluster node hardware docking unit. The interface is a physical interface specifically designed for multi-rack, multi-computing-power node clusters in smart computing centers and supercomputing centers, independent of all network synchronization protocols such as Transmission Control Protocol / Internet Protocol, Remote Direct Data Access, and unlimited bandwidth technology, used to achieve hardware-level global clock, task, and status synchronization. The global synchronization hardware pulse generator circuit produces a dedicated native hardware synchronization pulse signal, without any network protocol encapsulation, software parsing, or firmware processing. The point-to-point cascaded hardware transmission port... The interface supports pure hardware point-to-point cascading of multiple nodes, enabling full-domain hardware-level zero-delay transmission of synchronization pulse signals without any network switching, routing forwarding, or software relay. The synchronization hardware calibration module automatically calibrates the end-to-end synchronization signal deviation through fixed hardware logic that is non-reconfigurable and programmable by a non-field programmable gate array, without any software calibration algorithms, software intervention, or firmware adjustment ports, ensuring the synchronization accuracy of all nodes in the entire cluster. The cluster node hardware docking unit is directly connected to the hardware control unit of each computing node, and the synchronization pulse signal directly triggers the node synchronization action without any software intermediate forwarding or firmware intervention, realizing hardware-level collaborative operation of the computing cluster.

[0005] Furthermore, the interface is equipped with a synchronization anomaly hardware monitoring unit. When a node fails to synchronize, the node is directly marked by the hardware circuit and isolated from the synchronization domain, without affecting the synchronization process of the entire cluster, thus greatly improving the stability of cluster synchronization.

[0006] Furthermore, the point-to-point cascaded hardware transmission port supports hot-swappable hardware adaptation. Node plugging and unplugging actions do not require software drivers. The synchronization system automatically adapts to the addition or removal of nodes through fixed hardware logic, without the need for software configuration, thus achieving seamless expansion and maintenance of cluster nodes. Beneficial effects

[0007] Compared with the prior art, the present invention has the following outstanding substantive features and significant beneficial effects: 1. It achieves hardware-level global synchronization without network protocols or software dependencies. The synchronization signal is a native hardware pulse without any network protocol encapsulation or software parsing, which greatly reduces the latency caused by protocol encapsulation and software processing. The synchronization signal transmission latency is as low as less than 1 nanosecond, which is more than 6 orders of magnitude higher than the existing precision time protocol solution. It can meet the nanosecond-level synchronization accuracy requirements for training large artificial intelligence models. 2. The synchronization accuracy is significantly improved and it is not easily affected by external factors. The synchronization calibration is automatically completed through fixed hardware logic that cannot be reconfigured, without any software calibration algorithm or intervention. It is not easily affected by network fluctuations, bandwidth congestion, system load, and process scheduling. The synchronization deviation of all nodes in the entire cluster can be controlled within 1 nanosecond, achieving high-precision and consistent synchronization across the entire domain, which can greatly improve the efficiency and consistency of parallel computing tasks. 3. The architecture is extremely simple and the reliability is significantly improved. The synchronization signal is transmitted across the entire domain through point-to-point pure hardware cascading. There are no network switches, routers or software relays. The architecture is extremely simple and there is no risk of single point of failure. When a single node fails to synchronize, it can be automatically isolated by hardware without affecting the synchronization process across the entire domain. It can achieve long-term uninterrupted and stable operation, which solves the problems of complex architecture and easy crash of existing solutions. 4. Supports seamless hot-swappable expansion; node additions and removals require no software drivers or configuration. The synchronization system can automatically adapt to hardware and support full-domain synchronization of ultra-large-scale clusters with thousands of nodes. It has strong adaptability and can meet the synchronization management needs of ultra-large-scale computing power clusters in intelligent computing centers and supercomputing centers, providing underlying synchronization accuracy guarantees for high-performance parallel computing and artificial intelligence large-scale model training. Detailed Implementation

[0008] The technical solution of the present invention will be further described in detail below.

[0009] The intelligent computing center cluster hardware synchronization interface described in this embodiment is applied to the graphics processor cluster for training large-scale artificial intelligence models in the intelligent computing center. It provides full-domain hardware-level synchronization for 128 racks and 1024 graphics processor computing nodes within the cluster, achieving nanosecond-level precision clock synchronization, training task synchronization, and status synchronization.

[0010] The interface includes a global synchronous hardware pulse generation circuit, a point-to-point cascaded hardware transmission port, a synchronous hardware calibration module, and a cluster node hardware docking unit. The modules are directly connected to each other through onboard hardware lines without any software intermediaries.

[0011] The interface is a dedicated physical interface, independent of all network synchronization protocols such as Transmission Control Protocol / Internet Protocol, Remote Direct Data Access, and Infinite Bandwidth Technology. It does not rely on any existing network infrastructure and is specifically designed for the synchronization of intelligent computing center clusters, completely eliminating the limitations of network protocols on synchronization accuracy.

[0012] The global synchronization hardware pulse generation circuit is a high-precision temperature-controlled crystal oscillator hardware pulse generation circuit that generates a dedicated native hardware synchronization pulse signal with a frequency configurable from 10 MHz to 1 GHz. It has no network protocol encapsulation, software parsing, or firmware processing steps, and the jitter of the synchronization pulse signal is less than 100 picoseconds, providing a high-precision synchronization reference for the entire cluster.

[0013] The point-to-point cascaded hardware transmission port is a multi-channel onboard hardware cascade port that supports pure hardware point-to-point cascaded connection of multiple nodes and multiple racks. The synchronization pulse signal is transmitted with zero latency at the hardware level throughout the entire domain through a dedicated coaxial cable. There are no network switching, routing forwarding, or software relay links, and the transmission link latency is less than 5 nanoseconds per 100 meters. The port supports hot-swappable hardware adaptation. The insertion and removal of nodes / racks do not require software drivers. The synchronization system automatically recognizes the addition and removal of nodes and automatically adapts the synchronization link through fixed hardware logic. No software configuration is required, which realizes seamless expansion and online maintenance of the cluster.

[0014] The synchronization hardware calibration module implements fixed hardware logic that is non-reconfigurable and programmable by a dedicated integrated circuit and a non-field programmable gate array. It automatically calibrates the transmission delay and phase deviation of the full-link synchronization signal without any software calibration algorithm, software intervention, or firmware adjustment port. The calibration accuracy is less than 1 nanosecond, ensuring the synchronization accuracy of all nodes in the entire cluster. Regardless of the distance between nodes, it can achieve consistent nanosecond-level synchronization across the entire domain.

[0015] The cluster node hardware docking unit is directly connected to the hardware control unit of each graphics processor computing node. The synchronization pulse signal directly triggers the node's clock synchronization, task start-up, status acquisition and other synchronization actions without any software intermediate forwarding or firmware intervention. The trigger response time of the synchronization action is less than 1 nanosecond, which greatly reduces the delay and deviation caused by the software intermediate links.

[0016] The interface is equipped with a synchronization anomaly hardware monitoring unit to monitor the synchronization status of each node in real time. When a node fails to synchronize, the node is marked directly by the hardware circuit and isolated from the synchronization domain without affecting the synchronization process of the entire cluster. At the same time, a hardware alarm is triggered, which greatly improves the stability and reliability of cluster synchronization.

[0017] In the application architecture of this embodiment, the intelligent computing center cluster is equipped with one master synchronization interface, which is deployed in the core management cabinet of the cluster to generate a global synchronization reference pulse signal; each cabinet is equipped with one slave synchronization interface, which is cascaded with the master interface and the slave interfaces of adjacent cabinets through point-to-point cascading ports to form a global synchronization ring network; the graphics processor computing power nodes in each cabinet are directly connected to the slave interface of the cabinet through hardware docking units to achieve global synchronization of all 1024 nodes in the cluster.

[0018] The interface described in this embodiment can control the synchronization deviation of all nodes in the entire cluster to within 1 nanosecond, and the synchronization signal transmission delay is as low as 1 nanosecond. It is not easily affected by network fluctuations or system load, has no single point of failure risk, and can operate stably for a long time. It solves the problems of high latency, low accuracy, and poor reliability of existing network synchronization solutions, and provides a high-precision synchronization guarantee at the underlying level for large-scale artificial intelligence model training and high-performance parallel computing, which can improve the parallel efficiency of large-scale model training by more than 30%.

Claims

1. A hardware synchronization interface for an intelligent computing center cluster, characterized in that, It includes a global synchronization hardware pulse generation circuit, a point-to-point cascaded hardware transmission port, a synchronization hardware calibration module, and a cluster node hardware interface unit. The interface is a physical interface specifically designed for multi-rack, multi-computing-power node clusters in intelligent computing centers and supercomputing centers. It is independent of all network synchronization protocols such as transmission control protocols / Internet Protocol, remote direct data access, and unlimited bandwidth technology, and is used to achieve hardware-level global clock, task, and status synchronization. The global synchronization hardware pulse generation circuit generates a dedicated native hardware synchronization pulse signal without any network protocol encapsulation, software parsing, or firmware processing. The point-to-point cascaded hardware transmission port supports multi-node pure... The hardware point-to-point cascaded connection enables zero-delay, full-domain hardware-level transmission of synchronization pulse signals without any network switching, routing, or software relay. The synchronization hardware calibration module automatically calibrates the entire-link synchronization signal deviation through fixed hardware logic that is non-reconfigurable and programmable by a non-field-programmable gate array, without any software calibration algorithms, software intervention, or firmware adjustment ports, ensuring the synchronization accuracy of all nodes in the entire cluster. The cluster node hardware docking unit is directly connected to the hardware control unit of each computing node, and the synchronization pulse signal directly triggers the node synchronization action without any software intermediate forwarding or firmware intervention, realizing hardware-level collaborative operation of the computing cluster.

2. The intelligent computing center cluster hardware synchronization interface according to claim 1, characterized in that, The interface is equipped with a hardware monitoring unit for synchronization anomalies. When a node fails to synchronize, the node is directly marked by the hardware circuit without affecting the synchronization process of the entire cluster.

3. The intelligent computing center cluster hardware synchronization interface according to claim 1, characterized in that, The point-to-point cascaded hardware transmission port supports hot-swappable hardware adaptation. The plugging and unplugging action does not require software drivers, and the synchronization system automatically adapts to the addition or removal of nodes without the need for software configuration.