A multi-card NPU interconnection clock system, a multi-card NPU system and a clock control method thereof

CN122837579APending Publication Date: 2026-09-29奕算智能科技(上海)有限公司 +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611063264.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-17
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

但在高速NPU互连应用中,若仅依赖单一公共时钟直接分发,难以在长距离走线中兼顾信号完整性与本地TX时钟质量,因而无法彻底消除对硬件缓存的依赖

Benefits of technology

[0012]本发明公开的一种多卡NPU互连时钟系统、多卡NPU系统及其时钟控制方法,通过构建二级时钟分布网络实现各子卡时钟同频,以取消冗余缓存,并简化协议层处理逻辑,进而有效减少RDMA模块尺寸和数据传输延迟。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122837579A_ABST
    Figure CN122837579A_ABST
Patent Text Reader

Abstract

The application discloses a multi-card NPU interconnection clock system, a multi-card NPU system and a clock control method thereof, wherein the multi-card NPU interconnection clock system comprises a first clock module, a cache module and a second clock module; the first clock module and the cache module are arranged on a mainboard and are used for outputting a first reference clock as a global frequency reference; the cache module is used for synchronously distributing the first reference clock to eliminate frequency and phase differences between NPU subcards; and the second clock module is arranged on each NPU subcard and is used for locally generating a second reference clock based on the first reference clock as a reference clock of an RDMA channel in the NPU subcard. The multi-card NPU interconnection clock system is applied to the multi-card NPU system, so that the subcards can obtain consistent frequencies, the cache does not need to be added on a data path of the RDMA, and the processing of a protocol layer can be simplified, thereby reducing the size of the RDMA module and the data transmission delay during the multi-card NPU cooperative processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of integrated circuit technology, and in particular to a multi-card NPU interconnect clock system, a multi-card NPU system, and a clock control method thereof. Background Technology

[0002] Currently, multi-GPU neural network processor (NPU) systems commonly use Remote Direct Memory Access (RDMA) interfaces for inter-system connectivity, and their physical layer designs are widely based on Separate Clock Architecture, such as... Figure 1 As shown in the diagram. In this architecture, each NPU daughter card is independently configured with a local clock chip to obtain the best clock quality and clock integrity for the chip output signal.

[0003] Because the clock sources of each daughter card are independent, frequency differences inevitably arise during operation. These frequency differences accumulate over time and cannot be directly eliminated through hardware structure. Therefore, a cache unit must be forcibly inserted into the RDMA data path, directly increasing hardware routing complexity and silicon area footprint. The addition of the cache unit not only occupies valuable PCB trace space but also significantly increases the physical size of the RDMA module. Furthermore, the protocol layer needs to continuously perform clock skew detection and data resynchronization operations, thus increasing data transmission latency during multi-card NPU collaborative processing.

[0004] Chinese patent application CN118939177A discloses a technical solution to address the latency overhead caused by asynchronous clocks between different storage nodes by using a common synchronous clock source. However, in high-speed NPU interconnect applications, relying solely on a single common clock for direct distribution makes it difficult to balance signal integrity and local TX clock quality over long-distance traces, thus failing to completely eliminate the dependence on hardware cache. Furthermore, this solution focuses on latency optimization of the storage system and fails to effectively address the protocol layer processing burden and module size expansion issues caused by cumulative frequency deviations between NPU daughter cards. Therefore, it is difficult to directly migrate to low-latency collaboration scenarios involving multiple NPU cards. Summary of the Invention

[0005] To address some or all of the problems in existing technologies, and to overcome the limitations of independent clock architectures such as cache dependency, module size redundancy, and protocol processing latency, the first aspect of this invention provides a multi-card NPU interconnect clock system, comprising: The first clock module, located on the motherboard, is used to output a primary reference clock as a global frequency base. A cache module, which is mounted on the motherboard and communicatively connected to the first clock module and each of the second clock modules, is used to synchronously distribute the first-level reference clock to eliminate the frequency phase difference between each NPU daughter card. The second clock module is installed on each NPU daughter card and is used to generate a second-level reference clock locally based on the first-level reference clock, which serves as the reference clock for the RDMA channel in the NPU daughter card.

[0006] Furthermore, the cache module includes at least one level of buffer circuitry, which is used to perform signal shaping and pre-fanout driving on the first-level reference clock.

[0007] Based on the multi-card NPU interconnect clock system described above, a second aspect of the present invention provides a multi-card NPU system, comprising: Motherboard; As described above, a multi-card NPU interconnect clock system; The NPU daughter card includes a system-on-chip (SoC), the clock input of which is connected to the second clock module of the multi-card NPU interconnect clock system.

[0008] Furthermore, the system-on-a-chip includes: The Remote Direct Memory Access physical layer (RDMA PHY) has its clock input pin coupled to the output of the second clock module to lock the secondary reference clock generated by the second clock module as the reference clock for the RDMA channel. A third aspect of the present invention provides a clock control method for a multi-card NPU system as described above, comprising: The first-level reference clock is output through the oscillation of the first clock module; The primary reference clock is synchronously transmitted to each NPU daughter card via the caching module; The first-level reference clock is received by the second clock module on each NPU daughter card, and a second-level reference clock is generated based on the first-level reference clock and input to the corresponding NPU SoC chip as the operating reference clock for the RDMA channel.

[0009] Furthermore, the synchronization and transmission of the primary reference clock to each NPU daughter card via the caching module includes: Configure the gain parameters of the cache module to compensate for the signal attenuation of the primary reference clock on the motherboard traces; The processed primary reference clock is distributed in parallel to the clock input interfaces of each NPU daughter card.

[0010] Furthermore, generating a secondary reference clock based on the primary reference clock includes: Determine the frequency parameters of the primary reference clock; Adjust the phase-locked loop parameters of the second clock module to obtain the most stable clock output frequency; Output a phase-corrected secondary reference clock.

[0011] Furthermore, the clock control method further includes: Detect the frequency consistency of the secondary reference clock output by each NPU daughter card; When the frequency deviation of an NPU daughter card exceeds a preset threshold, the primary reference clock and the secondary reference clock are regenerated.

[0012] This invention discloses a multi-card NPU interconnect clock system, a multi-card NPU system and its clock control method. By constructing a two-level clock distribution network, the clocks of each sub-card are synchronized, thereby eliminating redundant buffers and simplifying the protocol layer processing logic, thus effectively reducing the size of the RDMA module and the data transmission latency. Attached Figure Description

[0013] To further illustrate the above and other advantages and features of the various embodiments of the present invention, a more specific description of the various embodiments of the present invention will be presented with reference to the accompanying drawings. It is to be understood that these drawings depict only typical embodiments of the invention and are therefore not intended to limit its scope. In the drawings, identical or corresponding parts will be indicated by identical or similar reference numerals for clarity.

[0014] Figure 1 This diagram illustrates the structure of a discrete clock architecture in the prior art. Figure 2 This diagram illustrates the structure of a multi-card NPU interconnect clock system according to an embodiment of the present invention. Figure 3 The diagram shows a flowchart of a clock control method for a multi-card NPU system according to an embodiment of the present invention. Detailed Implementation

[0015] In the following description, the invention is described with reference to various embodiments. However, those skilled in the art will recognize that the embodiments may be practiced without one or more specific details or in conjunction with other alternatives and / or additional methods or components. In other instances, well-known structures or operations are not shown or described in detail so as not to obscure the inventive points of the invention. Similarly, for illustrative purposes, specific numbers and configurations are set forth to provide a comprehensive understanding of embodiments of the invention. However, the invention is not limited to these specific details. Furthermore, it should be understood that the embodiments shown in the drawings are illustrative representations and are not necessarily drawn to scale.

[0016] In this specification, references to "an embodiment" or "this embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. The phrase "in one embodiment" appearing throughout this specification does not necessarily refer to the same embodiment in all instances.

[0017] It should be noted that the embodiments of the present invention describe the method steps in a specific order; however, this is only for illustrating the specific embodiment and not for limiting the order of the steps. On the contrary, in different embodiments of the present invention, the order of the steps can be adjusted according to actual needs.

[0018] To address the technical shortcomings of existing discrete clock architectures, which are prone to frequency accumulation deviations and require additional buffering and enhanced protocol layer processing in the RDMA path, this invention discloses a multi-card NPU interconnect clock system. This system employs a two-level clock distribution architecture. The motherboard outputs a first-level reference clock to each daughter card, and the daughter card's local clock chip generates a second-level reference clock based on the first-level reference clock, thus achieving multi-card clock synchronization at the hardware level. Based on this, the system can directly eliminate data path buffering and simplify protocol logic, thereby resolving the technical bottlenecks of module size expansion and high transmission latency.

[0019] The technical solution of the present invention will be further described below with reference to the accompanying drawings of the embodiments.

[0020] Figure 2 This diagram illustrates the structure of a multi-card NPU interconnect clock system according to an embodiment of the present invention. Figure 2 As shown, a multi-card NPU interconnect clock system includes a first clock module 201, a cache module 202, and a second clock module 203. The first clock module 201 and cache module 202 are mounted on the motherboard, while the second clock module 203 is mounted on each NPU daughter card. The first clock module 201 outputs a primary reference clock clk1 as a global frequency reference. The cache module 202 synchronously distributes the primary reference clock to eliminate frequency phase differences between the NPU daughter cards. The second clock module 203, based on the primary reference clock, locally generates a secondary reference clock clk2 as the reference clock for the RDMA channel in each NPU daughter card.

[0021] In one embodiment of the present invention, the buffer module 202 is electrically connected to the output of the first clock module 201 and extends to the physical interface of each NPU daughter card. It includes at least one level of buffer circuitry to perform signal shaping and fan-out driving on the first-level reference clock clk1, thereby maintaining the integrity of the first reference clock signal during long-distance motherboard routing. In another embodiment of the present invention, the output of the buffer module 202 is configured as a differential clock signal format to reduce crosstalk interference during motherboard routing.

[0022] In one embodiment of the present invention, the second clock module 203 adopts an independent oscillator architecture, and its input terminal only receives the first-level reference clock clk1 as a frequency and phase-locked source.

[0023] Based on the aforementioned multi-card NPU interconnect clock system, this invention also provides a multi-card NPU system, which includes the aforementioned multi-card NPU interconnect clock system. The NPU daughter card further includes a system-on-chip (SoC), the clock input of which is connected to the output of the second clock module of the multi-card NPU interconnect clock system, for receiving the secondary reference clock clk2 as the channel operating reference. The NPU SoCs on each NPU daughter card interconnect with each other through a high-speed interconnect channel, the clock domain of which is uniformly defined by the secondary reference clock clk2. As mentioned above, in one embodiment of this invention, the high-speed interconnect channel adopts the RDMA interface protocol, and its physical layer clock recovery circuit uses the secondary reference clock clk2 as the synchronization reference.

[0024] In one embodiment of the present invention, the system-on-a-chip integrates a Remote Direct Memory Access Physical Layer (RDMA PHY), a Remote Direct Memory Access Digital Link Layer (RDMA MAC), and a Remote Direct Memory Access Physical Coding Sublayer (RDMA PCS). The clock input pin of the RDMA PHY is coupled to the output of the second clock module. The RDMA MAC is used for data frame encapsulation, parsing, addressing, flow control, and lossless scheduling, while the RDMA PCS is used for encoding / decoding, synchronization, and FEC to convert MAC frames into bitstreams. As mentioned above, since each daughter card uses a consistent frequency reference, the generated secondary reference clock frequencies are theoretically consistent. Based on this, no additional buffer is needed, and the processing of the protocol layer can be simplified, thereby reducing the size of the RDMA module and data transmission latency.

[0025] Figure 3 This diagram illustrates a clock control method for a multi-GPU NPU system according to an embodiment of the present invention. Figure 3 As shown, a clock control method for a multi-card NPU system as described above includes: First, in step 301, the motherboard generates a primary reference clock. The primary reference clock clk1 is output via the first clock module on the motherboard. Next, in step 302, the reference clock is distributed. The primary reference clock is synchronously transmitted to the physical interfaces of each NPU daughter card through the buffer module. In one embodiment of the present invention, when distributing the reference clock, the primary reference clock clk1 is further fan-out driven and impedance matched through the buffer module to maintain signal integrity in long-distance motherboard traces. Specifically, the gain parameters of the buffer module are configured to compensate for signal attenuation of the primary reference clock on the motherboard traces, and then the processed primary reference clock is distributed in parallel to the clock input interfaces of each NPU daughter card. Finally, in step 403, the daughter card generates a secondary reference clock. The primary reference clock is received by the second clock module on each NPU daughter card, and a secondary reference clock clk2 is generated based on the primary reference clock and input to the corresponding NPU SoC chip as the operating reference clock for the RDMA channel. In one embodiment of the invention, generating the secondary reference clock clk2 based on the primary reference clock includes: determining the frequency parameters of the primary reference clock, then adjusting the phase-locked loop parameters of the second clock module to obtain the most stable clock output, and finally outputting the phase-corrected secondary reference clock. In one embodiment of the invention, the operating frequency of the phase-locked loop's frequency detector can be set, for example, to 1 / 8 of the primary reference clock frequency.

[0026] In one embodiment of the present invention, the frequency consistency of the secondary reference clock output by each NPU daughter card is further detected. When the frequency deviation of each NPU daughter card exceeds a preset threshold, the reconfiguration of the secondary generation step is triggered to regenerate the primary reference clock and the secondary reference clock.

[0027] In one embodiment of the present invention, after confirming that the secondary reference clock frequencies of each NPU daughter card are from the same source, the data read / write enable signal of the hardware cache unit of the RDMA data path in the NPU SoC chip can be turned off.

[0028] This invention discloses a multi-card NPU interconnect clock system, a multi-card NPU system and its clock control method. By constructing a two-level clock distribution network, the clocks of each sub-card are synchronized, thereby eliminating redundant buffers and simplifying the protocol layer processing logic, thus effectively reducing the size of the RDMA module and the data transmission latency.

[0029] Although various embodiments of the invention have been described above, it should be understood that they are presented by way of example only and not as limitations. It will be apparent to those skilled in the art that various combinations, modifications, and alterations can be made without departing from the spirit and scope of the invention. Therefore, the breadth and scope of the invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined solely by the appended claims and their equivalents.

Claims

1. A multi-card NPU interconnect clock system, characterized in that, include: The first clock module is located on the motherboard and is configured to output a primary reference clock as a global frequency reference. A cache module, which is located on the motherboard and is communicatively connected to the first clock module and each of the second clock modules, is configured to synchronously distribute the first-level reference clock to eliminate the frequency phase difference between each NPU daughter card; The second clock module is set on each NPU daughter card and is configured to generate a second-level reference clock locally based on the first-level reference clock, which serves as the reference clock for the RDMA channel in the NPU daughter card.

2. The multi-card NPU interconnect clock system as described in claim 1, characterized in that, The cache module includes at least one level of buffer circuitry, which is used to perform signal shaping and pre-fanout driving on the first-level reference clock.

3. A multi-GPU NPU system, characterized in that, include: Motherboard; The multi-card NPU interconnect clock system as described in claim 1 or 2; The NPU daughter card includes a system-on-a-chip (SoC), the clock input of which is connected to the second clock module of the multi-card NPU interconnect clock system.

4. The multi-card NPU system as described in claim 3, characterized in that, The system-on-a-chip includes a remote direct memory access physical layer, and the clock input pin of the remote direct memory access physical layer is coupled to the output of the second clock module to lock the secondary reference clock generated by the second clock module as the reference clock of the RDMA channel.

5. A clock control method for a multi-card NPU system as described in claim 4, characterized in that, include: The first-level reference clock is output through the oscillation of the first clock module; The primary reference clock is synchronously transmitted to each NPU daughter card via the caching module; The first-level reference clock is received by the second clock module on each NPU daughter card, and a second-level reference clock is generated based on the first-level reference clock and input to the corresponding NPU SoC chip as the operating reference clock for the RDMA channel.

6. The clock control method as described in claim 5, characterized in that, The primary reference clock is synchronously transmitted to each NPU daughter card via the caching module, including: Configure the gain parameters of the cache module to compensate for the signal attenuation of the primary reference clock on the motherboard traces; The processed primary reference clock is distributed in parallel to the clock input interfaces of each NPU daughter card.

7. The clock control method as described in claim 5, characterized in that, Generating a secondary reference clock based on the primary reference clock includes: Determine the frequency parameters of the primary reference clock; Adjust the phase-locked loop parameters of the second clock module to obtain the most stable clock output frequency; Output a phase-corrected secondary reference clock.

8. The clock control method as described in claim 5, characterized in that, Also includes: Detect the frequency consistency of the secondary reference clock output by each NPU daughter card; When the frequency deviation of an NPU daughter card exceeds a preset threshold, the primary reference clock and the secondary reference clock are regenerated.

Citation Information

Patent Citations

  • System and method for storage

    CN118939177A