Parallel interface data transmission method based on 3D package chip communication interface system
Patent Information
- Application Number
- CN202611059494.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-07-16
AI Technical Summary
但这些校准手段往往需要模拟电路辅助,或者在数字实现中占用大量逻辑资源与校准时间,难以兼顾精度、速度和硬件开销
[0048]本发明提供的并行接口数据传输方法,将并行接口信号通过分级校准,第一级延迟调整以第一级延迟单元为调整粒度,对并行接口信号每个数据字节通过自动比较操作确定每一数据字节的第一延迟窗口,并根据所述第一延迟窗口设置所述第一级延迟单元使能数量;第二级延迟调整以第二级延迟单元为调整粒度,对每个数据字节的每个比特通过自动比较操作确定每一比特的第二延迟窗口,并根据所述第二延迟窗口设置所述第二级延迟单元的使能数量。该方法中,通过第一级延迟调整,调整延迟链路中第一级延迟单元的导通/旁路状态,调整信号之间较大的偏差,保证clock可以采样到数据。通过第二级延迟调整,调整导通状态下的第二级延迟单元的导通/旁路状态,实现调整每个byte中每个bit的细节偏差。通过两级校准,达到调整信号延迟目的,让并行接口信号对齐到最佳采样窗口,提供最佳的setup/hold裕量。通过多级校准机制,实现了校准精度和开销的平衡,解决了并行传输的时序问题,保障了高速并行通信的可靠性。
Smart Images

Figure CN122570410B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of AI chip technology, and in particular to a parallel interface data transmission method based on a 3D packaged chip communication interface system. Background Technology
[0002] Existing high-performance cloud computing chips (such as AI inference chips and multimedia processing chips) all have a module responsible for communicating with the external host-side main control CPU chip. This module is responsible for receiving tasks and model networks issued by the main control CPU chip. This module typically uses a high-speed serial interface such as PCIe (Rapid Interconnect Standard) or CXL (Rapid Interconnect Standard). Figure 1 As shown, the host 10 is equipped with a high-speed serializer / deserializer analog circuit (SerDes) 11, and the high-performance computing chip (xPU) 20 integrates a PCIE / CXL controller IP core. This IP core includes a high-speed serializer / deserializer analog circuit 21, which is connected to the host through PCB traces and receives inference tasks and neural network model data issued by the host.
[0003] To achieve higher integration density, greater bandwidth, lower latency, and optimized power consumption, high-performance computing chips employ 3D packaging. 3D packaging utilizes through-silicon via (TSV) technology to vertically stack multiple dies directly, such as... Figure 2 As shown, multi-layer memory dies 51 are vertically stacked on computing dies 52, with PCIe-EP located on the computing die. However, in this scheme, the 3D-packaged high-performance computing chip needs to reserve space for TSVs in the chip's physical layout. As a sensitive analog circuit, placing TSVs around high-speed SerDes can introduce mechanical stress, signal interference, and impedance discontinuities, affecting the performance of the analog circuit. Moreover, if PCIe IP is used in the chip, it often requires physical adjustments. The PCIe / CXL interface protocol is complex, and the design threshold for high-speed SerDes analog circuits is extremely high. In-house development requires significant manpower and time investment, and the risk of tape-out is high. Furthermore, purchased PCIe IP is usually a black box, difficult to modify and complex to verify, significantly increasing the risk of project delays or tape-out failures.
[0004] Furthermore, the proportion of inference demands for small-to-medium-sized models with fewer than 10 billion parameters in current cloud-based AI inference scenarios continues to increase. These inference tasks can be independently handled by a single 3D-packaged chip, without the need for multi-chip collaborative computation. However, the aforementioned 3D-packaged chips are generally designed to meet all scenario requirements and typically integrate high-speed SerDes interfaces. For lightweight inference scenarios with small-to-medium-sized models on a single chip, this is a typical over-design. Not only is the high bandwidth capability completely unnecessary, but it also requires additional 3D packaging adaptation risks associated with SerDes circuitry.
[0005] In existing technologies, to reduce reliance on high-speed SerDes analog circuits, some solutions attempt to replace serial interfaces with parallel interfaces. However, parallel interfaces face severe timing challenges during high-speed transmission: due to differences in PCB trace lengths, fluctuations in chip manufacturing processes, and changes in operating temperature, the phase relationship between clock and data signals dynamically drifts, resulting in insufficient setup and hold time margins. Simultaneously, inconsistent transmission delays between multi-bit data lines cause data skew, further compressing the effective sampling window. To address these issues, traditional parallel interfaces typically require the integration of complex timing calibration logic, such as adjusting the clock phase using delay-locked loops (DLLs) or phase interpolators, or employing automatic data sampling point alignment algorithms at the receiving end. However, these calibration methods often require analog circuitry or consume significant logic resources and calibration time in digital implementations, making it difficult to balance accuracy, speed, and hardware overhead. More importantly, when parallel interfaces are used in 3D packaged chips, the mismatch between the distribution of through-silicon vias (TSVs) inside the chip and the interlayer interconnects introduces additional signal offsets, making it difficult for traditional timing calibration methods to converge and limiting the practical bandwidth of parallel interfaces in 3D packaging scenarios. Summary of the Invention
[0006] To address the shortcomings of the existing technologies, this invention provides a parallel interface data transmission method based on a 3D packaged chip communication interface system. This method uses graded calibration to achieve deskipation and level adjustment of parallel interface signals between a 3D packaged chip (excluding high-speed serializer / deserializer analog circuits) and an FPGA bridge chip. This ensures the signal quality of the parallel signals, maximizes the setup / hold margin, solves the timing problem of parallel transmission, and guarantees the reliability of high-speed parallel communication.
[0007] This invention also provides a parallel interface data transmission method based on a 3D packaged chip communication interface system, applied to a 3D packaged chip communication interface system. The 3D packaged chip communication interface system includes a 3D packaged chip and an FPGA bridge chip. The 3D packaged chip is only equipped with a digital parallel interface to exclude high-speed serializer / deserializer analog circuitry. The FPGA bridge chip is connected to the 3D packaged chip via the digital parallel interface. The FPGA bridge chip is the transmitting end, and the 3D packaged chip is the receiving end.
[0008] The sending end performs the following steps:
[0009] The first-level delay adjustment is performed with the first-level delay unit as the adjustment granularity. For each data byte of the parallel interface signal, the first delay window of each data byte is determined by automatic comparison operation, and the number of first-level delay units enabled is set according to the first delay window.
[0010] The second-level delay adjustment is performed with the second-level delay unit as the adjustment granularity. For each bit of each data byte, the second delay window of each bit is determined by an automatic comparison operation, and the number of enabled second-level delay units is set according to the second delay window.
[0011] The parallel interface signal is configured with a delay link, which includes multiple sets of cascaded first-level delay units. Each set of first-level delay units includes multiple sets of cascaded second-level delay units. Each delay unit supports two states: on and off. In the on state, the signal passes through the delay unit and generates a fixed delay. In the off state, the signal bypasses the delay unit directly without delay.
[0012] In one embodiment of the present invention, when performing the first-level delay adjustment, starting from the first minimum delay state, the first-level delay units are enabled sequentially, one first-level delay unit is enabled at a time, and an automatic comparison operation is performed for each enabled number, until the first maximum delay state is reached; wherein the first minimum delay state is when all first-level delay units are in a bypass state, and the first maximum delay state is when all first-level delay units are in an enabled state; and / or
[0013] When performing the second-level delay adjustment, for the second-level delay units that are in the enabled state of the first-level delay units, starting from the second minimum delay state, all second-level delay units are put in the bypass state, and the second-level delay units are enabled sequentially, one second-level delay unit at a time. An automatic comparison operation is performed for each number of enabled units until the second maximum delay state is reached. The second minimum delay state is when all second-level delay units are in the bypass state, and the second maximum delay state is when all second-level delay units are in the enabled state.
[0014] In one embodiment of the present invention, determining a first delay window or a second delay window through an automatic comparison operation includes:
[0015] When the correct data is obtained for the first time through automatic comparison, the delay position corresponding to the current number of enabled functions is recorded as the left boundary of the delay window;
[0016] Continue increasing the number of enabled delay units until the automatic comparison operation first encounters a data acquisition error. Then, use the delay position corresponding to the previous number of enabled units as the right boundary of the delay window.
[0017] In one embodiment of the present invention, the automatic comparison operation includes:
[0018] The "write-then-read" mode includes: writing a test sequence to a specified location in the memory of the receiving end, reading back the data from the specified location on the sending end, and comparing the read-back data with the test sequence. If the two match, it is determined that the correct data has been obtained; otherwise, it is determined that the data acquisition is incorrect.
[0019] The "read-only" mode includes: the transmitting end reads data from a preset address in the memory of the receiving end and compares the read-back data with the expected data;
[0020] The "write-only" mode includes writing a test sequence to a specified location in the memory on the receiving end without the sending end reading back the data.
[0021] In one embodiment of the present invention, the parallel interface signal includes an input signal and an output signal, the input signal includes parallel input data and ECC-encoded input data corresponding to the parallel input data, and the output signal includes parallel output data and ECC-encoded output data corresponding to the parallel output data.
[0022] If a single data byte prevents the determination of the first delay window, exception handling is performed in the following order of priority:
[0023] Fix the delay of the output signal corresponding to the data byte, adjust the delay of the input signal corresponding to the data byte separately, and re-verify the delay window of the input signal of the data byte through the "read-only" mode;
[0024] Fix the delay of the input signal corresponding to the data byte, adjust the delay of the output signal corresponding to the data byte separately, and re-verify the delay window of the output signal of the data byte through the "write-only" mode;
[0025] Adjust the delay of the parallel interface clock signal and re-traverse all data bytes to determine the first delay window;
[0026] Reduce the operating clock frequency of the parallel interface and re-traverse all data bytes to determine the first delay window;
[0027] And / or,
[0028] If a single bit prevents the determination of the second delay window, exception handling is performed in the following order of priority:
[0029] Fix the delay of the output signal corresponding to the bit, adjust the delay of the input signal corresponding to the bit separately, and re-verify the delay window of the input signal of the bit through the "read-only" mode;
[0030] Fix the delay of the input signal corresponding to the bit, adjust the delay of the output signal corresponding to the bit separately, and re-verify the delay window of the output signal of the bit through the "write-only" mode;
[0031] Adjust the delay of the parallel interface clock signal and re-traverse all bits to determine the second delay window;
[0032] Reduce the operating clock frequency of the parallel interface and re-traverse all bits to determine the second delay window.
[0033] In one embodiment of the present invention, setting the number of first-level delay units enabled according to the first delay window includes: calculating the first intersection window of the first delay windows of all data bytes, and configuring the number of first-level delay units enabled for each data byte such that the delay of the data byte is equal to the delay value corresponding to the center point of the first intersection window.
[0034] And / or,
[0035] The step of setting the number of second-level delay units enabled according to the second delay window includes: calculating the second intersection window of the second delay windows of all bits, and configuring the number of second-level delay units enabled for each bit so that the delay of the bit is equal to the delay value corresponding to the center point of the second intersection window.
[0036] In one embodiment of the present invention, if the first intersection window of the first delay windows of all data bytes is less than a first window threshold and greater than or equal to a third window threshold, then the delay of the clock signal is adjusted, and the first-level delay adjustment is re-executed, wherein the third window threshold is less than the first window threshold; and / or,
[0037] If the second intersection window of the second delay windows of all bits is less than the second window threshold, then the delay of the clock signal is adjusted, and the second-level delay adjustment is re-executed.
[0038] In one embodiment of the present invention, if the first intersection window of the first delay window of all data bytes is less than the third window threshold, then the second-level delay adjustment is performed with the second-level delay unit as the adjustment granularity.
[0039] In one embodiment of the present invention, it further includes:
[0040] Using the calibrated delay link, the transmitting end scrambles and ECC-encodes the parallel interface signal before sending it, while the receiving end performs ECC verification and descrambles the received parallel interface signal before processing.
[0041] In one embodiment of the present invention, the steps of using a calibrated delay link, scrambling and ECC encoding the parallel interface signal at the transmitting end before transmission, and performing ECC verification and descrambling on the received parallel interface signal at the receiving end, include:
[0042] On the transmitting end, the original parallel interface signal to be transmitted is grouped into units of 32 bits each, and each group of data is scrambled.
[0043] Generate an ECC checksum for each group of scrambled data, and send the scrambled parallel data and the corresponding ECC checksum synchronously.
[0044] At the receiving end, the received parallel data and ECC check code are grouped and matched in units of 32 bits of data + 6 bits of check code.
[0045] The error check and correction are performed on each group of data according to the ECC check code. If there is only a 1-bit data error, the error correction is completed automatically. If there are 2 or more errors, an interrupt signal is triggered to notify the sender to retransmit.
[0046] For each set of data that has passed verification or completed error correction, descrambling is performed to restore the original parallel interface signal.
[0047] As can be seen from the above solutions, the advantages of the present invention are:
[0048] The parallel interface data transmission method provided by this invention performs hierarchical calibration on the parallel interface signal. The first-level delay adjustment uses a first-level delay unit as the adjustment granularity. For each data byte of the parallel interface signal, an automatic comparison operation determines a first delay window for each data byte, and the number of enabled first-level delay units is set according to the first delay window. The second-level delay adjustment uses a second-level delay unit as the adjustment granularity. For each bit of each data byte, an automatic comparison operation determines a second delay window for each bit, and the number of enabled second-level delay units is set according to the second delay window. In this method, the first-level delay adjustment adjusts the conduction / bypass state of the first-level delay units in the delay link, adjusting large deviations between signals to ensure that the clock can sample data. The second-level delay adjustment adjusts the conduction / bypass state of the second-level delay units in the conduction state, achieving adjustment of the detailed deviation of each bit in each byte. Through two-level calibration, the signal delay is adjusted, aligning the parallel interface signal to the optimal sampling window and providing optimal setup / hold margin. This multi-level calibration mechanism achieves a balance between calibration accuracy and overhead, solves the timing problems of parallel transmission, and ensures the reliability of high-speed parallel communication. Attached Figure Description
[0049] Figure 1This is a schematic diagram of the structure of a communication interface system composed of a main control CPU chip and a high-performance computing chip in the prior art.
[0050] Figure 2 A schematic diagram of a high-performance computing chip in 3D packaging;
[0051] Figure 3 This is a schematic diagram of the structure of a 3D packaged chip communication interface system provided in an embodiment of the present invention;
[0052] Figure 4 A schematic diagram of the specific structure of a 3D packaged chip communication interface system provided in yet another embodiment;
[0053] Figure 5 A schematic diagram of the main-side link control unit;
[0054] Figure 6 This is a schematic diagram of the specific structure of the delay link;
[0055] Figure 7 This is a schematic diagram of the side-link control unit.
[0056] Figure 8 This is a flowchart illustrating a parallel interface data transmission method according to an embodiment of the present invention.
[0057] Figure 9 A flowchart illustrating a parallel interface data transmission method according to another embodiment of the present invention;
[0058] Figure 10 This is a schematic diagram of L1 delay unit adjustment in write-before-read mode.
[0059] The attached figures are labeled as follows:
[0060] 10 - Host computer, 20 - High-performance computing chip, 11, 21 - High-speed serializer / deserializer analog circuit;
[0061] 51-Memory die, 52-Compute die;
[0062] 31-Host, 311-PCIE root endpoint module, 312-First serializer / deserializer analog circuit, 313-High-speed serial interface;
[0063] 32-3D packaged chip, 321-digital parallel interface, 322-slave link module, 3221-slave-side link control unit, 3222-parallel to AXI unit, 32211-second interface conversion module, 32212-second ECC module, 32213-SRAM;
[0064] 33-FPGA bridge chip, 331-first side, 332-second side, 3321-standard high-speed serial interface, 3311-parallel general-purpose I / O interface, 333-main link module, 3331-main-side link control unit, 33311-delay link, L1-first-level delay unit, L2-second-level delay unit, 33312-first interface conversion module, 33313-first ECC module, 3332-AXI to parallel unit, 334-PCIE endpoint module, 335-second serializer / deserializer analog circuit. Detailed Implementation
[0065] It should be noted that, in this invention, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.
[0066] In the absence of further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0067] Example 1:
[0068] See Figure 3 , Figure 3 A schematic diagram of the structure of a 3D packaged chip communication interface system provided in an embodiment of the present invention is shown.
[0069] A 3D packaged chip communication interface system includes: a host 31, a 3D packaged chip 32, and an FPGA bridge chip 33.
[0070] The 3D packaged chip 32 is equipped with only a digital parallel interface 321 to exclude the high-speed serializer / deserializer analog circuit, i.e., it does not have a high-speed serializer / deserializer SerDes analog circuit. This digital parallel interface 321 is responsible for outputting / inputting task data issued by the host CPU chip of the host 31 in a parallel manner.
[0071] The first side 331 of the FPGA bridge chip 33 is the parallel interface side, which is connected to the digital parallel interface 321 of the 3D packaged chip through the parallel general-purpose I / O interface 3311 to receive / send parallel data.
[0072] The second side 332 of the FPGA bridge chip is a high-speed serial interface side, integrating a standard high-speed serial interface 3321, which is connected to the high-speed serial interface 313 of the host 31. In one embodiment, the standard high-speed serial interface adopts a PCIe or CXL serial interface.
[0073] In this embodiment, an FPGA bridge chip 33 is added between the 3D package chip 32 and the host 31. The physical layer implementation of the high-speed interface is transferred from the 3D package chip to the external FPGA bridge chip. The parallel port of the FPGA is used to replace the PCIe serial port. The high-speed serial interfaces such as PCIe / CXL are completely removed from the 3D package chip, while the digital parallel interface is retained. Inter-chip communication is carried out entirely using digital parallel interface signals. The parallel digital signals are not sensitive to the physical changes introduced by TSV. There is no need to make complex adjustments to the module to adapt to TSV. This fundamentally avoids the design risks of high-speed analog circuits and the TSV adaptation problem, and reduces IP costs.
[0074] Furthermore, in this embodiment, the FPGA bridge chip has a built-in PCIe interface, which reduces a lot of development work. The FPGA uses digital logic to implement signal deskewing and level adjustment, which solves the timing problem of parallel transmission and ensures the reliability of high-speed parallel communication.
[0075] In a preferred implementation, the multi-layer memory dies of the 3D packaged chip 32 are vertically stacked on top of the computing dies, and the FPGA and xPU are not packaged together.
[0076] In a preferred embodiment, the FPGA bridging chip 33 and the 3D packaging chip 32 are encapsulated together. The FPGA bridging chip 33 and the 3D packaging chip 32 are arranged adjacently on the same packaging substrate, and the physical distance between them is less than a preset distance threshold. They are set close to each other to avoid excessively long wiring.
[0077] In one implementation, the 3D packaged chip 32 can be an AI inference chip or a multimedia processing chip, such as a GPU, TPU, NPU, LPU, etc., and the present invention does not make any specific limitation.
[0078] In a preferred implementation, the FPGA bridging chip configures timing calibration logic to deskew and adjust the levels of the parallel interface signals, resolving timing issues in parallel transmission and ensuring the reliability of high-speed parallel communication. See details... Figure 4 As shown, Figure 4This is a schematic diagram of the specific structure of a 3D packaged chip communication interface system provided in one embodiment. The first side 331 of the FPGA bridge chip 33 integrates a main link module 333, which includes a main-side link control unit 3331. The 3D packaged chip 32 integrates a slave link module 322, which includes a slave-side link control unit 3221. The slave link module 322 is connected to the main link module 333 via a digital parallel interface. The main-side link control unit 3331 is configured with timing calibration logic, ECC encoding / decoding logic, and scrambling / descrambling logic to perform timing calibration and data fault tolerance processing for the parallel interface signals. Timing calibration is controlled and implemented by the main-side link control unit. The slave-side link control unit 3221 is only configured with minimal ECC verification logic and descrambling / descrambling logic, and does not require timing calibration.
[0079] Further reference Figure 4 As shown, the host 31 includes a PCIe root endpoint module (PCIe-RC) 311, which includes a first serializer / deserializer analog circuit 312. The second side of the FPGA bridge chip 33 also includes a PCIe endpoint module (PCIe-EP) 334, which includes a second serializer / deserializer analog circuit 335. The first serializer / deserializer analog circuit 312 and the second serializer / deserializer analog circuit 335 are connected via a high-speed serial link, enabling serial communication between the host 31 and the FPGA bridge chip 33.
[0080] Further reference Figure 4 As shown, the main link module 333 also includes an AXI-to-parallel conversion unit 3332, used to convert the AXI interface of the PCIe endpoint module 334 into an APB interface. The slave link module 322 also includes a parallel-to-AXI conversion unit 3222, which converts the APB interface into an AXI interface and accesses other logic through the on-chip bus of the 3D packaged chip 32.
[0081] refer to Figure 5 As shown, Figure 5 This is a schematic diagram of the main-side link control unit 3331. In one embodiment, the main-side link control unit 3331 includes a delay link 33311, a first interface conversion module 33312, and a first ECC module 33313. The first interface conversion module 33312 is a parallel interface that converts an APB interface to a PADN interface. The first ECC module 33313 is used to perform ECC encoding on the output data of the first interface conversion module 33312 by word, ECC verification decoding on the input data by word, and functions such as scrambling / descrambling logic, controlling the delay link, and timing calibration. The timing calibration logic is implemented based on the delay link.
[0082] refer to Figure 6 As shown, Figure 6 This is a schematic diagram of the specific structure of the delay link. The delay link 33311 includes multiple sets of first-level delay units L1, and each set of first-level delay units L1 includes multiple second-level delay units L2. In one specific implementation, the delay link 33311 includes multiple sets of cascaded first-level delay units L1, and each set of first-level delay units L1 includes multiple cascaded second-level delay units L2. Some bits in the parallel interface signal can each correspond to one delay link. Furthermore, in another implementation, the delay link 33311 includes multiple sets of parallel first-level delay units L1, and each set of first-level delay units L1 includes multiple cascaded second-level delay units L2. Moreover, the specific structure of the delay link 33311 is not limited to this and can be used with a selector to achieve conduction control.
[0083] Timing calibration can be achieved within the FPGA bridge chip through programmable delay links. The timing calibration logic can employ hierarchical calibration to ensure accurate timing. For example, coarse calibration is first performed with the first-level delay unit as the adjustment granularity, automatically comparing each byte of the parallel interface signal to determine the first delay window for each data byte. Then, fine calibration is performed with the second-level delay unit as the adjustment granularity, automatically comparing each bit of each data byte to determine the second delay window for each bit. The detailed logic for timing calibration will be described in detail in the subsequent embodiment two.
[0084] In this embodiment, timing calibration can be achieved through a programmable delay link inside the FPGA bridge chip, which solves the timing problem of parallel transmission and ensures the reliability of high-speed parallel communication.
[0085] refer to Figure 7 As shown, Figure 7 This is a schematic diagram of the slave link control unit 3221. The slave link control unit 3221 includes a second interface conversion module 32211, a second ECC module 32212, and an SRAM on-chip cache module 32213. The second interface conversion module 32211 is used to convert the pad parallel interface to an app interface. The second ECC module 32212 is used to perform ECC verification and descrambling on the input data of the parallel interface, and to perform ECC encoding and scrambling on the output data by word. The slave link control unit 3221 integrates an SRAM on-chip cache module 32213, which is used as a buffer during normal operation and mapped to the register address space, so that the space of this buffer can be accessed during initialization.
[0086] In this embodiment, the 3D packaged chip 32 only needs to process parallel interface signals, which are digital signals. Its communication module is entirely composed of digital logic. The slave link module 322 includes a slave-side link control unit 3221 and a parallel-to-AXI unit 3222. The slave-side link control unit 3221 is equipped with a second interface conversion module 32211, a second ECC module 32212, and an SRAM on-chip cache module 32213. No delay link needs to be set up, and no timing calibration logic is configured. Only the simplest ECC verification logic and descrambling / scrambling logic are configured. This makes the entire design flow of the AI inference chip simpler and completely unrestricted by the placement of the TSV.
[0087] In this embodiment, the 3D packaged chip is a single-chip structure with a single computing die stacked with multiple high-bandwidth memory dies. It is configured with only the digital parallel interface as the only chip-level interface for communication with the outside world, which is particularly suitable for small-to-medium-scale lightweight scenarios with independent single-chip inference.
[0088] In addition, in one embodiment, the FPGA bridge chip 33 may adopt an independent FPGA bridge chip structure that is additionally set between the 3D packaged chip 32 and the main control CPU chip of the host 31; or it may adopt existing FPGA resources in the system, such as a heterogeneous main control CPU chip with integrated FPGA hard core.
[0089] Example 2:
[0090] refer to Figure 8 As shown, Figure 8 A flowchart illustrating a parallel interface data transmission method provided in one embodiment is shown.
[0091] A parallel interface data transmission method is applied to the 3D packaged chip communication interface system of Embodiment 1 above. The 3D packaged chip communication interface system includes a 3D packaged chip and an FPGA bridge chip. The 3D packaged chip only has a digital parallel interface to exclude high-speed serializer / deserializer analog circuitry. The FPGA bridge chip and the 3D packaged chip are connected through the digital parallel interface. The FPGA bridge chip is the transmitting end, and the 3D packaged chip is the receiving end. In this embodiment, for parallel communication between the FPGA bridge chip and the 3D packaged chip, the FPGA bridge chip is the transmitting end, and the 3D packaged chip is the receiving end. All timing calibration logic can be deployed on the FPGA side. When adapting to scenarios without SerDes 3D packaged chips, the 3D packaged chip does not need to integrate additional calibration circuitry or SerDes analog circuitry.
[0092] The method includes the following steps:
[0093] Step S1: Perform first-level delay adjustment with the first-level delay unit as the adjustment granularity. For each data byte of the parallel interface signal, determine the first delay window of each data byte through automatic comparison operation, and set the number of first-level delay units enabled according to the first delay window.
[0094] Step S2: Perform second-level delay adjustment with the second-level delay unit as the adjustment granularity. For each bit of each data byte, determine the second delay window for each bit through an automatic comparison operation, and set the number of enabled second-level delay units according to the second delay window.
[0095] Each bit of the parallel interface signal is configured with a delay link. The delay link includes multiple sets of cascaded first-level delay units. Each set of first-level delay units includes multiple sets of cascaded second-level delay units. Each delay unit supports two states: on and off. In the on state, the signal generates a fixed delay after passing through the delay unit. In the off state, the signal bypasses the delay unit directly without delay.
[0096] In this embodiment, the parallel interface signal is calibrated in stages. First, a first-stage delay adjustment adjusts the on / bypass state of the first-stage delay unit in the delay link, where the on state is enabled and the bypass state is bypassed. This adjusts large deviations between signals to ensure that the range of successfully sampled data is found within the sampling period of the clock signal. Then, a second-stage delay adjustment adjusts the on / bypass state of the second-stage delay unit of the first-stage delay unit in the on state, thereby adjusting the detailed deviation of each bit in each byte. Through two-stage calibration, the signal delay is adjusted, aligning the parallel interface signal to the optimal sampling window and providing optimal setup / hold margin. This multi-stage calibration mechanism achieves a balance between calibration accuracy and overhead, solves the timing problems of parallel transmission, and ensures the reliability of high-speed parallel communication.
[0097] The parallel interface signals include input signals and output signals. The input signals include parallel input data and ECC-encoded input data corresponding to the parallel input data. The output signals include parallel output data and ECC-encoded output data corresponding to the parallel output data. The parallel input data is the data input to the transmitting FPGA, including read data. The parallel output data is the data output downstream by the FPGA, including address and write data.
[0098] In a preferred implementation, in step S2, if the first intersection window of the first delay windows of all data bytes after the first-level delay adjustment is less than the third window threshold, then a second-level delay adjustment is further triggered, and the second-level delay adjustment is performed with the second-level delay unit as the adjustment granularity for fine-tuning. If the first intersection window of the first delay windows of all data bytes after the first-level delay adjustment is greater than or equal to the third window threshold, then the second-level delay adjustment does not need to be triggered, and timing calibration is performed only through the first-level delay adjustment. The first-level delay adjustment is a coarse calibration, and the second-level delay adjustment is a fine calibration. The calibration level is flexibly selected according to the size of the first intersection window after coarse calibration: when the first intersection window is sufficient, fine calibration is skipped, and calibration time is compressed; when the first intersection window is insufficient, fine calibration is automatically triggered to ensure timing accuracy, balancing calibration efficiency and transmission reliability.
[0099] In a preferred implementation, in step S1, when performing the first-level delay adjustment, starting from the first minimum delay state, the first-level delay units are enabled sequentially, one first-level delay unit is enabled at a time, and an automatic comparison operation is performed for each enabled number until the first maximum delay state is reached; wherein the first minimum delay state is when all first-level delay units are in a bypass state, and the first maximum delay state is when all first-level delay units are in an enabled state.
[0100] In step S2, when performing the second-level delay adjustment, for the second-level delay units of the first-level delay units that are in the enabled state, starting from the second minimum delay state, all second-level delay units are made to be in the bypass state, and the second-level delay units are enabled sequentially, one second-level delay unit at a time. An automatic comparison operation is performed for each number of enabled units until the second maximum delay state is reached. The second minimum delay state is when all second-level delay units are in the bypass state, and the second maximum delay state is when all second-level delay units are in the enabled state.
[0101] In a preferred implementation, in step S1, determining the first delay window through an automatic comparison operation includes:
[0102] When the correct data is obtained for the first time through automatic comparison, the delay position corresponding to the current number of enabled functions is recorded as the left boundary of the first delay window;
[0103] Continue to increase the number of enabled first delay units until the automatic comparison operation first encounters a data acquisition error. Then, use the delay position corresponding to the previous number of enabled units as the right boundary of the first delay window.
[0104] In step S2, the second delay window is determined through an automatic comparison operation, which is similar to the first delay window operation. Specifically, when the correct data is obtained for the first time through the automatic comparison operation, the delay position corresponding to the current number of enabled units is recorded as the left boundary of the second delay window. The number of enabled units of the second delay unit is increased until the automatic comparison operation first fails to obtain data. Then, the delay position corresponding to the previous number of enabled units is used as the right boundary of the second delay window.
[0105] In a preferred implementation, the automatic comparison operation adopts a "write-before-read" mode. Specifically, a test sequence is written to a designated location in the memory of the receiving end. The sending end reads back data from the designated location and compares it with the test sequence. If they match, it is determined that the correct data has been obtained, i.e., the correct test sequence can be read; otherwise, it is determined that the data obtained is incorrect. It should be noted that the test sequences set in the two comparison operations—one for determining the first delay window of each data byte of the parallel interface signal through automatic comparison, and the other for determining the second delay window of each bit of each data byte through automatic comparison—can be different. In this case, the delays of the input and output signals corresponding to the data byte are adjusted synchronously, and after each adjustment, a "write-before-read" operation is performed on a designated location in the SRAM of the slave-side link control unit. The master-side link control unit writes a fixed test sequence to a specified location in the SRAM of the slave-side link control unit via the output signal corresponding to the data byte. After writing, the master-side link control unit reads back the data from the same SRAM address via the input signal corresponding to the data byte. It automatically compares the read-back data with the test sequence stored in the SRAM of the slave-side link control unit. If they match, the data acquisition is considered correct, meaning the correct test sequence can be read, indicating that the delays of both the input and output signals are within a reasonable range. Otherwise, the data acquisition is considered incorrect, meaning the correct test sequence was not read. In this "write-then-read" mode, the write path and the read-back path share the same delay parameter.
[0106] For each bit of the signal, control is performed from the maximum delay to the minimum delay. The arrival time of each path is finely adjusted to determine which range of data can be stable and correct. The middle position of the overlapping part of each effective window is found. This middle position is the optimal setup / hold margin position.
[0107] In a preferred implementation, the automatic comparison operation employs a "read-only" mode, comprising: the transmitting end reading data from a preset address in the receiving end's memory and comparing the readback data with the expected data. In this "read-only" mode, the write path delay parameter is fixed, and only the readback path delay parameter is adjusted.
[0108] In a preferred implementation, the automatic comparison operation employs a "write-only" mode, comprising: writing a test sequence to a specified location in the receiver's memory while the sender does not read back data. In this case, the correctness of the write operation can be determined based on the write confirmation information fed back from the receiver, but this invention does not impose specific limitations. In the "read-only" mode, the readback path delay parameter is fixed, and only the write path delay parameter is adjusted.
[0109] In a preferred implementation, if a single data byte cannot determine the first delay window, exception handling is performed in the following order of priority: fix the delay of the output signal corresponding to the data byte, adjust the delay of the input signal corresponding to the data byte separately, and then re-verify whether the data byte can determine the delay window of the input signal through "read-only" mode; fix the delay of the input signal corresponding to the data byte, adjust the delay window of the output signal corresponding to the data byte separately, and then re-verify whether the data byte can determine the delay window of the output signal through "write-only" mode.
[0110] Adjust the delay of the parallel interface clock signal and re-traverse all data bytes to determine the first delay window; reduce the operating clock frequency of the parallel interface and re-traverse all data bytes to determine the first delay window.
[0111] If a single bit cannot determine the second delay window, exception handling is performed in the following order of priority: Fix the delay of the output signal corresponding to the bit, adjust the delay of the input signal corresponding to the bit separately, and re-verify whether the bit can determine the second delay window; Fix the delay of the input signal corresponding to the data byte, adjust the delay of the output signal corresponding to the bit separately, and re-verify whether the bit can determine the second delay window; Adjust the delay of the parallel interface clock signal and re-traverse all bits to determine the second delay window; Reduce the operating clock frequency of the parallel interface and re-traverse all bits to determine the second delay window.
[0112] Furthermore, in one embodiment, in step S1, setting the number of first-level delay units enabled according to the first delay window includes: calculating the first intersection window of the first delay windows of all data bytes, and configuring the number of first-level delay units enabled for each data byte such that the delay of the data byte is equal to the delay value corresponding to the center point of the first intersection window.
[0113] In step S2, setting the number of second-level delay units enabled according to the second delay window includes: calculating the second intersection window of the second delay windows of all bits, and configuring the number of second-level delay units enabled for each bit so that the delay of the bit is equal to the delay value corresponding to the center point of the second intersection window.
[0114] If the first intersection window of the first delay windows of all data bytes is less than the first window threshold and greater than or equal to the third window threshold, then the clock signal delay is adjusted, and the first-level delay adjustment is re-executed, where the third window threshold is less than the first window threshold. If the second intersection window of the second delay windows of all bits is less than the second window threshold, then the clock signal delay is adjusted, and the second-level delay adjustment is re-executed.
[0115] In one embodiment, it is preferable to perform a first-level delay adjustment. If the first intersection window of the first delay windows of all data bytes is less than the third window threshold, then a second-level delay adjustment is triggered, and the second-level delay adjustment is performed with the second-level delay unit as the adjustment granularity.
[0116] Example 3:
[0117] refer to Figure 9 As shown, Figure 9 A flowchart illustrating a parallel interface data transmission method according to another embodiment is shown.
[0118] A parallel interface data transmission method is applied to the 3D packaged chip communication interface system of Embodiment 1 above. The 3D packaged chip communication interface system includes a 3D packaged chip and an FPGA bridge chip. The 3D packaged chip only has a digital parallel interface to exclude high-speed serializer / deserializer analog circuitry. The FPGA bridge chip and the 3D packaged chip are connected through the digital parallel interface. The FPGA bridge chip is the transmitting end, and the 3D packaged chip is the receiving end. For parallel communication between the FPGA bridge chip and the 3D packaged chip, with the FPGA bridge chip as the transmitting end and the 3D packaged chip as the receiving end, all timing calibration logic can be deployed on the FPGA side. When adapting to scenarios without a SerDes 3D packaged chip, the 3D packaged chip does not need to integrate additional calibration circuitry or SerDes analog circuitry.
[0119] The method includes the following steps:
[0120] Step S1: Perform first-level delay adjustment with the first-level delay unit as the adjustment granularity. For each data byte of the parallel interface signal, determine the first delay window of each data byte through automatic comparison operation, and set the number of first-level delay units enabled according to the first delay window.
[0121] Step S2: Perform second-level delay adjustment with the second-level delay unit as the adjustment granularity. For each bit of each data byte, determine the second delay window for each bit through an automatic comparison operation, and set the number of enabled second-level delay units according to the second delay window.
[0122] Each bit of the parallel interface signal is configured with a delay link. The delay link includes multiple sets of cascaded first-level delay units. Each set of first-level delay units includes multiple sets of cascaded second-level delay units. Each delay unit supports two states: on and off. In the on state, the signal generates a fixed delay after passing through the delay unit. In the off state, the signal bypasses the delay unit directly without delay.
[0123] The specific technical details of steps S1 and S2 are the same as those in Embodiment 1 above, and will not be repeated in this embodiment.
[0124] Step S3: Using the calibrated delay link, the transmitting end scrambles and ECC-encodes the parallel interface signal before sending it, and the receiving end performs ECC verification and descrambling on the received parallel interface signal before processing.
[0125] The process involves using a calibrated delay link, where the transmitting end scrambles and ECC-encodes the parallel interface signal before transmission, and the receiving end performs ECC verification and descrambling on the received parallel interface signal before processing. The specific steps are as follows:
[0126] On the transmitting end, the original parallel interface signal to be transmitted is grouped into units of 32 bits each, and each group of data is scrambled. An ECC check code is generated for each group of scrambled data, and the scrambled parallel data and the corresponding ECC check code are sent synchronously.
[0127] At the receiving end, the received parallel data and ECC checksum are grouped and matched in units of 32 bits of data + 6 bits of checksum. Error checking and correction are performed on each group of data according to the ECC checksum. If there is only 1 bit error, the error correction is completed automatically. If there are 2 or more errors, an interrupt signal is triggered to notify the sending end to retransmit. Descrambling is performed on each group of data after the checksum passes or the error correction is completed to restore the original parallel interface signal.
[0128] In a specific example, the parallel interface signals are shown in Table 1. The direction is based on the FPGA. The FPGA bridging chip provides output to the 3D package chip, and the 3D package chip provides input to the FPGA bridging chip.
[0129] The parallel interface signals include input signals and output signals. The input signals include parallel input data `parallel_i` and the corresponding ECC-encoded input data `ecc_i`. The output signals include parallel output data `parallel_o` and the corresponding ECC-encoded output data `ecc_o`. Here, `parallel_i` represents the data input from downstream to the FPGA, including read data (`read-data`), and `parallel_o` represents the data output from the FPGA downstream, including address `addr` and write data (`write-data`).
[0130] Table 1 Parallel Interface Signal Definitions
[0131]
[0132] First, the main link control unit adjusts the clk signal to the initial position, which can be set to 1 / 2 of the maximum delay. The maximum delay is the delay duration when all first and second level delay units are in the on state, i.e., the second maximum delay state.
[0133] First, disable the ECC function and perform timing calibration. Both parallel_o and parallel_i signals have corresponding delay links; adjust the delay of the corresponding delay links for parallel_o and parallel_i.
[0134] The master link control unit writes a specific set of test sequences into the SRAM of the slave link control unit on the receiving end side. Under different delay states, it determines whether the sequences can be correctly read from the SRAM in the specified order.
[0135] First, adjust the first byte, namely parallel_o[7:0] and parallel_i[7:0]. Adjust sequentially from the first minimum delay state to the first maximum delay state, using the first delay unit L1 as the unit; that is, enable N first-level delay units sequentially. These 16 bits of the first delay unit L1 are enabled synchronously, with the same adjustment step size. Then, for each data byte of the parallel interface signal, determine the first delay window for each data byte through an automatic comparison operation. The automatic comparison operation is implemented on the transmitting side master-side link control unit. After each adjustment, a "write-then-read" operation is performed on a specified location in the SRAM of the slave-side link control unit. "Write-then" means the master-side link control unit writes a fixed test sequence to the specified location in the SRAM of the slave-side link control unit via parallel_o. "Read-then" means that after writing, the master-side link control unit reads the data back from the same SRAM address via parallel_i. The written data reaches the slave-side link control unit via parallel_o, and then returns from the slave-side link control unit to the FPGA for sampling via parallel_i. The primary link control unit automatically compares the read-back data with the test sequence stored in the SRAM of the secondary link control unit. If they match, the data acquisition is considered correct, meaning the correct test sequence can be read, indicating that the delays in both parallel_o and parallel_i directions are within a reasonable range. Otherwise, the data acquisition is considered incorrect, meaning the correct test sequence was not read. In this "write-before-read" mode, the write path and read-back path share the same delay parameters.
[0136] When correct data is obtained for the first time through automatic comparison, i.e., the correct test sequence can be read, it is considered that the left boundary of the first delay window of parallel_o[7:0] and parallel_i[7:0] has been obtained simultaneously. The delay position corresponding to the current number of enabled units is recorded as the left boundary of the first delay window. The number of enabled units of the first delay unit continues to increase until the automatic comparison operation first encounters a data reading error. The delay position corresponding to the previous number of enabled units is then used as the right boundary of the first delay window. Thus, the first delay window of parallel_o[7:0] and parallel_i[7:0] is obtained. The first delay window of other bytes of parallel_o and parallel_i is found by following the above steps.
[0137] like Figure 10 As shown, Figure 10The diagram shows the L1 delay unit adjustment of the first data byte's parallel input data (parallel_i[7:0]), where the horizontal axis represents N L1-level delay units, i.e., L1_1 to L1_N. The √ symbol in the diagram marks the automatic comparison operation that passed at that delay position. Taking parallel_i[0] as an example, the √ is distributed from LI_3 to LI_5, indicating that the reasonable delay window for this byte has a left boundary of LI_3, i.e., the automatic comparison operation passed for the first time; the right boundary is LI_5. If we continue to increase the delay window, LI_6 will fail for the first time, so LI_5 is the right boundary.
[0138] Furthermore, the adjustment of parallel_o is based on the signal at the first stage flip-flop of parallel_o sampled by the master-side link control unit 3331 and the slave-side link control unit 3221; the adjustment of parallel_i is based on the signal at the first stage flip-flop of parallel_i and ecc_i sampled by the master-side link control unit 3331 of the FPGA, which meets the timing requirements; that is, all timing calibration judgments are located in the master-side link control unit 3331, reducing the logic circuit complexity of the slave-side link control unit 3221.
[0139] Furthermore, during the first delay adjustment process, if any byte cannot find the first delay window, that is, no matter how the first delay unit L1 is adjusted, the "write-then-read" method cannot read the correct test sequence. There are two possible reasons for this: 1) the signals for the parallel_i byte and the parallel_o byte cannot use the same delay parameters; 2) the delay of the clk signal needs adjustment.
[0140] At this point, exception handling will be executed in the following order of priority:
[0141] 1) Keep the delay parameter of parallel_o unchanged, and only adjust the delay parameter of parallel_i corresponding to this data byte. Then, read data only from the slave link control unit (i.e., not write-before-read, read-only), and re-verify whether the slave link control unit can correctly sample the data on parallel_i. If the data on parallel_i can be sampled correctly, it means that the direction of parallel_i itself is not the problem, and the direction of parallel_o may be problematic. Then continue to adjust the delay of parallel_o to find the delay window of parallel_o.
[0142] 2) If a suitable delay window still cannot be found, adjust the delay of the parallel_o corresponding to the data byte individually in "write-only" mode and re-verify. In "write-only" mode, the delay parameter of the input signal parallel_i remains fixed.
[0143] 3) If a suitable delay window still cannot be found, it is assumed that the delay of the clk signal is inappropriate. Adjust the delay of the parallel interface clock signal clk, and readjust from parallel_o[7:0] and parallel_i[7:0] to restore "write first, read later". Verify whether the data read back by the master link control unit is consistent with the test sequence stored in the SRAM of the slave link control unit.
[0144] 4) If a suitable first delay window cannot be found for parallel_o and parallel_i corresponding to a certain data byte, it is considered that "it cannot be satisfied at this operating frequency". Then, the operating clock frequency of the parallel interface is reduced, and all data bytes are re-traversed to determine the first delay window.
[0145] Once all data bytes of parallel_o and parallel_i have found the first delay window, calculate the first intersection window of the first delay windows of all data bytes. If the first intersection window of the first delay windows of all data bytes is greater than or equal to the first window threshold, then all first delay units are considered appropriate. At this time, configure the number of first-level delay units enabled in parallel_o and parallel_i to make the delay of the data byte equal to the delay value corresponding to the center point of the first intersection window, that is, configure the L1 level delay parameters of parallel_o and parallel_i as the center point of the first intersection window.
[0146] Continue as Figure 10 As shown, the overlapping area of each level √ from parallel_i[0] to parallel_i[7] is the "intersection window", and the center point of the intersection is set as the final delay parameter.
[0147] Furthermore, after the first-stage delay unit configuration of the parallel_o and parallel_i signals is completed, the ECC function is activated to start adjusting the first-stage delay unit configuration of the ecc_o and ecc_i signals. Since ecc_o and ecc_i are 12 bits, 6 bits of the signal are adjusted each time. The adjustment method is the same as that of parallel_o and parallel_i mentioned above, and will not be repeated here.
[0148] As the clock frequency increases, the first intersection window will inevitably shrink. When the first intersection window is smaller than the third window threshold, a second-level delay adjustment needs to be initiated for finer delay adjustments. The adjustment process for the second-level delay adjustment is the same as that for the first-level delay adjustment, and will not be described again here.
[0149] Furthermore, utilizing the aforementioned calibrated delay link, the transmitting end scrambles and ECC-encodes the parallel interface signal before transmission, while the receiving end performs ECC verification and descrambling on the received parallel interface signal. Specifically, the master-side link control unit adds ECC checksums (transmitted via the ECC_O signal) to the addr and write_data transmitted on the 64-bit parallel_o signal. ECC verification is performed in the ECC function of the slave-side link control unit; automatic error correction is applied for a 1-bit error, and if more than 1 bit of error occurs, the AI inference chip xPU notifies the FPGA bridge chip via an interrupt signal. ECC is performed every 32 bits in the 64-bit parallel_o data. The master-side link control unit performs ECC verification on the read_data transmitted on the parallel_i signal (transmitted via the ECC_i signal). The slave-side link control unit performs ECC on every 32 bits of the 64-bit parallel_i data. Simultaneously, the master-side link control unit scrambles the addr and write_data transmitted on the 64-bit parallel_o signal, applying scrambling to every 32 bits of the 64-bit data using the PRBS31 algorithm. The slave-side link control unit then descrambles the data. The master-side link control unit descrambles the read_data transmitted on the parallel_i signal, while the slave-side link control unit scrambles every 32 bits of the 64-bit data in parallel_i.
[0150] Furthermore, 3D packaged chips are typically manufactured using advanced processes. These advanced processes, such as PCI-X, require special I / O units to meet the 3.3V I / O requirement of their interface, and PCI-X data lines are 64-bit, not supporting 32-bit. 64-bit pins have a large number of pins, and power consumption is directly proportional to the bit width, resulting in high power consumption. The design described in this invention allows for the selection of both 32-bit and 64-bit, providing flexible bit width configuration capabilities and optimizing chip area and power consumption. The parallel interface supports 32-bit mode, using only the lower 32 bits of parallel_o and parallel_i, and the lower 6 bits of ecc_o and ecc_i. This reduces the number of PCB connections and further increases the clock frequency. This invention circumvents the PCI-X requirement, which has a 3.3V requirement. This invention can use ordinary digital I / O (1.8V or lower), completely avoiding the 3.3V requirement, allowing advanced process 3D packaged chips to directly use standard, low-voltage I / O units, simplifying chip design and manufacturing processes.
[0151] In summary, the 3D packaged chip communication interface system and corresponding parallel interface data transmission method provided by this invention add an FPGA bridge chip 33 between the 3D packaged chip 32 and the host 31, transferring the physical layer implementation of the high-speed interface from the 3D packaged chip to the external FPGA bridge chip. The parallel port of the FPGA replaces the PCIe serial port, completely removing high-speed serial interfaces such as PCIe / CXL from the 3D packaged chip, retaining only the digital parallel interface. Since the 3D packaged chip only needs to process parallel digital signals, its communication module is entirely composed of digital logic, including only interface conversion and on-chip buffer modules, and a simple control state machine, without the need for analog components. This simplifies the entire chip design process and is completely unaffected by the placement of the TSV. The parallel digital signals are insensitive to physical changes introduced by the TSV, eliminating the need for complex adjustments to adapt to the TSV. Furthermore, signal deskipation and level adjustment are implemented within the FPGA through a programmable delay chain, solving the timing problem of parallel transmission and ensuring the reliability of high-speed parallel communication. Furthermore, parallel transmission is implemented within the FPGA, and the parallel signal adjustment is divided into a first-level delay adjustment and a second-level delay adjustment. The first-level delay adjustment is a coarse adjustment to address inter-group deviations. The second-level delay adjustment is a fine adjustment to address intra-group deviations. Through these two levels of adjustment, a balance between calibration accuracy and overhead is achieved.
[0152] It should be noted that this method implementation can be implemented in conjunction with the above-described system implementation. The relevant technical details mentioned in the above system implementation remain valid in this method implementation, and will not be repeated here to avoid repetition.
[0153] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.
Claims
1. A parallel interface data transmission method based on a 3D packaged chip communication interface system, characterized in that, This invention relates to a communication interface system for 3D packaged chips. The system includes a 3D packaged chip and an FPGA bridge chip. The 3D packaged chip is equipped with only a digital parallel interface to exclude high-speed serializer / deserializer analog circuitry. The FPGA bridge chip is connected to the 3D packaged chip via the digital parallel interface. The FPGA bridge chip is the transmitting end, and the 3D packaged chip is the receiving end. The sending end performs the following steps: The first-level delay adjustment is performed with the first-level delay unit as the adjustment granularity. For each data byte of the parallel interface signal, the first delay window of each data byte is determined by automatic comparison operation, and the number of first-level delay units enabled is set according to the first delay window. The second-level delay adjustment is performed with the second-level delay unit as the adjustment granularity. For each bit of each data byte, the second delay window of each bit is determined by an automatic comparison operation, and the number of enabled second-level delay units is set according to the second delay window. The parallel interface signal is configured with a delay link, which includes multiple sets of cascaded first-level delay units. Each set of first-level delay units includes multiple cascaded second-level delay units. Each delay unit supports two states: on and off. In the on state, the signal passes through the delay unit and generates a fixed delay. In the off state, the signal directly bypasses the delay unit without delay. The process of determining the first or second delay window through automatic comparison includes: When the correct data is obtained for the first time through automatic comparison, the delay position corresponding to the current number of enabled functions is recorded as the left boundary of the delay window; Continue to increase the number of enabled delay units until the automatic comparison operation first encounters a data acquisition error, at which point the delay position corresponding to the previous number of enabled units is used as the right boundary of the delay window. The automatic comparison operation includes: The "write-then-read" mode includes: writing a test sequence to a specified location in the memory of the receiving end, reading back the data from the specified location on the sending end, and comparing the read-back data with the test sequence. If the two match, it is determined that the correct data has been obtained; otherwise, it is determined that the data acquisition is incorrect. The "read-only" mode includes: the transmitting end reads data from a preset address in the memory of the receiving end and compares the read-back data with the expected data; The "write-only" mode includes writing a test sequence to a specified location in the memory on the receiving end without the sending end reading back the data.
2. The method according to claim 1, characterized in that, When performing the first-level delay adjustment, starting from the first minimum delay state, the first-level delay units are enabled sequentially, one unit at a time. An automatic comparison operation is performed for each enabled unit until the first maximum delay state is reached; wherein the first minimum delay state is when all first-level delay units are in a bypass state, and the first maximum delay state is when all first-level delay units are in an enabled state; and / or, When performing the second-level delay adjustment, for the second-level delay units that are in the enabled state of the first-level delay units, starting from the second minimum delay state, all second-level delay units are put in the bypass state, and the second-level delay units are enabled sequentially, one second-level delay unit at a time. An automatic comparison operation is performed for each number of enabled units until the second maximum delay state is reached. The second minimum delay state is when all second-level delay units are in the bypass state, and the second maximum delay state is when all second-level delay units are in the enabled state.
3. The method according to claim 2, characterized in that, in, The parallel interface signal includes an input signal and an output signal. The input signal includes parallel input data and ECC-encoded input data corresponding to the parallel input data. The output signal includes parallel output data and ECC-encoded output data corresponding to the parallel output data. If a single data byte prevents the determination of the first delay window, exception handling is performed in the following order of priority: Fix the delay of the output signal corresponding to the data byte, adjust the delay of the input signal corresponding to the data byte separately, and re-verify the delay window of the input signal of the data byte through the "read-only" mode; Fix the delay of the input signal corresponding to the data byte, adjust the delay of the output signal corresponding to the data byte separately, and re-verify the delay window of the output signal of the data byte through the "write-only" mode; Adjust the delay of the parallel interface clock signal and re-traverse all data bytes to determine the first delay window; Reduce the operating clock frequency of the parallel interface and re-traverse all data bytes to determine the first delay window; And / or, If a single bit prevents the determination of the second delay window, exception handling is performed in the following order of priority: Fix the delay of the output signal corresponding to the bit, adjust the delay of the input signal corresponding to the bit separately, and re-verify the delay window of the input signal of the bit through the "read-only" mode; Fix the delay of the input signal corresponding to the bit, adjust the delay of the output signal corresponding to the bit separately, and re-verify the delay window of the output signal of the bit through the "write-only" mode; Adjust the delay of the parallel interface clock signal and re-traverse all bits to determine the second delay window; Reduce the operating clock frequency of the parallel interface and re-traverse all bits to determine the second delay window.
4. The method according to claim 1, characterized in that, The step of setting the number of first-level delay units enabled according to the first delay window includes: calculating the first intersection window of the first delay windows of all data bytes, and configuring the number of first-level delay units enabled for each data byte to make the delay of the data byte equal to the delay value corresponding to the center point of the first intersection window; And / or, The step of setting the number of second-level delay units enabled according to the second delay window includes: calculating the second intersection window of the second delay windows of all bits, and configuring the number of second-level delay units enabled for each bit so that the delay of the bit is equal to the delay value corresponding to the center point of the second intersection window.
5. The method according to claim 4, characterized in that, If the first intersection window of the first delay windows of all data bytes is less than the first window threshold and greater than or equal to the third window threshold, then adjust the clock signal delay and re-execute the first-level delay adjustment, wherein the third window threshold is less than the first window threshold; and / or, If the second intersection window of the second delay windows of all bits is less than the second window threshold, then the delay of the clock signal is adjusted, and the second-level delay adjustment is re-executed.
6. The method according to claim 5, characterized in that, If the first intersection window of the first delay window of all data bytes is less than the third window threshold, then the second-level delay adjustment is performed with the second-level delay unit as the adjustment granularity.
7. The method according to claim 1, characterized in that, Also includes: Using the calibrated delay link, the transmitting end scrambles and ECC-encodes the parallel interface signal before sending it, while the receiving end performs ECC verification and descrambles the received parallel interface signal before processing.
8. The method according to claim 7, characterized in that, The process of using a calibrated delay link, where the transmitting end scrambles and ECC-encodes the parallel interface signal before transmission, and the receiving end performs ECC verification and descrambling on the received parallel interface signal, includes the following steps: On the transmitting end, the original parallel interface signal to be transmitted is grouped into units of 32 bits each, and each group of data is scrambled. Generate an ECC checksum for each group of scrambled data, and send the scrambled parallel data and the corresponding ECC checksum synchronously. At the receiving end, the received parallel data and ECC check code are grouped and matched in units of 32 bits of data + 6 bits of check code. The error check and correction are performed on each group of data according to the ECC check code. If there is only a 1-bit data error, the error correction is completed automatically. If there are 2 or more errors, an interrupt signal is triggered to notify the sender to retransmit. For each set of data that has passed verification or completed error correction, descrambling is performed to restore the original parallel interface signal.