Self-adaptive fast punch-through clock network suitable for large-scale programmable device
By designing an adaptive fast through clock network, the problem of solidification of clock network structure in large-scale programmable devices is solved, efficient transmission and quality maintenance of clock signals are achieved, power consumption and clock jitter are reduced, and high-performance clock needs are met.
Patent Information
- Application Number
- CN202510077690.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, the clock network structure is relatively "cured", which is difficult to meet the demand for high-performance clocks of large-scale programmable devices, resulting in increased clock skew and jitter, and high power consumption.
An adaptive fast-through clock network is designed, including clock area, horizontal clock backbone, vertical clock backbone, horizontal clock buffer and vertical clock buffer. Through flexible wiring and buffer design, efficient transmission and quality maintenance of clock signals are achieved.
It effectively reduces power consumption, skew and jitter during clock transmission, provides high-performance clock resources, and meets the diverse clock performance requirements of large-scale programmable devices.
Smart Images

Figure CN120029767A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an adaptive fast-through clock network suitable for large-scale programmable devices, belonging to the technical field of integrated circuits. Background Art
[0002] In integrated circuit design, the clock signal is the benchmark for data signal transmission in the entire circuit. It is directly related to whether the timing of large-scale programmable devices such as programmable SOCs and FPGAs meets the design requirements. It is also the signal with the fastest flip speed, the largest driving load, and the longest transmission distance. It has an important impact on the function, performance, and stability of the chip. The clock network plays an important role in minimizing chip area and power consumption, and reducing clock skew and jitter while meeting synchronous timing.
[0003] In the previous design of programmable devices, the clock network is generally divided into regional clock networks and global clock networks. The clock source of the regional clock network is placed at the edge of the device, and the clock source of the global clock network is placed at the center of the device. However, this clock network structure is relatively "rigid". As the scale of the circuit increases, the clock signal has to drive millions or even tens of millions of drivers. Under this structure, the clock skew and clock delay at the end of the device continue to increase. At the same time, higher requirements are also placed on the frequency and power consumption of the clock, which becomes a difficulty and focus in the clock network design. Therefore, it is necessary to design a more flexible clock network to solve the above problems and meet the needs of large-scale programmable devices for high-performance clocks. Summary of the invention
[0004] The technical problem solved by the present invention is: to overcome the deficiencies of the prior art and provide an adaptive fast-through clock network suitable for large-scale programmable devices, which can fully reduce the power consumption, skew and jitter in the clock transmission process and provide high-performance clock resources suitable for large-scale programmable devices.
[0005] The objective of the present invention is achieved through the following technical solutions:
[0006] The present invention discloses an adaptive fast through-clock network suitable for large-scale programmable devices, comprising: a clock region, a horizontal clock trunk, a vertical clock trunk, a horizontal clock buffer and a vertical clock buffer; wherein,
[0007] Several clock regions are rectangular regions and are evenly arranged on the programmable device;
[0008] The horizontal clock trunk includes a horizontal distribution clock line and a horizontal wiring clock line; the horizontal distribution clock line is used to transmit the clock signal in the horizontal direction within the clock area, and the horizontal wiring clock line is used to transmit the clock signal in the horizontal direction across different clock areas;
[0009] A vertical clock trunk, including a vertical wiring clock line and a vertical distribution clock line; the vertical distribution clock line is used to transmit a clock signal in a vertical direction within a clock region, and the vertical wiring clock line is used to transmit a clock signal in a vertical direction across different clock regions;
[0010] Several horizontal clock stems traverse the center of each row of the clock region;
[0011] A plurality of vertical clock stems running vertically through the center of each column of the clock region;
[0012] A vertical clock buffer is arranged at the intersection of a horizontal clock trunk and a vertical clock trunk, and is used to implement signal switching between a horizontal distribution clock and a horizontal wiring clock on the horizontal clock trunk and a vertical wiring clock and a vertical distribution clock on the vertical clock trunk;
[0013] A horizontal clock buffer is provided on the horizontal clock trunk between two vertical clock buffers; it is used to maintain the clock signal quality during the clock signal transmission process on the horizontal distribution clock line and the horizontal routing clock line.
[0014] Furthermore, in the above network, a resource column clock line and a leaf buffer are included; wherein,
[0015] Several resource column clock lines are vertically connected to the horizontal distribution clock line to form a leaf node; the clock signal is transmitted to the leaf node by the horizontal distribution clock line, a leaf clock is generated through a leaf buffer, and the leaf clock signal is transmitted to the resource column of the programmable device through the resource column clock line.
[0016] Furthermore, in the above network, the horizontal wiring clock line intersects with the vertical wiring clock line to form a root node in each clock area; the root node of each clock area realizes bidirectional switching between the clock signal transmitted on the horizontal wiring clock line and the clock signal on the vertical wiring clock line.
[0017] Furthermore, in the above network, an area covering 50 to 70 CLB resources is divided into a clock area.
[0018] Further, in the above network, the horizontal clock buffer includes a first group of circuits and a second group of circuits;
[0019] The first group of circuits includes two-input NAND gates NAND0, NAND1, NAND2, inverters INV0, INV1, INV2, INV3, two-input NOR gate NOR0, PMOS tubes PP0, PP1 and NMOS tube NN0; wherein,
[0020] The first clock signal D is connected to the two-input NOR gate NOR0 and the two-input NAND gate NAND2 respectively;
[0021] The input enable signal OE1 is connected to the first input terminal of the two-input NAND gate NAND0; the input enable signal OE2 is connected to the second input terminal of the two-input NAND gate NAND0 and the first input terminal of NAND1; the input enable signal OE3 is connected to the second input terminal of the two-input NAND gate NAND1;
[0022] The output terminal of the two-input NAND gate NAND0 is connected to the input terminal of the two-input NOR gate INV0;
[0023] The output end of the two-input NOR gate INV0 is connected to the gate of the PMOS tube PP0;
[0024] The output end of the two-input NAND gate NAND1 is connected to the input end of the two-input NOR gate NOR0 and the input end of the inverter INV1; the output end of the inverter INV1 is connected to the input end of the two-input NAND gate NAND2;
[0025] The output end of the two-input NOR gate NOR0 is connected to the gate of the PMOS tube PP1 through the inverter INV2;
[0026] The output end of the two-input NAND gate NAND2 is connected to the gate of the NMOS tube NN0 through the inverter INV3;
[0027] The source of the PMOS tube PP0 is connected to a high level, and the drain is connected to the second clock signal ZN;
[0028] The source of the PMOS tube PP1 is connected to a high level, the source of the NMOS tube NN0 is grounded, and the drain of the PMOS tube PP1 is connected to the drain of the NMOS tube NN0 and the second clock signal ZN;
[0029] The second clock signal ZN is connected to the input end of the second group of circuits; the output end of the second group of circuits is connected to the first clock signal D.
[0030] Furthermore, in the above network, the second group of circuits includes two-input NAND gates NAND3, NAND4, NAND5, inverters INV4, INV5, INV6, INV7, two-input NOR gate NOR1, PMOS tubes PP2, PP3 and NMOS tube NN1; wherein,
[0031] The second clock signal ZN is connected to the two-input NOR gate NOR1 and the two-input NAND gate NAND5 respectively;
[0032] The input enable signal OE4 is connected to the first input terminal of the two-input NAND gate NAND3; the input enable signal OE5 is connected to the second input terminal of the two-input NAND gate NAND3 and the first input terminal of NAND4; the input enable signal OE6 is connected to the second input terminal of the two-input NAND gate NAND4;
[0033] An output terminal of a two-input NAND gate NAND3 is connected to an input terminal of a two-input NOR gate INV4;
[0034] The output end of the two-input NOR gate INV4 is connected to the gate of the PMOS tube PP2;
[0035] The output end of the two-input NAND gate NAND4 is connected to the input end of the two-input NOR gate NOR1 and the input end of the inverter INV5; the output end of the inverter INV5 is connected to the input end of the two-input NAND gate NAND5;
[0036] The output end of the two-input NOR gate NOR1 is connected to the gate of the PMOS tube PP3 through the inverter INV6;
[0037] The output end of the two-input NAND gate NAND5 is connected to the gate of the NMOS tube NN1 through the inverter INV7;
[0038] The source of the PMOS tube PP2 is connected to a high level, and the drain is connected to the first clock signal D;
[0039] The source of the PMOS tube PP3 is connected to a high level, the source of the NMOS tube NN1 is grounded, and the drain of the PMOS tube PP3 and the drain of the NMOS tube NN1 are connected to the first clock signal D as the output signal of the second group of circuits.
[0040] Furthermore, in the above network, when the input enable signals OE1, OE2, OE3 are 1, 1, 1 respectively and OE4, OE5, OE6 are 1, 1, 0, the clock signal is transmitted from the first clock signal D to the second clock signal ZN, and the second clock signal ZN is output; when the input enable signals OE1, OE2, OE3 are 1, 1, 0 and OE4, OE5, OE6 are 1, 1, 1, the clock signal is propagated from the second clock signal ZN to the first clock signal D, and the first clock signal D is output.
[0041] Furthermore, in the above network, the vertical clock buffer includes a two-input NAND gate NAND0, an inverter INV0, a NOR gate NOR0, PMOS tubes PP0-PP3, PP7, NMOS tubes NN0-NN2 and a Schmitt inverter; wherein,
[0042] The first clock signal D is connected to the input terminals of the two-input NAND gate NAND0 and the NOR gate NOR0;
[0043] The input enable signal OE1 is connected to the input terminal of NAND0;
[0044] The input enable signal OE2 is connected to the input terminal of NOR0 through INV0;
[0045] An input enable signal OE3 is connected to the gates of PP1, PP2 and PP3;
[0046] The output of NAND0 is connected to the gate of PP0; the source of PP0 is connected to a high level;
[0047] The output of NOR0 is connected to the gates of NNO and NN1, and the sources of NNO and NN1 are grounded;
[0048] The drains of NN0 and NN1 are connected to the drain of PP0;
[0049] The source of PP1 is connected to the high level, the drain of PP1 is connected to the source of PP2, the drain of PP2 is connected to the source of PP3, the drain of PP3 is connected to the drain of PP0 and the drains of NN1 and NN2, and is connected to the bidirectional clock signal Z1;
[0050] The gate and source of NN2 are grounded;
[0051] The bidirectional clock signal Z1 is connected to the input terminal of the Schmitt inverter;
[0052] The output of the Schmitt inverter is connected to the gates of PP7 and NN6;
[0053] The source of PP7 is connected to high level, and the source of NN6 is grounded;
[0054] The drains of PP7 and NN6 are connected to the output clock signal Z2;
[0055] Further, in the above network, the Schmidt inverter includes PMOS tubes PP4 to PP6 and NMOS tubes NN3 to NN5; wherein the gates of the PMOS tubes PP4 and PP5 and the NMOS tubes NN3 and NN4 are commonly connected to the clock signal Z1; the source of the PMOS tube PP4 is connected to a high level, the drain of the PMOS tube PP4 is connected to the source of the PMOS tube PP5 and the source of the PMOS tube PP6, and the drain of the PMOS tube PP6 is grounded; the source of the NMOS tube NN4 is grounded, the drain of the NMOS tube NN4 is connected to the source of the NMOS tube NN3 and the source of the NN5, and the drain of the NMOS tube NN5 is connected to a high level; the drain of the PMOS tube PP5 is connected to the drain of the NMOS tube NN3, and is connected to the gate of the PMOS tube PP6 and the gate of the NMOS tube NN5, which together serve as the output signal of the Schmidt inverter.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] (1) The present invention proposes an adaptive fast-through clock network suitable for large-scale programmable devices. Through fast-through routing clocks and flexibly distributed distribution clocks, the quality of clock transmission can be guaranteed, and the clock network transmission skew and clock jitter in large-scale programmable devices can be effectively reduced;
[0058] (2) In different usage scenarios, the present invention can design different clock signal transmission paths in a targeted manner by flexibly selecting wiring clocks, distribution clocks, and interconnection networks to meet the diverse clock performance requirements in large-scale programmable devices;
[0059] (3) The clock buffer designed by the present invention adopts a gated clock. By controlling the clock enable control terminal, only the power required for the clock signal from the clock source to the clock load is consumed, and the remaining unused clock networks are all in a closed state, effectively reducing the power consumption of the clock network;
[0060] (4) The root node of the present invention is adaptively placed. During the configuration of the programmable device, the clock root node can be automatically placed in the nearest clock area, effectively reducing the transmission delay of the clock network;
[0061] (5) The present invention uses a fast-through wiring clock, and the wiring clock line does not carry any load, so the clock signal can be quickly transmitted to the required clock area with minimum loss and generate a distribution clock;
[0062] (6) The present invention uses sufficient clock buffers, and each clock management unit has 8 clock multiplexing buffers (BUFGCTRL), 24 global clock buffers (BUFGCE), 4 clock frequency division buffers (BUFGCE_DIV), and non-user clock buffers automatically inserted by the configuration software, including horizontal clock buffers, vertical clock buffers and leaf clock buffers, to achieve refined clock management;
[0063] (7) The present invention adopts a "root node-leaf node" clock network architecture, and the clock line includes a wiring clock and a distribution clock, with two directions: horizontal and vertical. The wiring clock does not have any load and can achieve a fast penetration effect. It is mainly used to transmit the clock signal to the root node of the required clock area with high quality; the distribution clock connects the corresponding leaf nodes, and is mainly used to distribute the clock to the logical resource leaf clock. During the clock transmission process, the wiring clock first transmits the clock signal to the root node with minimal loss, and then the distribution clock transmits the clock signal to the corresponding leaf node, thereby generating a leaf clock and providing it to the corresponding logical resource. In this way, the clock network transmission skew and clock jitter can be effectively reduced, and only the power required from the clock source to the clock load is consumed, effectively reducing the power consumption of the clock network. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a schematic diagram of an adaptive fast punch-through clock network architecture applicable to large-scale programmable devices proposed by the present invention;
[0065] Figure 2 It is a schematic diagram of the relationship between resource utilization and jitter performance in clock region division;
[0066] Figure 3 It is a schematic diagram of the design of a horizontal clock buffer which is automatically inserted during the horizontal clock transmission process;
[0067] Figure 4 It is a schematic diagram of the design of clock paths, clock root nodes, and leaf nodes in the clock network within a clock region;
[0068] Figure 5 It is a schematic diagram of the design of a vertical clock buffer that is automatically inserted during the vertical clock transmission process;
[0069] Figure 6 It is a typical application clock route diagram of the clock signal from the clock source to the required logic resources in the programmable device architecture. DETAILED DESCRIPTION
[0070] The specific implementation of the present invention is described below in conjunction with the accompanying drawings so that those skilled in the art can better understand the present invention. It should be noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.
[0071] The word "exemplary" is used exclusively herein to mean "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise noted.
[0072] The present invention discloses an adaptive fast through-clock network suitable for large-scale programmable devices, comprising: a clock region, a horizontal clock trunk, a vertical clock trunk, a horizontal clock buffer and a vertical clock buffer; wherein,
[0073] Several clock regions are rectangular regions and are evenly arranged on the programmable device;
[0074] The horizontal clock trunk includes a horizontal distribution clock line and a horizontal wiring clock line; the horizontal distribution clock line is used to transmit the clock signal in the horizontal direction within the clock area, and the horizontal wiring clock line is used to transmit the clock signal in the horizontal direction across different clock areas;
[0075] A vertical clock trunk, including a vertical wiring clock line and a vertical distribution clock line; the vertical distribution clock line is used to transmit a clock signal in a vertical direction within a clock region, and the vertical wiring clock line is used to transmit a clock signal in a vertical direction across different clock regions;
[0076] Several horizontal clock stems traverse the center of each row of the clock region;
[0077] A plurality of vertical clock stems running vertically through the center of each column of the clock region;
[0078] A vertical clock buffer is arranged at the intersection of a horizontal clock trunk and a vertical clock trunk, and is used to implement signal switching between a horizontal distribution clock and a horizontal wiring clock on the horizontal clock trunk and a vertical wiring clock and a vertical distribution clock on the vertical clock trunk;
[0079] A horizontal clock buffer is provided on the horizontal clock trunk between two vertical clock buffers; it is used to maintain the clock signal quality during the clock signal transmission process on the horizontal distribution clock line and the horizontal routing clock line.
[0080] Preferably, it includes a resource column clock line and a leaf buffer; wherein,
[0081] Several resource column clock lines are vertically connected to the horizontal distribution clock line to form a leaf node; the clock signal is transmitted to the leaf node by the horizontal distribution clock line, a leaf clock is generated through a leaf buffer, and the leaf clock signal is transmitted to the resource column of the programmable device through the resource column clock line.
[0082] Preferably, the horizontally wired clock line intersects with the vertically wired clock line to form a root node in each clock region; the root node of each clock region realizes bidirectional switching between the clock signal transmitted on the horizontally wired clock line and the clock signal on the vertically wired clock line.
[0083] Preferably, an area covering 50 to 70 CLB resources is divided into a clock area.
[0084] Preferably, the horizontal clock buffer comprises a first group of circuits and a second group of circuits;
[0085] The first group of circuits includes two-input NAND gates NAND0, NAND1, NAND2, inverters INV0, INV1, INV2, INV3, two-input NOR gate NOR0, PMOS tubes PP0, PP1 and NMOS tube NN0; wherein,
[0086] The first clock signal D is connected to the two-input NOR gate NOR0 and the two-input NAND gate NAND2 respectively;
[0087] The input enable signal OE1 is connected to the first input terminal of the two-input NAND gate NAND0; the input enable signal OE2 is connected to the second input terminal of the two-input NAND gate NAND0 and the first input terminal of NAND1; the input enable signal OE3 is connected to the second input terminal of the two-input NAND gate NAND1;
[0088] The output terminal of the two-input NAND gate NAND0 is connected to the input terminal of the two-input NOR gate INV0;
[0089] The output end of the two-input NOR gate INV0 is connected to the gate of the PMOS tube PP0;
[0090] The output end of the two-input NAND gate NAND1 is connected to the input end of the two-input NOR gate NOR0 and the input end of the inverter INV1; the output end of the inverter INV1 is connected to the input end of the two-input NAND gate NAND2;
[0091] The output end of the two-input NOR gate NOR0 is connected to the gate of the PMOS tube PP1 through the inverter INV2;
[0092] The output end of the two-input NAND gate NAND2 is connected to the gate of the NMOS tube NN0 through the inverter INV3;
[0093] The source of the PMOS tube PP0 is connected to a high level, and the drain is connected to the second clock signal ZN;
[0094] The source of the PMOS tube PP1 is connected to a high level, the source of the NMOS tube NN0 is grounded, and the drain of the PMOS tube PP1 is connected to the drain of the NMOS tube NN0 and the second clock signal ZN;
[0095] The second clock signal ZN is connected to the input end of the second group of circuits; the output end of the second group of circuits is connected to the first clock signal D.
[0096] Preferably, the second group of circuits includes two-input NAND gates NAND3, NAND4, NAND5, inverters INV4, INV5, INV6, INV7, two-input NOR gate NOR1, PMOS tubes PP2, PP3 and NMOS tube NN1; wherein,
[0097] The second clock signal ZN is connected to the two-input NOR gate NOR1 and the two-input NAND gate NAND5 respectively;
[0098] The input enable signal OE4 is connected to the first input terminal of the two-input NAND gate NAND3; the input enable signal OE5 is connected to the second input terminal of the two-input NAND gate NAND3 and the first input terminal of NAND4; the input enable signal OE6 is connected to the second input terminal of the two-input NAND gate NAND4;
[0099] An output terminal of a two-input NAND gate NAND3 is connected to an input terminal of a two-input NOR gate INV4;
[0100] The output end of the two-input NOR gate INV4 is connected to the gate of the PMOS tube PP2;
[0101] The output end of the two-input NAND gate NAND4 is connected to the input end of the two-input NOR gate NOR1 and the input end of the inverter INV5; the output end of the inverter INV5 is connected to the input end of the two-input NAND gate NAND5;
[0102] The output end of the two-input NOR gate NOR1 is connected to the gate of the PMOS tube PP3 through the inverter INV6;
[0103] The output end of the two-input NAND gate NAND5 is connected to the gate of the NMOS tube NN1 through the inverter INV7;
[0104] The source of the PMOS tube PP2 is connected to a high level, and the drain is connected to the first clock signal D;
[0105] The source of the PMOS tube PP3 is connected to a high level, the source of the NMOS tube NN1 is grounded, and the drain of the PMOS tube PP3 and the drain of the NMOS tube NN1 are connected to the first clock signal D as the output signal of the second group of circuits.
[0106] Preferably, when the input enable signals OE1, OE2, OE3 are 1, 1, 1 respectively and OE4, OE5, OE6 are 1, 1, 0, the clock signal is transmitted from the first clock signal D to the second clock signal ZN, and the second clock signal ZN is output; when the input enable signals OE1, OE2, OE3 are 1, 1, 0 and OE4, OE5, OE6 are 1, 1, 1, the clock signal is propagated from the second clock signal ZN to the first clock signal D, and the first clock signal D is output.
[0107] Preferably, the vertical clock buffer includes a two-input NAND gate NAND0, an inverter INV0, a NOR gate NOR0, PMOS tubes PP0-PP3, PP7, NMOS tubes NN0-NN2 and a Schmitt inverter; wherein,
[0108] The first clock signal D is connected to the input terminals of the two-input NAND gate NAND0 and the NOR gate NOR0;
[0109] The input enable signal OE1 is connected to the input terminal of NAND0;
[0110] The input enable signal OE2 is connected to the input terminal of NOR0 through INV0;
[0111] An input enable signal OE3 is connected to the gates of PP1, PP2 and PP3;
[0112] The output of NAND0 is connected to the gate of PP0; the source of PP0 is connected to a high level;
[0113] The output of NOR0 is connected to the gates of NNO and NN1, and the sources of NNO and NN1 are grounded;
[0114] The drains of NN0 and NN1 are connected to the drain of PP0;
[0115] The source of PP1 is connected to the high level, the drain of PP1 is connected to the source of PP2, the drain of PP2 is connected to the source of PP3, the drain of PP3 is connected to the drain of PP0 and the drains of NN1 and NN2, and is connected to the bidirectional clock signal Z1;
[0116] The gate and source of NN2 are grounded;
[0117] The bidirectional clock signal Z1 is connected to the input terminal of the Schmitt inverter;
[0118] The output of the Schmitt inverter is connected to the gates of PP7 and NN6;
[0119] The source of PP7 is connected to high level, and the source of NN6 is grounded;
[0120] The drains of PP7 and NN6 are connected to the output clock signal Z2;
[0121] Preferably, the Schmitt inverter comprises PMOS tubes PP4 to PP6 and NMOS tubes NN3 to NN5; wherein the gates of the PMOS tubes PP4 and PP5 and the NMOS tubes NN3 and NN4 are commonly connected to the clock signal Z1; the source of the PMOS tube PP4 is connected to a high level, the drain of the PMOS tube PP4 is connected to the source of the PMOS tube PP5 and the source of the PMOS tube PP6, and the drain of the PMOS tube PP6 is grounded; the source of the NMOS tube NN4 is grounded, the drain of the NMOS tube NN4 is connected to the source of the NMOS tube NN3 and the source of the NMOS tube NN5, and the drain of the NMOS tube NN5 is connected to a high level; the drain of the PMOS tube PP5 is connected to the drain of the NMOS tube NN3, and is connected to the gate of the PMOS tube PP6 and the gate of the NMOS tube NN5, which together serve as the output signal of the Schmitt inverter.
[0122] Example
[0123] The programmable device as a whole adopts a structured and expandable stacked module layout, which strengthens the hierarchical characteristics of the overall architectural design. It aims to be able to be designed quickly and flexibly for different applications. Under this design concept, the system architecture is divided into multiple aspects such as power network, clock network, interconnection network, configuration circuit architecture, etc.
[0124] The overall clock network architecture of programmable devices divides the device into several clock regions in the horizontal and vertical directions. Each clock region contains resources such as CLB, DSP, BRAM, etc., and these resources are combined by column. The horizontal clock backbone of the clock network includes horizontal routing clocks and horizontal distribution clock lines, which cross the center of each row of clock regions; the vertical clock backbone of the clock network includes vertical routing clocks and vertical distribution clocks, which cross the clock region columns.
[0125] The clock can be quickly transmitted in each clock area through horizontal and vertical routing clocks between clock areas, and distributed to the programmable logic resources through horizontal distribution clock lines and vertical distribution clocks after entering the required clock area. Several leaf buffers are distributed within the clock area to convert the horizontal distribution clock lines into leaf clocks and transmit them to specific programmable logic resources.
[0126] In each clock region, the horizontal routing clock and the horizontal distribution clock line are located in the horizontal center of the clock region. The intersection area of the vertical clock and the horizontal clock is the clock root node, which is defined as the starting position of the accumulated clock skew, that is, the clock skew at this point is zero. At the root node of each clock region, bidirectional switching between the horizontal routing clock and the vertical routing clock can be achieved, and the horizontal routing clock or the vertical routing clock can be converted into a horizontal distribution clock line, and then the clock signal is transmitted to the resource column.
[0127] In each clock region, both horizontal clock and vertical clock can be transmitted bidirectionally. The clock signal can be transmitted to the leaf node by the horizontal distribution clock line, and the leaf clock is generated by the leaf buffer and finally transmitted to the programmable device resource column.
[0128] Clock signals can be transmitted between different clock regions through horizontal routing clocks, vertical routing clocks, horizontal distribution clock lines, and vertical distribution clocks. Routing clocks do not have load circuits and are mainly used to quickly send clocks to the clock root node, achieve fast clock penetration between clock regions, and ensure the quality of clock signals; distribution clocks connect corresponding leaf nodes and are mainly used to distribute clocks to the required programmable device logic resources.
[0129] In order to meet the diverse clock performance requirements in large-scale programmable devices, there are three transmission paths for clock signals:
[0130] The first type, in typical cases, the clock source is transmitted to the root node of the clock region through the horizontal and vertical wiring clocks, and then transmitted from the clock root node to the vertical distribution clock, and then transmitted to the upper and lower required clock regions through the vertical distribution clock, and then converted into a horizontal distribution clock line, and transmitted to the left and right required clock regions through the horizontal distribution clock line, and finally transmitted to the leaf buffer by the horizontal distribution clock line, and then transmitted to each logic resource through the leaf buffer;
[0131] Second, in applications where the insertion delay of the clock buffer needs to be reduced, the clock source can be directly transmitted to the leaf buffer via the horizontal distribution clock line, and then transmitted to each logic resource via the leaf buffer;
[0132] The third type is that in applications with lower requirements on clock performance and quality, the data can be directly transmitted to the leaf buffer of the clock area through the interconnection network, and then transmitted to each logic resource through the leaf buffer.
[0133] As mentioned above, along the routing clock, the clock signal can be transmitted to the root node of the distribution clock, from which the clock signal is transmitted as one or more vertical distribution clocks. From the vertical distribution clock, the clock signal is transmitted as one or more horizontal distribution clock lines. Finally, from the horizontal distribution clock line, the clock signal is transmitted to one or more leaf clocks. For all signal sources, the clock signal can be transmitted directly to the routing clock, or directly transmitted to at least one clock root through interconnection, or indirectly transmitted to at least one clock root using one or more routing clocks.
[0134] The clock network also has a corresponding clock buffer design. The clock buffer is divided into a user clock buffer and a non-user clock buffer. The user clock buffer is mainly used for each clock source to drive each clock transmission path, which is optional and controllable by the user; the non-user clock buffer is mainly used for mutual driving between each clock transmission path, which is selected and controlled by the programmable device configuration software layout and routing.
[0135] The user buffer includes a clock multiplexer buffer (BUFGCTRL), a global clock buffer (BUFGCE), and a clock divider buffer (BUFGCE_DIV). BUFGCTRL, BUFGCE, and BUFGCE_DIV are located in a clock management unit column. The clock column in each clock region contains 8 BUFGCTRLs, 24 BUFGCEs, and 4 BUFGCE_DIVs.
[0136] BUFGCTRL selects one of the input signals I0 and I1 to the output signal O through six control signals S0, S1, CE0, CE1, IGNORE0, and IGNORE1. BUFGCTRL is mainly used for glitch-free switching between two clock sources, or to cut off the clock from a failed clock source.
[0137] BUFGCE selects whether to transfer the clock from the input signal I to the output signal O through the enable control signal CE. BUFGCE is a general global clock buffer.
[0138] BUFGCE_DIV selects whether to transfer the clock from input signal I to output signal O by enabling control signal CE, and CLR is an asynchronous reset signal. BUFGCE_DIV has a frequency division function and can achieve integer frequency division from 1 to 8.
[0139] The non-user clock buffer includes a horizontal clock buffer, a vertical clock buffer, and a leaf clock buffer. All three clock buffers are tri-state clock buffers. When the clock enable signal CE is valid, the clock is transmitted from input I to output O; when CE is invalid, output O is in a high impedance state, blocking the transmission of the clock on the clock transmission path.
[0140] During the transmission of horizontal wiring clock and horizontal distribution clock line, a horizontal clock buffer needs to be added after a certain transmission length to maintain the clock signal quality; when the horizontal wiring clock is converted into a horizontal distribution clock line, a vertical wiring clock or a vertical distribution clock, and during the transmission of the vertical wiring clock and the vertical distribution clock, a vertical clock buffer needs to be added to realize the switching of different clock types and maintain the clock signal quality.
[0141] Figure 1 The schematic diagram of the adaptive fast-through clock network architecture for large-scale programmable devices proposed by the present invention is shown in the figure. The overall clock architecture is a grid structure, and the programmable device is divided into several clock areas in the horizontal and vertical directions. Each large square in the figure is a clock area. The horizontal clock trunk of the clock network includes 24 horizontal wiring clocks and 24 horizontal distribution clock lines, which cross the center of each row of clock areas, and the vertical clock trunk includes vertical wiring clocks and vertical distribution clocks that cross the clock area columns. During the transmission of horizontal wiring clocks and horizontal distribution clock lines, a horizontal clock buffer needs to be added after a certain transmission length, such as the small triangle in the figure, to maintain the quality of the clock signal; when the horizontal wiring clock is converted into a horizontal distribution clock line, a vertical wiring clock or a vertical distribution clock, and during the transmission of the vertical wiring clock and the vertical distribution clock, a vertical clock buffer needs to be added, such as the small square in the figure, to achieve the switching of different clock line types and maintain the quality of the clock signal.
[0142] Figure 2This is a diagram showing the relationship between resource utilization and jitter performance in clock region division. The division of clock regions is based on the number of CLB resources. The number of CLB resources in a clock region affects resource utilization and terminal jitter performance. The more CLB resources in a clock region, the greater the load of the root node driving circuit in the clock region and the greater the terminal jitter; the fewer CLB resources in a clock region, the more clock regions will be divided, the more root nodes will need to be introduced, and the resource utilization will decrease. After simulating the load capacity of the driving circuit in the root node, it is determined that a clock region containing 60 CLB resources is the most appropriate, at which time the terminal jitter performance is 50ps and the resource utilization is 98%.
[0143] Figure 3 The horizontal clock buffer is designed to be automatically inserted during the horizontal clock transmission process. The horizontal clock buffer is composed of two groups of circuits with the same structure and has a bidirectional transmission function. When the input enable signals OE1, OE2, OE3 are 1, 1, 1 and OE4, OE5, OE6 are 1, 1, 0, the clock signal is transmitted from D to ZN; when the input enable signals OE1, OE2, OE3 are 1, 1, 0 and OE4, OE5, OE6 are 1, 1, 1, the clock signal is transmitted from ZN to D. The horizontal clock buffer can be automatically inserted by the programmable device configuration software during the transmission of the horizontal wiring clock and the horizontal distribution clock line to maintain the quality of the horizontal clock signal during transmission and enhance the driving capability.
[0144] The first group of circuits includes three two-input NAND gates NAND0, NAND1, NAND2, four inverters INV0, INV1, INV2, INV3, a two-input NOR gate NOR0, two PMOS tubes PP0, PP1, and an NMOS tube NN0.
[0145] The connection relationship is as follows: the first clock signal D is sent to NOR0 and NAND2 respectively, the input enable signal OE1 is sent to NAND0, the input enable signal OE2 is sent to NAND0 and NAND1, and the input enable signal OE3 is sent to NAND1. The output of NAND0 is sent to the gate of PP0 through INV0, and the output of NAND1 is sent to NOR0 in one way and to NAND2 in the other way through INV1. The output of NOR0 is sent to the gate of PP1 through INV2, and the output of NAND2 is sent to the gate of NN0 through INV3. The source of PP0 is connected to a high level, and the drain is connected to the second clock signal ZN. The source of PP1 is connected to a high level, the source of NN0 is grounded, PP1 is connected to the drain of NN0, and then the second clock signal ZN is sent.
[0146] The second group of circuits includes three two-input NAND gates NAND3, NAND4, NAND5, four inverters INV4, INV5, INV6, INV7, a two-input NOR gate NOR1, two PMOS tubes PP2, PP3, and an NMOS tube NN1.
[0147] The connection relationship is as follows: the input clock signal ZN is sent to NOR1 and NAND5 respectively, the input enable signal OE4 is sent to NAND3, the input enable signal OE5 is sent to NAND3 and NAND4, and the input enable signal OE6 is sent to NAND4. The output of NAND3 is sent to the gate of PP2 through INV4, and the output of NAND4 is sent to NOR1 in one way and to NAND5 in the other way through INV5. The output of NOR1 is sent to the gate of PP3 through INV6, and the output of NAND5 is sent to the gate of NN1 through INV7. The source of PP2 is connected to a high level, and the drain is connected to the output clock signal D. The source of PP3 is connected to a high level, the source of NN1 is grounded, PP3 is connected to the drain of NN1, and then the output clock signal D is sent.
[0148] Figure 4 Design schematic diagram of clock path, clock root node and leaf node in clock network in a clock region. Horizontal routing clock and horizontal distribution clock line are located in the horizontal center of the clock region. The intersection of vertical clock and horizontal clock is the clock root node. Both horizontal clock and vertical clock can be transmitted bidirectionally in the clock region. The clock signal can be transmitted to the leaf node by the horizontal distribution clock line, and the leaf clock is generated by the leaf buffer and finally transmitted to the resource column of the programmable device. At the root node of each clock region, bidirectional switching between horizontal routing clock and vertical routing clock can be realized, and the horizontal routing clock or vertical routing clock can be converted to the horizontal distribution clock line, and then the clock signal is transmitted to the resource column.
[0149] Figure 5 The vertical clock buffer is designed for automatic insertion in the clock root node. The vertical clock buffer includes a two-input NAND gate NAND0, an inverter INV0, a NOR gate NOR0, eight PMOS tubes PP0~PP7, and seven NMOS tubes NN0~NN6.
[0150] The connection relationship is as follows: the first clock signal D is sent to NAND0 and NOR0 respectively, the input enable signal OE1 is sent to NAND0, the input enable signal OE2 is sent to NOR0 through INV0, and the input enable signal OE3 is sent to the gates of PP1, PP2, and PP3. The output of NAND0 is sent to the gate of PP0, and the source of PP0 is connected to a high level. The output of NOR0 is sent to the gates of NNO and NN1, the sources of NNO and NN1 are grounded, and the drains of NN0, NN1, and PP0 are connected. The source of PP1 is connected to a high level, the drain of PP1 is connected to the source of PP2, the drain of PP2 is connected to the source of PP3, the drain of PP3 is connected to the drains of PP0, NN1, and NN2, and the bidirectional clock signal Z1 is sent. The gate and source of NN2 are both grounded. The bidirectional clock signal Z1 is then connected to the input of the Schmitt inverter, the output of the Schmitt inverter is connected to the gates of PP7 and NN6, the source of PP7 is connected to a high level, the source of NN6 is grounded, and the drains of PP7 and NN7 are connected to feed the output clock signal Z2 together.
[0151] The Schmitt inverter is composed of three PMOS tubes PP4, PP5, PP6 and three NMOS tubes NN3, NN4, NN5. The gates of PP4, PP5, NN3, and NN4 are connected to the clock signal Z1. The source of PP4 is connected to a high level, the drain is connected to the source of PP5 and PP6, and the drain of PP6 is grounded. The source of NN4 is grounded, the drain is connected to the source of NN3 and NN5, and the drain of NN5 is connected to a high level. PP5 is connected to the drain of NN3, and is connected to the gates of PP6 and NN5, which together serve as the output signal of the Schmitt inverter.
[0152] The vertical clock buffer is effective when the input enable signals OE1, OE2, and OE3 are all 1, the first clock signal D is the horizontal wiring clock, the bidirectional clock signal Z1 is the vertical wiring clock or the vertical distribution clock, and the output clock signal Z2 is the horizontal wiring clock or the horizontal distribution clock line. The vertical clock buffer can be automatically inserted by the programmable device configuration software at the intersection of the horizontal wiring clock, the horizontal distribution clock line and the vertical wiring clock, and the vertical distribution time to achieve the conversion of different clock line types and clock line directions, and reduce clock jitter through the Schmitt inverter to enhance the quality of the clock signal.
[0153] Figure 6This is a typical application clock flow diagram of the clock signal from the clock source to the required logic resources in the programmable device architecture. The clock signal enters from the constraint pin of the IO column in the programmable device and enters the clock network through the global clock buffer in the clock management unit column. First, it is converted into a horizontal wiring clock. After a certain transmission distance, the programmable device configuration software automatically inserts a horizontal clock buffer to maintain the signal quality and continues to transmit it to the vertical clock buffer. One signal is converted into a horizontal distribution clock line and transmitted to the leaf buffer to provide to the programmable device resources. The other signal is converted into a vertical wiring clock for cross-clock region propagation. The vertical clock needs to pass through the vertical clock buffer when it propagates across clock regions. When it is transmitted to the required clock region, it is simultaneously converted into a horizontal distribution clock line and provided to the leaf node and the vertical wiring clock for further transmission. When the vertical wiring clock enters the last level of the clock region that needs to be transmitted, it is converted into a vertical distribution clock, and converted into a horizontal distribution clock line in the vertical clock buffer, and transmitted to the leaf buffer to provide to the required programmable device resources.
[0154] The vertical clock buffer can be automatically inserted by the programmable device configuration software at the intersection of the horizontal wiring clock, horizontal distribution clock line and the vertical wiring clock, vertical distribution time, to achieve the conversion of different clock line types and clock line directions, and reduce clock jitter through the Schmitt inverter to enhance the quality of the clock signal.
[0155] All the above clock buffers contain a clock enable control terminal. With the gated clock technology, when the module is not working, the clock signal of this part of the module is prohibited from flipping through the control signal. Only when the control signal is valid can the clock of this part of the module work normally. By controlling the clock enable terminal, refined clock management is achieved and the power consumption of the clock network is reduced.
[0156] Although the above describes the illustrative specific embodiments of the present invention to facilitate those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations using the concept of the present invention are protected.
Claims
1. An adaptive fast-through clock network suitable for large-scale programmable devices, characterized in that: include: a clock region, a horizontal clock stem, a vertical clock stem, a horizontal clock buffer, and a vertical clock buffer; wherein, Several clock regions are rectangular regions and are evenly arranged on the programmable device; The horizontal clock trunk includes a horizontal distribution clock line and a horizontal wiring clock line; the horizontal distribution clock line is used to transmit the clock signal in the horizontal direction within the clock area, and the horizontal wiring clock line is used to transmit the clock signal in the horizontal direction across different clock areas; A vertical clock trunk, including a vertical wiring clock line and a vertical distribution clock line; the vertical distribution clock line is used to transmit a clock signal in a vertical direction within a clock region, and the vertical wiring clock line is used to transmit a clock signal in a vertical direction across different clock regions; Several horizontal clock stems traverse the center of each row of the clock region; A plurality of vertical clock stems running vertically through the center of each column of the clock region; A vertical clock buffer is arranged at the intersection of a horizontal clock trunk and a vertical clock trunk, and is used to implement signal switching between a horizontal distribution clock and a horizontal wiring clock on the horizontal clock trunk and a vertical wiring clock and a vertical distribution clock on the vertical clock trunk; A horizontal clock buffer is provided on the horizontal clock trunk between two vertical clock buffers; it is used to maintain the clock signal quality during the clock signal transmission process on the horizontal distribution clock line and the horizontal routing clock line.
2. The adaptive fast through-clock network suitable for large-scale programmable devices according to claim 1, characterized in that: including resource column clock lines and leaf buffers; wherein, Several resource column clock lines are vertically connected to the horizontal distribution clock line to form a leaf node; the clock signal is transmitted to the leaf node by the horizontal distribution clock line, a leaf clock is generated through a leaf buffer, and the leaf clock signal is transmitted to the resource column of the programmable device through the resource column clock line.
3. The adaptive fast through-clock network suitable for large-scale programmable devices according to claim 1, characterized in that: The horizontally wired clock line intersects with the vertically wired clock line to form a root node in each clock region; the root node of each clock region realizes bidirectional switching between the clock signal transmitted on the horizontally wired clock line and the clock signal on the vertically wired clock line.
4. The adaptive fast punch-through clock network suitable for large-scale programmable devices according to claim 1, characterized in that: An area covering 50 to 70 CLB resources is divided into clock regions.
5. The adaptive fast through-clock network suitable for large-scale programmable devices according to claim 1, characterized in that: The horizontal clock buffer includes a first group of circuits and a second group of circuits; The first group of circuits includes two-input NAND gates NAND0, NAND1, NAND2, inverters INV0, INV1, INV2, INV3, two-input NOR gate NOR0, PMOS tubes PP0, PP1 and NMOS tube NN0; wherein, The first clock signal D is connected to the two-input NOR gate NOR0 and the two-input NAND gate NAND2 respectively; The input enable signal OE1 is connected to the first input terminal of the two-input NAND gate NAND0; the input enable signal OE2 is connected to the second input terminal of the two-input NAND gate NAND0 and the first input terminal of NAND1; the input enable signal OE3 is connected to the second input terminal of the two-input NAND gate NAND1; The output terminal of the two-input NAND gate NAND0 is connected to the input terminal of the two-input NOR gate INV0; The output end of the two-input NOR gate INV0 is connected to the gate of the PMOS tube PP0; The output end of the two-input NAND gate NAND1 is connected to the input end of the two-input NOR gate NOR0 and the input end of the inverter INV1; the output end of the inverter INV1 is connected to the input end of the two-input NAND gate NAND2; The output end of the two-input NOR gate NOR0 is connected to the gate of the PMOS tube PP1 through the inverter INV2; The output end of the two-input NAND gate NAND2 is connected to the gate of the NMOS tube NN0 through the inverter INV3; The source of the PMOS tube PP0 is connected to a high level, and the drain is connected to the second clock signal ZN; The source of the PMOS tube PP1 is connected to a high level, the source of the NMOS tube NN0 is grounded, and the drain of the PMOS tube PP1 is connected to the drain of the NMOS tube NN0 and the second clock signal ZN; The second clock signal ZN is connected to the input end of the second group of circuits; the output end of the second group of circuits is connected to the first clock signal D.
6. The adaptive fast through-clock network suitable for large-scale programmable devices according to claim 5, characterized in that: The second group of circuits includes two-input NAND gates NAND3, NAND4, NAND5, inverters INV4, INV5, INV6, INV7, two-input NOR gate NOR1, PMOS tubes PP2, PP3 and NMOS tube NN1; wherein, The second clock signal ZN is connected to the two-input NOR gate NOR1 and the two-input NAND gate NAND5 respectively; The input enable signal OE4 is connected to the first input terminal of the two-input NAND gate NAND3; the input enable signal OE5 is connected to the second input terminal of the two-input NAND gate NAND3 and the first input terminal of NAND4; the input enable signal OE6 is connected to the second input terminal of the two-input NAND gate NAND4; An output terminal of a two-input NAND gate NAND3 is connected to an input terminal of a two-input NOR gate INV4; The output end of the two-input NOR gate INV4 is connected to the gate of the PMOS tube PP2; The output end of the two-input NAND gate NAND4 is connected to the input end of the two-input NOR gate NOR1 and the input end of the inverter INV5; the output end of the inverter INV5 is connected to the input end of the two-input NAND gate NAND5; The output end of the two-input NOR gate NOR1 is connected to the gate of the PMOS tube PP3 through the inverter INV6; The output end of the two-input NAND gate NAND5 is connected to the gate of the NMOS tube NN1 through the inverter INV7; The source of the PMOS tube PP2 is connected to a high level, and the drain is connected to the first clock signal D; The source of the PMOS tube PP3 is connected to a high level, the source of the NMOS tube NN1 is grounded, and the drain of the PMOS tube PP3 and the drain of the NMOS tube NN1 are connected to the first clock signal D as the output signal of the second group of circuits.
7. The adaptive fast through-clock network suitable for large-scale programmable devices according to claim 6, characterized in that: When the input enable signals OE1, OE2, OE3 are 1, 1, 1 respectively and OE4, OE5, OE6 are 1, 1, 0, the clock signal is transmitted from the first clock signal D to the second clock signal ZN, and the second clock signal ZN is output; when the input enable signals OE1, OE2, OE3 are 1, 1, 0 and OE4, OE5, OE6 are 1, 1, 1, the clock signal is propagated from the second clock signal ZN to the first clock signal D, and the first clock signal D is output.
8. The adaptive fast through-clock network suitable for large-scale programmable devices according to claim 1, characterized in that: The vertical clock buffer includes a two-input NAND gate NAND0, an inverter INV0, a NOR gate NOR0, PMOS tubes PP0-PP3, PP7, NMOS tubes NN0-NN2 and a Schmitt inverter; wherein, The first clock signal D is connected to the input terminals of the two-input NAND gate NAND0 and the NOR gate NOR0; The input enable signal OE1 is connected to the input terminal of NAND0; The input enable signal OE2 is connected to the input terminal of NOR0 through INV0; An input enable signal OE3 is connected to the gates of PP1, PP2 and PP3; The output of NAND0 is connected to the gate of PP0; the source of PP0 is connected to a high level; The output of NOR0 is connected to the gates of NNO and NN1, and the sources of NNO and NN1 are grounded; The drains of NN0 and NN1 are connected to the drain of PP0; The source of PP1 is connected to the high level, the drain of PP1 is connected to the source of PP2, the drain of PP2 is connected to the source of PP3, the drain of PP3 is connected to the drain of PP0 and the drains of NN1 and NN2, and is connected to the bidirectional clock signal Z1; The gate and source of NN2 are grounded; The bidirectional clock signal Z1 is connected to the input terminal of the Schmitt inverter; The output of the Schmitt inverter is connected to the gates of PP7 and NN6; The source of PP7 is connected to high level, and the source of NN6 is grounded; The drains of PP7 and NN6 are connected to the output clock signal Z2.
9. The adaptive fast punch-through clock network suitable for large-scale programmable devices according to claim 8, characterized in that: The Schmitt inverter comprises PMOS tubes PP4-PP6 and NMOS tubes NN3-NN5; wherein the gates of the PMOS tubes PP4 and PP5 and the NMOS tubes NN3 and NN4 are commonly connected to the clock signal Z1; the source of the PMOS tube PP4 is connected to a high level, the drain of the PMOS tube PP4 is connected to the source of the PMOS tube PP5 and the source of the PMOS tube PP6, and the drain of the PMOS tube PP6 is grounded; the source of the NMOS tube NN4 is grounded, the drain of the NMOS tube NN4 is connected to the source of the NMOS tube NN3 and the source of the NMOS tube NN5, and the drain of the NMOS tube NN5 is connected to a high level; the drain of the PMOS tube PP5 is connected to the drain of the NMOS tube NN3, and is connected to the gate of the PMOS tube PP6 and the gate of the NMOS tube NN5, which are jointly used as the output signal of the Schmitt inverter.