Data channel clock tree circuit, UCIe core and chip
By using a data channel clock tree circuit in the chip technology and generating a delayed clock signal with a preset phase difference using a delay buffer group, the problem of excessive current change rate when the potential of multiple data channel signals reverses is solved, a more stable power supply network is achieved, and the power integrity of the chip is optimized.
Patent Information
- Application Number
- CN202522573435.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2035-12-04
AI Technical Summary
In chip technology, the signal potential reversal between multiple data channels generates a huge instantaneous current change rate, resulting in excessive power supply noise and affecting the power integrity of the chip.
A data channel clock tree circuit is adopted, which generates a delayed clock signal with a preset phase difference through a delay buffer group to control the timing of the signal potential reversal of the data channel, so that the signal potential reversal times of multiple data channels are staggered, thereby reducing the instantaneous changes in the power supply network current.
It reduces the rate of change of power supply network current, decreases instantaneous power supply noise, and optimizes the power integrity of the chip.
Smart Images

Figure CN223784676U_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the chip technical field, in particular to a data channel clock tree circuit, a UCIe chiplet and a chip. BACKGROUND
[0002] With the continuous development of chip integrated circuit technology, a single large chip in a single form is limited by the lithography mask area, yield collapse and cost exponential rise. At the same time, the evolution of advanced process nodes is increasingly expensive and the marginal benefit is seriously diminishing. In this case, the single large chip of the system level is split into multiple miniaturized chiplets along the physical and functional boundaries, and different processes, different manufacturers, different functions of the chiplets or dies are reassembled into a large chip using advanced packaging and standardized interconnection technology, so as to continue to improve the integration and performance while reducing the yield loss and the rising rate of wafer cost, which has become an important means to continue Moore's law.
[0003] This chiplet technology needs to interconnect chiplets of different processes, different manufacturers and different functions, so ensuring good communication between different chiplets becomes a crucial link in the chiplet technology. CONTENT OF THE UTILITY MODEL
[0004] Therefore, the present disclosure provides a data channel clock tree circuit, a UCIe chiplet and a chip to help reduce the current fluctuation when the signal potential of the data channel of the chiplet is reversed, help reduce the current change rate, and further help optimize the power integrity of the chip.
[0005] According to an aspect of an embodiment of the present disclosure, a data channel clock tree circuit is provided, comprising:
[0006] a clock source;
[0007] a delay buffer group, the delay buffer group is two, both of the delay buffer groups are coupled to the clock source to generate at least two delay clock signals respectively, and both of the delay buffer groups are respectively coupled to different data channels in a plurality of data channels through different clock signal lines to send the at least two delay clock signals generated respectively to the data channels coupled to the clock signal lines respectively.
[0008] In a possible implementation, each of the delay buffer groups comprises:
[0009] The clock signal output end of any one of the at least two delay buffers is coupled to at least one of the plurality of data channels through a clock signal line associated with the any one of the at least two delay buffers, and the clock signal outputted by the clock source sequentially passes through the at least two delay buffers in series to generate the at least two delay clock signals, which are respectively provided to at least one of the plurality of data channels.
[0010] In a possible implementation, the driving capability of each of the delay buffers in the two delay buffer groups is the same.
[0011] In a possible implementation, the plurality of data channels are arranged in parallel to form a channel array.
[0012] The clock source is arranged at a center position of the channel array.
[0013] The two delay buffer groups are respectively extended to the direction of the data channels at the two side edges of the channel array with the clock source as the center.
[0014] In a possible implementation, each of the delay buffer groups includes at least two delay buffers in series, and the number of the delay buffers included in the two delay buffer groups is the same, and the delay buffers included in the two delay buffer groups are symmetrically arranged with respect to the center line of the channel array.
[0015] In a possible implementation, the plurality of data channels are divided into a plurality of channel groups, and the number of data channels included in each of the channel groups is equal.
[0016] Each of the delay buffers in the two delay buffer groups is associated with each of the channel groups, and the associated delay buffer and the channel group are arranged adjacently, and between the associated delay buffer and the channel group, the clock signal output end of the delay buffer is coupled to each of the data channels in the channel group through a clock signal line associated with the delay buffer, to provide a delay clock signal for each of the data channels in the channel group.
[0017] In a possible implementation, in the plurality of channel groups, the delay clock signals obtained from the channel groups close to the center position of the channel array to the channel groups far away from the center position of the channel array are symmetrically delayed with respect to the center position of the channel array by a set delay step, and the set delay step is the delay of one delay buffer to the clock signal.
[0018] In one possible implementation, the data channel is at least one of a UCIe data channel, a PCIe data channel, and an XSR data channel.
[0019] According to another aspect of the embodiments of the present disclosure, a UCIe die is provided, comprising the data channel clock tree circuit according to any one of the preceding embodiments.
[0020] According to another aspect of the embodiments of the present disclosure, a chip is provided, comprising the data channel clock tree circuit according to any one of the preceding embodiments.
[0021] As can be seen from the above solutions, the data channel clock tree circuit, the UCIe die and the chip of the present disclosure achieve the sending of different delay clock signals to different data channels, so that the clock signals between the multiple data channels are not strictly synchronized and there is a preset phase difference due to the delay of the delay buffer group. Because the timing of the inversion of the signal potential of the data channel is driven and controlled by the transition edge of the clock signal, when the data channel clock tree circuit, the UCIe die and the chip of the present disclosure are adopted, the clock signals between the multiple data channels are not strictly synchronized and there is a preset delay, so that when the signal potential of the multiple data channels is inverted, the inversion timing of the signal potential of the multiple data channels has a preset delay difference, and further, at any signal potential inversion timing, only a part of the data channels performs signal potential inversion, and the execution of the signal potential inversion between the data channels is staggered due to the delay clock signal, compared with the simultaneous execution of the signal potential inversion of all the data channels at the same timing, the instantaneous change of the power supply network current is reduced, the current change rate is reduced, the power supply network current is more stable, and further, the instantaneous power supply noise is reduced, and the overall power supply integrity of the chip is optimized. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 is a layout structure diagram of a channel array composed of multiple data channels of a die;
[0023] Figure 2 is a structural diagram of a data channel clock tree circuit according to an illustrative embodiment;
[0024] Figure 3 is a structural diagram of a specific application scenario of a data channel clock tree circuit according to an illustrative embodiment;
[0025] Figure 4 is a structural diagram of an application scenario of a data channel clock tree circuit in the case of a certain number of data channels and delay buffers.
[0026] In the drawings, the names of the components represented by the respective reference numerals are as follows:
[0027] 100, channel array, L1, first data channel, L2, second data channel, L8, eighth data channel, L9, ninth data channel, L16, sixteenth data channel, L17, seventeenth data channel, L24, twenty-fourth data channel, L25, twenty-fifth data channel, L32, thirty-second data channel, L33, thirty-third data channel, L40, fortieth data channel, L41, forty-first data channel, L48, forty-eighth data channel, L49, forty-ninth data channel, L56, fifty-sixth data channel, L57, fifty-seventh data channel, L64, sixty-fourth data channel, LN / 2, N / 2th data channel, LN / 2+1, N / 2+1th data channel, LN, Nth data channel, LG11, first_1 channel group, LG12, first_2 channel group, LG13, first_3 channel group, LG14, first_4 channel group, LG1M, first_M channel group, LG21, second_1 channel group, LG22, second_2 channel group, LG23, second_3 channel group, LG24, second_4 channel group, LG2M, second_M channel group, 200, clock source, 201, first delay buffer group, Buf11, first_1 delay buffer, Buf12, first_2 delay buffer, Buf13, first_3 delay buffer, Buf14, first_4 delay buffer, Buf1M, first_M delay buffer, Ck11, first_1 clock signal line, Ck12, first_2 clock signal line, Ck13, first_3 clock signal line, Ck14, first_4 clock signal line, Ck1M, first_M clock signal line, 202, second delay buffer group, Buf21, second_1 delay buffer, Buf22, second_2 delay buffer, Buf23, second_3 delay buffer, Buf24, second_4 delay buffer, Buf2M, second_M delay buffer, Ck21, second_1 clock signal line, Ck22, second_2 clock signal line, Ck23, second_3 clock signal line, Ck24, second_4 clock signal line, Ck2M, second_M clock signal line, 203, multiple data channels. DETAILED DESCRIPTION
[0028] In order to make the purposes, technical solutions and advantages of the present disclosure clearer, further detailed description will be made to the present disclosure with reference to the drawings and examples.
[0029] It should be noted that the terms "first", "second" and the like in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence.
[0030] As used in the specification and claims of this disclosure, “coupled” or “connected” can mean either a direct connection between the first device and the second device or an indirect connection through one or more intervening devices or connection means.
[0031] Chiplet technology is a technology that combines multiple chiplets (small chips) together to form a large chip architecture, aiming to improve chip performance, reduce cost, and shorten development cycle. By combining different chiplets together, chiplet technology can achieve similar functions to large chips, but also has higher flexibility and scalability.
[0032] The characteristics of chiplet technology mainly include the following aspects.
[0033] First, chiplet technology helps to reduce complexity. Breaking down complex circuits into a series of individual modules that can easily work together makes the design, manufacture, and measurement of circuits easier to implement, thereby reducing the complexity of chips to a lower level.
[0034] Second, chiplet technology has high scalability. Chiplet technology supports the integration of multiple functional chips (chiplets), and the addition of chiplets with related functions according to application requirements, which helps to accelerate system expansion and significantly improve system performance.
[0035] In addition, chiplet technology also has scalability. It can package multiple devices (chiplets) together, helping to simplify the design process.
[0036] Finally, chiplet technology helps to reduce costs. It can reuse the same chiplet hardware in the same large chip and / or different large chips, thereby helping to reduce design complexity. In addition, different chiplets can be designed and produced by different design and production teams, so chiplet technology naturally has the advantage of parallel design, manufacturing, and testing, which can also help to shorten development time and reduce production costs.
[0037] Chiplet technology includes a variety of different packaging forms and interface standards.
[0038] First, in terms of packaging form, chiplets can adopt various packaging methods, such as 2.5D packaging, 3D packaging, etc. Among them, 2.5D packaging refers to placing multiple chiplets on a packaging substrate and interconnecting them through wiring on the substrate. 3D packaging refers to vertically stacking multiple chiplets together and interconnecting them through through-silicon vias or other means.
[0039] Secondly, in terms of interface standards, chiplets can use different interface standards for interconnection, such as UCIe (Universal Chiplet Interconnect Express), PCIe (Peripheral Component Interconnect Express), XSR (Extra Short Reach), etc. In addition, chiplets can also use self-defined interface standards for interconnection, such as AMD's Infinity Fabric, etc.
[0040] In addition, in terms of implementation, chiplets can use different implementations for interconnection, such as IP (Intellectual Property) core-based multiplexing, FPGA-based multiplexing, etc. IP core-based multiplexing refers to abstracting the functional modules in multiple chiplets into IP cores, and realizing different functions through the combination of IP cores.
[0041] Finally, in terms of application fields, chiplets can be applied in various fields, such as computers, communications, medical treatment, etc. In the computer field, chiplets can be used to build high-performance computer clusters, servers, etc.; in the communication field, chiplets can be used to build high-speed communication networks, data centers, etc.; in the medical field, chiplets can be used to build high-precision medical devices, intelligent medical systems, etc.
[0042] In summary, chiplet technology has many different characteristics and application fields, and can meet the needs of different fields.
[0043] Currently, chiplets mainly use standard packaging or 2.5D / 3D advanced packaging. With the continuous advancement of manufacturing and packaging technology, the demand for chiplet bandwidth is also increasing, but the power consumption of chiplets, including the occupied chip width, is expected to be smaller and smaller, which leads to an increase in bandwidth in the limited area of the chip, and in turn causes the power density per unit area of the chiplet to be much higher than that of other serial interfaces (such as Ethernet, PCIe, etc.), and the power density per unit area of the chiplet is also significantly higher than that of other parallel interfaces (such as DDR (Double Data Rate), GDDR (Graphics Double Data Rate), LPDDR (Low-Power Double Data Rate), HBM (High Bandwidth Memory), etc.). Therefore, in application scenarios that require high bandwidth, due to the excessive power density, the power supply noise is too large, which leads to a large number of unstable problems of chip systems.
[0044] Power supply noise is usually caused by the instantaneous current value I multiplied by the equivalent resistance of the power supply network. Therefore, if it is necessary to reduce the power supply noise, it is necessary to reduce the instantaneous current value I or optimize the design of the power plane to reduce the impedance. However, in the core particle, due to the size limitation of the core particle, it is difficult to optimize the power plane of the core particle, but reducing the current means reducing the speed and sacrificing the bandwidth, which is not desirable.
[0045] The power supply noise has different frequency components, and the power supply noise in the working state has the most obvious influence on stability and performance, has the greatest influence on the bit error rate, and has the greatest influence on the stability of the entire link.
[0046] In view of the above situation of the core particle, the embodiment of the present disclosure provides a data lane clock tree circuit, a UCIe core particle and a chip. When the core particle is normally working, the phase of the control clock signal is changed, and the delay of the clock signal is controlled to help reduce the current change rate, thereby optimizing the overall power integrity (PI) of the chip.
[0047] The communication between the core particles is realized based on a plurality of data lanes established between the core particles. Figure 1 is a layout structure diagram of a lane array composed of a plurality of data lanes of a core particle. As shown in Figure 1 , the lane array 100 includes a plurality of data lanes, such as Figure 1 the first data lane L1, the second data lane L2, …, and the Nth data lane LN shown in , wherein the number of N is not limited. According to the latest UCIe protocol standard, the number of lanes of the core particle can be 64. In one application scenario, it can be necessary to simultaneously perform signal potential toggle (rise or fall) on the first data lane L1 to the Nth data lane LN. Whether the signal potential rises or falls simultaneously will generate a huge instantaneous current, resulting in a large DIDT (di / dt, derivative of current with respect to time, current change rate), causing a huge instantaneous power supply noise, which causes a huge impact on the power integrity.
[0048] Figure 2 is a structure diagram of a data lane clock tree circuit according to an illustrative embodiment, as shown in Figure 2 , in the illustrative embodiment, the data lane clock tree circuit of the embodiment of the present disclosure includes a clock source 200 and a delay buffer group. The delay buffer group is two, such as Figure 2 the first delay buffer group 201 and the second delay buffer group 202 shown in , both of which are coupled to the clock source 200 to generate at least two delayed clock signals, such as Figure 2The first delay buffer group 201 is coupled to the clock source 200 and the second delay buffer group 202 is also coupled to the clock source 200, the first delay buffer group 201 generates at least two delay clock signals, and the second delay buffer group 202 generates at least two delay clock signals. The two delay buffer groups are respectively coupled to different data channels in the plurality of data channels 203 through different clock signal lines, respectively, to send the at least two delay clock signals generated by the two delay buffer groups to the data channels coupled to the clock signal lines, for example Figure 2 In the figure, the first delay buffer group 201 is coupled to a part of the data channels on the left side in the plurality of data channels 203 through different clock signal lines, to send the at least two delay clock signals generated by the first delay buffer group 201 to the data channels coupled to the clock signal lines, respectively, and the second delay buffer group 202 is coupled to a part of the data channels on the right side in the plurality of data channels 203 through different clock signal lines, to send the at least two delay clock signals generated by the second delay buffer group 202 to the data channels coupled to the clock signal lines, respectively.
[0049] The data channel clock tree circuit of the embodiment of the present disclosure realizes that different delay clock signals are sent to different data channels, so that the clock signals between the plurality of data channels are not strictly synchronized and there is a preset phase difference due to the delay of the delay buffer group. Because the timing of the inversion of the signal potential of the data channel is driven and controlled by the transition edge of the clock signal, when the data channel clock tree circuit of the embodiment of the present disclosure is used, the clock signals between the plurality of data channels are not strictly synchronized and there is a preset delay, so that when the signal potential of the plurality of data channels is inverted, because the clock signals between the plurality of data channels are not strictly synchronized and there is a preset delay, the inversion time of the signal potential of the plurality of data channels has a preset delay difference, and further, only a part of the data channels performs signal potential inversion at any signal potential inversion time, the execution of the signal potential inversion between the data channels is staggered due to the delay clock signal, compared with the simultaneous execution of the signal potential inversion of all the data channels at the same time, the instantaneous change of the power supply network current is reduced, the DIDT is reduced, the power supply network current is more stable, and further, it is helpful to reduce the instantaneous power supply noise and optimize the overall power integrity of the chip.
[0050] Figure 3 It is a structure schematic diagram of one specific application scenario of the data channel clock tree circuit according to an illustrative embodiment. As shown in the figure, Figure 3 and in combination with Figure 2As shown, in the illustrative embodiment, each delay buffer group includes at least two delay buffers connected in series, wherein the clock signal output terminal of any one of the at least two delay buffers is coupled to at least one of the plurality of data lanes 203 via a clock signal line associated with the any one of the at least two delay buffers, and the clock signal outputted by the clock source 200 generates at least two delayed clock signals in sequence via the at least two delay buffers connected in series, and the at least two delayed clock signals are provided to at least one of the plurality of data lanes 203, respectively.
[0051] For example Figure 3 As shown, the first delay buffer group 201 includes a first_1 delay buffer Buf11, a first_2 delay buffer Buf12, a first_M delay buffer Buf1M connected in series, and a total of M delay buffers, the clock signal input terminal of the first_1 delay buffer Buf11 is coupled to the clock source 200, the clock signal input terminal of the first_2 delay buffer Buf12 is coupled to the clock signal output terminal of the first_1 delay buffer Buf11, and so on. The clock signal output terminal of the first_1 delay buffer Buf11 is coupled to at least one of the plurality of data lanes (i.e., the lane array 100) via a first_1 clock signal line Ck11 associated with the first_1 delay buffer Buf11, the clock signal output terminal of the first_2 delay buffer Buf12 is coupled to at least one of the plurality of data lanes (i.e., the lane array 100) via a first_2 clock signal line Ck12 associated with the first_2 delay buffer Buf12, and so on. The clock signal output terminal of the first_M delay buffer Buf1M is coupled to at least one of the plurality of data lanes (i.e., the lane array 100) via a first_M clock signal line Ck1M associated with the first_M delay buffer Buf1M. The clock signal outputted by the clock source 200 generates M delayed clock signals in sequence via the first_1 delay buffer Buf11, the first_2 delay buffer Buf12, the first_M delay buffer Buf1M connected in series, and the M delayed clock signals are provided to at least one of the plurality of data lanes (i.e., the lane array 100), respectively.
[0052] For example Figure 3As shown, the second delay buffer group 202 includes a second-first delay buffer Buf21, a second-second delay buffer Buf22, ... a second-M delay buffer Buf2M connected in series, for a total of M delay buffers. The clock signal input terminal of the second-first delay buffer Buf21 is coupled to the clock source 200, the clock signal input terminal of the second-second delay buffer Buf22 is coupled to the clock signal output terminal of the second-first delay buffer Buf21, and so on. The clock signal output of the 2_1 delay buffer Buf21 is coupled to at least one of the multiple data channels (i.e., channel array 100) through the 2_1 clock signal line Ck21 associated with the 2_1 delay buffer Buf21. The clock signal output of the 2_2 delay buffer Buf22 is coupled to at least one of the multiple data channels (i.e., channel array 100) through the 2_2 clock signal line Ck22 associated with the 2_2 delay buffer Buf22. ... The clock signal output of the 2_M delay buffer Buf2M is coupled to at least one of the multiple data channels (i.e., channel array 100) through the 2_M clock signal line Ck2M associated with the 2_M delay buffer Buf2M. The clock signal emitted by the clock source 200 is sequentially passed through the second-1st delay buffer Buf21, the second-2nd delay buffer Buf22... the second-M delay buffer Buf2M to generate M delayed clock signals. The M delayed clock signals are respectively provided to at least one of the multiple data channels (i.e., the channel array 100).
[0053] In the illustrative embodiment, the individual delay buffers in the two delay buffer groups have the same driving capability. For example... Figure 3 As shown, the driving capabilities of delay buffers 1_1 to 1_M, delay buffers 1_M and 2_1 to 2_M are the same, so the delay steps of the generated delay clock signals are all equal, which helps to uniformize the inversion of signal potential between different data channels in the time domain, and thus helps to smooth power supply noise.
[0054] In the illustrative embodiment, multiple data channels 203 are arranged in parallel to form a channel array 100. A clock source 200 is positioned at the center line of the channel array 100. Two delay buffer groups extend from the clock source 200 towards the data channels on both sides of the channel array 100. Figure 3 As shown in the example, there are N data channels. For ease of explanation, N is a complex number. In an example based on the UCIe protocol, N could be, for example, 64. Figure 3As shown, N data channels are arranged side by side from left to right to form the channel array 100, the leftmost one is the first data channel L1, and the rightmost one is the Nth data channel LN. The clock source 200 is arranged at the middle line of the channel array 100, i.e. the clock source 200 is arranged at the middle line between the N / 2th data channel LN / 2 and the N / 2+1th data channel LN / 2. The first delay buffer group 201 and the second delay buffer group 202 are respectively extended to the first data channel L1 and the Nth data channel LN on both sides of the channel array 100 with the clock source 200 as the center.
[0055] In the illustrative embodiment, each delay buffer group includes at least two delay buffers connected in series, and the number of delay buffers included in the two delay buffer groups is the same, and the delay buffers included in the two delay buffer groups are symmetrically arranged at the middle line of the channel array 100. For example Figure 3 In the illustrative embodiment, the first delay buffer group 201 and the second delay buffer group 202 each include the same number of M delay buffers connected in series, and the delay buffers included in the first delay buffer group 201 and the second delay buffer group 202 are symmetrically arranged at the middle line of the channel array 100. Specifically, the first_1 delay buffer Buf11 and the second_1 delay buffer Buf21 are symmetrically arranged at the middle line of the channel array 100, the first_2 delay buffer Buf12 and the second_2 delay buffer Buf22 are symmetrically arranged at the middle line of the channel array 100, and so on, and the first_M delay buffer Buf1M and the second_M delay buffer Buf2M are symmetrically arranged at the middle line of the channel array 100. This symmetrical arrangement is conducive to the symmetry of the layout and physical wiring, and also conducive to the consistency control of the clock signal delay between the first delay buffer group 201 and the second delay buffer group 202.
[0056] In the illustrative embodiment, the plurality of data channels are divided into a plurality of channel groups, and the number of data channels included in each channel group is equal. Each delay buffer in the two delay buffer groups is respectively associated with each channel group, and the associated delay buffer and channel group are arranged adjacent to each other, and the clock signal output end of the delay buffer is coupled to each data channel in the channel group through the clock signal line associated with the delay buffer to provide a delayed clock signal to each data channel in the channel group.
[0057] For example Figure 3In the shown embodiment, the plurality of data channels are divided into 2M channel groups (including the 1_1 channel group LG11, the 1_2 channel group LG12, …, the 1_M channel group LG1M, and the 2_1 channel group LG21, the 2_2 channel group LG22, …, the 2_M channel group LG2M), the number of data channels included in each channel group is equal, then each channel group includes N / 2M data channels, taking N=64 and M=4 as an example, then each channel group includes 8 data channels. The 1_1 delay buffer Buf11 is associated with and disposed adjacent to the 1_1 channel group LG11, between the 1_1 delay buffer Buf11 and the 1_1 channel group LG11, the clock signal output end of the 1_1 delay buffer Buf11 is coupled to each data channel in the 1_1 channel group LG11 through the 1_1 clock signal line Ck11 associated with the 1_1 delay buffer Buf11, to provide a delayed clock signal to each data channel in the 1_1 channel group LG11; the 1_2 delay buffer Buf12 is associated with and disposed adjacent to the 1_2 channel group LG12, between the 1_2 delay buffer Buf12 and the 1_2 channel group LG12, the clock signal output end of the 1_2 delay buffer Buf12 is coupled to each data channel in the 1_2 channel group LG12 through the 1_2 clock signal line Ck12 associated with the 1_2 delay buffer Buf12, to provide a delayed clock signal to each data channel in the 1_2 channel group LG12, …, the 1_M delay buffer Buf1M is associated with and disposed adjacent to the 1_M channel group LG1M, between the 1_M delay buffer Buf1M and the 1_M channel group LG1M, the clock signal output end of the 1_M delay buffer Buf1M is coupled to each data channel in the 1_M channel group LG1M through the 1_M clock signal line Ck1M associated with the 1_M delay buffer Buf1M, to provide a delayed clock signal to each data channel in the 1_M channel group LG1M; the 2_1 delay buffer Buf21 is associated with and disposed adjacent to the 2_1 channel group LG21, between the 2_1 delay buffer Buf21 and the 2_1 channel group LG21, the clock signal output end of the 2_1 delay buffer Buf21 is coupled to each data channel in the 2_1 channel group LG21 through the 2_1 clock signal line Ck21 associated with the 2_1 delay buffer Buf21, to provide a delayed clock signal to each data channel in the 2_1 channel group LG21;The 2_2 delay buffer Buf22 is associated with and adjacently arranged with the 2_2 channel group LG22, between the 2_2 delay buffer Buf22 and the 2_2 channel group LG22, the clock signal output end of the 2_2 delay buffer Buf22 is coupled to each data channel in the 2_2 channel group LG22 through the 2_2 clock signal line Ck22 associated with the 2_2 delay buffer Buf22, to provide a delayed clock signal to each data channel in the 2_2 channel group LG22… The 2_M delay buffer Buf2M is associated with and adjacently arranged with the 2_M channel group LG2M, between the 2_M delay buffer Buf2M and the 2_M channel group LG2M, the clock signal output end of the 2_M delay buffer Buf2M is coupled to each data channel in the 2_M channel group LG2M through the 2_M clock signal line Ck2M associated with the 2_M delay buffer Buf2M, to provide a delayed clock signal to each data channel in the 2_M channel group LG2M. The total number of delay buffers in the 1st delay buffer group 201 and the 2nd delay buffer group 202 and the number of channel groups are both 2M.
[0058] Figure 4 The structural diagram of the application scenario of the data channel clock tree circuit in the case of a certain number of data channels and delay buffers. The application scenario takes N=64 and M=4 as an example, the data channels are 64, that is, the channel array contains 64 data channels, the 1st delay buffer group 201 includes the 1_1 delay buffer Buf11, the 1_2 delay buffer Buf12, the 1_3 delay buffer Buf13 and the 1_4 delay buffer Buf14, the 2nd delay buffer group 202 includes the 2_1 delay buffer Buf21, the 2_2 delay buffer Buf22, the 2_3 delay buffer Buf23 and the 2_4 delay buffer Buf24, the number of channel groups is the same as the number of delay buffers, because the 1st delay buffer group 201 and the 2nd delay buffer group 202 each have M=4 delay buffers, a total of 8 delay buffers, so the number of channel groups is 8, therefore the 64 data channels are divided into 8 channel groups, each channel group contains 8 data channels. Figure 4 In the application scenario, based on Figure 3The data channel numbers shown in the order of increasing from left to right are associated with the first 1_4 channel group LG14 containing the first data channel L1 to the eighth data channel L8 associated with the first 1_4 delay buffer Buf14, the first 1_3 channel group LG13 containing the ninth data channel L9 to the sixteenth data channel L16 associated with the first 1_3 delay buffer Buf13, the first 1_2 channel group LG12 containing the seventeenth data channel L17 to the twenty-fourth data channel L24 associated with the first 1_2 delay buffer Buf12, the first 1_1 channel group LG11 containing the twenty-fifth data channel L25 to the thirty-second data channel L32 associated with the first 1_1 delay buffer Buf11, the second 2_1 channel group LG21 containing the thirty-third data channel L33 to the fortieth data channel L40 associated with the second 2_1 delay buffer Buf21, the second 2_2 channel group LG22 containing the forty-first data channel L41 to the forty-eighth data channel L48 associated with the second 2_2 delay buffer Buf22, the second 2_3 channel group LG23 containing the forty-ninth data channel L49 to the fifty-sixth data channel L56 associated with the second 2_3 delay buffer Buf23, and the second 2_4 channel group LG24 containing the fifty-seventh data channel L57 to the sixty-fourth data channel L64 associated with the second 2_4 delay buffer Buf24.The clock signal output end of the 1st 1 delay buffer Bufl l is coupled to the 25th data channel L25 to the 32nd data channel L32 included in the 1st 1 lane group LGl l through the 1st 1 clock signal line Ckl l, the clock signal output end of the 1st 2 delay buffer Bufl 2 is coupled to the 17th data channel L17 to the 24th data channel L24 included in the 1st 2 lane group LGl 2 through the 1st 2 clock signal line Ck 12, the clock signal output end of the 1st 3 delay buffer Bufl 3 is coupled to the 9th data channel L9 to the 16th data channel L16 included in the 1st 3 lane group LGl 3 through the 1st 3 clock signal line Ckl 3, the clock signal output end of the 1st 4 delay buffer Bufl 4 is coupled to the 1st data channel L1 to the 8th data channel L8 included in the 1st 4 lane group LGl 4 through the 1st 4 clock signal line Ckl 4, the clock signal output end of the 2nd 1 delay buffer Buf21 is coupled to the 33rd data channel L33 to the 40th data channel L40 included in the 2nd 1 lane group LG21 through the 2nd 1 clock signal line Ck21, the clock signal output end of the 2nd 2 delay buffer Buf22 is coupled to the 41st data channel L41 to the 48th data channel L48 included in the 2nd 2 lane group LG22 through the 2nd 2 clock signal line Ck22, the clock signal output end of the 2nd 3 delay buffer Buf23 is coupled to the 49th data channel L49 to the 56th data channel L56 included in the 2nd 3 lane group LG23 through the 2nd 3 clock signal line Ck23, and the clock signal output end of the 2nd 4 delay buffer Buf24 is coupled to the 57th data channel L57 to the 64th data channel L64 included in the 2nd 4 lane group LG24 through the 2nd 4 clock signal line Ck24.
[0059] Thus, different channel groups obtain different delay clock signals through different associated delay buffers, and the delay is gradually increased in the delay buffer delay step from the central region of the channel array 100 to the two side regions, so in the actual application scenario, the phase of the clock signal of the channel group in the central region of the channel array 100 is relatively in advance, and from the data channel in the central region of the channel array 100 to the two side regions of the channel array 100, the phase of the clock signal of different channel groups is sequentially delayed, in a hypothetical case, if all data signals of the data channel need to be inverted from low to high at the same time, the data signal of the channel group in the central region of the channel array 100 will be inverted from low to high in advance due to the effect of the data channel clock tree circuit implemented by the present disclosure, and from the data channel in the central region of the channel array 100 to the two side regions of the channel array 100, the data signal of different channel groups will be sequentially inverted from low to high in the delay buffer delay step, which avoids the instantaneous impact on the power supply network when multiple data channels 203 are inverted from low to high at the same time, reduces the instantaneous change of the power supply network current, reduces the DIDT, makes the power supply network current more stable, helps to reduce the instantaneous power noise, and helps to optimize the overall power integrity of the chip.
[0060] In the illustrative embodiment, from the channel group close to the center line position of the channel array 100 to the channel group far from the center line position of the channel array 100, the delay clock signal obtained is symmetrically delayed by a set delay step from the center line position of the channel array 100, and the set delay step is the delay of a delay buffer to the clock signal.
[0061] For example Figure 3 In the figure, the channel group close to the center line position of the channel array 100 is the first 1 channel group LG11 and the second 1 channel group LG21, and the channel group farthest from the center line position of the channel array 100 is the first M channel group LG1M and the second M channel group LG2M. From the first 1 channel group LG11 and the second 1 channel group LG21 to the first M channel group LG1M and the second M channel group LG2M, the delay clock signal obtained by each channel group is symmetrically delayed by a set delay step from the center line position of the channel array 100, and the set delay step is the delay of a delay buffer to the clock signal.
[0062] The data channel clock tree circuit of the embodiment of the present disclosure can be applied to related chiplet protocols of multiple data channels, based on commonly used UCIe protocols, PCIe protocols, and XSR protocols of chiplets, in the illustrative embodiment, the data channel is at least one of a UCIe data channel, a PCIe data channel, and an XSR data channel.
[0063] In the illustrative embodiments, the implementation of at least one of the clock source 200, the delay buffer set, and the plurality of data lanes can be a combination of hardware, firmware, and software (i.e., a program) according to different designs.
[0064] In hardware form, at least one of the clock source 200, the delay buffer set, and the plurality of data lanes can be implemented as logic circuits on an integrated circuit, e.g., various logic blocks, modules, and circuits of one or more hardware controllers, microcontrollers, hardware processors, microprocessors, application-specific integrated circuits (ASICs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), central processing units (CPUs), or other processing units that perform the functions of at least one of the clock source 200, the delay buffer set, and the plurality of data lanes. The functions of at least one of the clock source 200, the delay buffer set, and the plurality of data lanes can be implemented in hardware circuitry, e.g., various logic blocks, modules, and circuits of an integrated circuit, using hardware description languages (e.g., Verilog HDL or VHDL) or other suitable programming languages.
[0065] In software or firmware form, the functions of at least one of the clock source 200, the delay buffer bank, and the plurality of data lanes can be implemented as programming codes. For example, the clock source 200, the delay buffer bank, and the plurality of data lanes can be implemented using a general programming language (e.g., C, C++, or assembly language) or other suitable programming language. The programming codes can be recorded, stored in a non-transitory machine-readable storage medium. In some embodiments, the non-transitory machine-readable storage medium includes, for example, a semiconductor memory and / or a storage device. An electronic device (e.g., a CPU, a hardware controller, a microcontroller, a hardware processor, or a microprocessor) can read and execute the programming codes from the non-transitory machine-readable storage medium, thereby implementing the functions of at least one of the clock source 200, the delay buffer bank, and the plurality of data lanes.
[0066] In illustrative embodiments, the data lane clock tree circuit of the present disclosure is applicable to a SoC chip, etc., where the SoC chip can be any one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a TPU (Tensor Processing Unit), a NPU (Neural network Processing Unit), a DPU (Deep learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit).
[0067] In illustrative embodiments, a UCIe die is also provided, including the data lane clock tree circuit of any one of the above embodiments.
[0068] In illustrative embodiments, a chip is also provided, including the data lane clock tree circuit of any one of the above embodiments.
[0069] The above only describes the preferred embodiments of the present disclosure and is not intended to limit the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the scope of protection of the present disclosure.
Claims
1. A data lane clock tree circuit, comprising: include: Clock source; The delay buffer group comprises two delay buffer groups, each of which is coupled to the clock source to generate at least two delayed clock signals. Furthermore, each of the two delay buffer groups is coupled to different data channels among multiple data channels via different clock signal lines, so as to send the at least two delayed clock signals generated therefrom to the data channels coupled to each of the clock signal lines.
2. The data lane clock tree circuit of claim 1, wherein, Each of the aforementioned delay buffer groups includes: At least two delay buffers are connected in series, wherein the clock signal output terminal of any one of the at least two delay buffers is coupled to at least one of the plurality of data channels through a clock signal line associated with the at least one delay buffer, and the clock signal emitted by the clock source is sequentially transmitted through the at least two delay buffers connected in series to generate the at least two delayed clock signals, and the at least two delayed clock signals are respectively provided to at least one of the plurality of data channels.
3. The data channel clock tree circuit according to claim 2, characterized in that: Each of the delay buffers in the two delay buffer groups has the same driving capability.
4. The data channel clock tree circuit according to claim 1, characterized in that: The multiple data channels are arranged in parallel to form a channel array; The clock source is positioned at the center line of the channel array; The two delay buffer groups extend from the clock source toward the data channels on both sides of the channel array.
5. The data channel clock tree circuit according to claim 4, characterized in that: Each of the delay buffer groups includes at least two delay buffers connected in series, and the number of delay buffers in both delay buffer groups is the same, with the delay buffers in both delay buffer groups arranged symmetrically about the centerline of the channel array.
6. The data channel clock tree circuit according to claim 5, characterized in that: The multiple data channels are divided into multiple channel groups, and each channel group contains an equal number of data channels; Each of the two delay buffer groups is associated with each of the channel groups, the associated delay buffers and the channel groups are arranged adjacent to each other, and the clock signal output of the associated delay buffer is coupled to each data channel in the channel group through a clock signal line associated with the delay buffer to provide a delayed clock signal to each data channel in the channel group.
7. The data channel clock tree circuit according to claim 6, characterized in that: In the plurality of channel groups, the delayed clock signal obtained from the channel group near the center line of the channel array to the channel group away from the center line of the channel array is delayed symmetrically with respect to the center line of the channel array by a set delay step, wherein the set delay step is a delay of the clock signal by one of the delay buffers.
8. The data lane clock tree circuit of claim 1, wherein: the data lane is at least one of a UCIe data lane, a PCIe data lane, an XSR data lane.
9. A UCie core particle, characterized in that, a data lane clock tree circuit as claimed in any one of claims 1 to 7.
10. A chip, characterized by a data lane clock tree circuit as claimed in any one of claims 1 to 7.