Energy recovery adiabatic flip-flop and resonator-based Petint clock generator

Through adiabatic reversible calculation and Bennett clock technology, SCRL logic and slow tilt change clocking are used to solve the heat limitation caused by energy dissipation in traditional CMOS circuits, achieving more efficient energy utilization and performance improvement.

CN120380702APending Publication Date: 2025-07-25INDIANA INTEGRATED CIRCUITS LLC +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380048428.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-06-21
Filing Date
2023-06-21
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The performance growth of modern microprocessors is limited by heat generation. Traditional CMOS circuits consume energy during switching lead to heat limitations, limiting speed improvements.

Method used

Using adiabatic reversible computing technology, using the orbital charge recovery logic (SCRL) and Bennett clock, the slow tilt change clock is used as a voltage source to reduce energy dissipation and avoid energy loss through the energy recovery step.

Benefits of technology

The power consumption and heat generation in the computing device are significantly reduced, and more efficient energy utilization is achieved, which promotes the speed of the computing device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120380702A_ABST
    Figure CN120380702A_ABST
Patent Text Reader

Abstract

A method includes, during a time period A, in a computer storage element having first and second power inputs separated by an array of transistors, the computer storage element configured to store computer data bits, move an input of the array of transistors, a logic value "1" or "0" to a master latch; during a time period B, recovering at least a portion of the energy containing the input using the first clock; during a time period C, moving the input in the main latch to the isolation stage; recovering at least a portion of the energy containing the input from the isolation stage using a second clock during a time period D; during a time period E, recovering at least a portion of the energy containing the input from the slave latch using a third clock; and moving at least a portion of the input from the isolation stage to the slave latch during the time period F.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 354,109, filed on June 21, 2022, the disclosure of which is incorporated herein by reference.

[0003] This disclosure relates to energy - recycling adiabatic flip - flops and resonator - based Bennett clock generators. Background Art

[0004] Despite exponential progress over the past five decades, the growth in computing device performance is coming to an end. Modern microprocessors are limited by heat generation and have been stuck at speeds of 4 GHz since 2004. Traditional Complementary MOSFET logic (CMOS) circuits consume energy in the form of heat every time they switch.

[0005] Accordingly, those of ordinary skill in the art to which this invention pertains continue to research and develop in the area of energy recycling for computing devices. Summary of the Invention

[0006] Generally, in some non - limiting embodiments or examples, a computer storage element and a method of using the same are provided. In some non - limiting embodiments, the computer storage element can be a flip - flop or a memory. In some non - limiting embodiments, compared with existing usage methods, the computer storage element can be operated in a way that reduces, minimizes, or avoids power from entering the computer storage element and thereby reduces, minimizes, or avoids power consumption and the resulting heat in the computer storage element.

[0007] In one example, the disclosed method includes: during a time period A, in a computer storage element having first and second power inputs separated by a transistor array of the computer storage element, the computer storage element being configured to store a computer data bit, moving an input, a logical value of “1” or “0”, of the transistor array to a master latch; during a time period B after the time period A, recovering at least a portion of the energy containing the input using a first clock; during a time period C after the time period B, moving the input in the master latch to an isolation stage; during a time period D after the time period C, recovering at least a portion of the energy containing the input from the isolation stage using a second clock; during a time period E after the time period D, recovering at least a portion of the energy containing the input from a slave latch using a third clock; and during a time period F after the time period E, moving at least a portion of the input from the isolation stage to the slave latch. Brief Description of the Drawings

[0008] The present disclosure will be described with reference to the following drawings, in which like reference numerals always represent like components.

[0009] Figure 1A is a schematic diagram of a conventional CMOS inverter;

[0010] Figure 1B is Figure 1A the timing diagram of the CMOS inverter of

[0011] Figure 2A is a schematic diagram of an adiabatic inverter implemented using split-rail charge recovery logic (SCRL);

[0012] Figure 2B is Figure 2A the timing diagram of the adiabatic inverter of

[0013] Figure 3A is a schematic diagram of three SCRL inverters connected in series;

[0014] Figure 3B is Figure 3A the timing diagram of the Bennett clock of the three SCRL inverters of

[0015] FIG. 4A is a schematic diagram of an adiabatic SCRL circuit connected in series with timing elements;

[0016] Figure 4B is the timing diagram of the timing elements of FIG. 4A;

[0017] Figure 5A is a schematic diagram of an adiabatic master-slave flip-flop according to the principles of the present invention;

[0018] Figure 5B is Figure 5A the timing diagram of the adiabatic master-slave flip-flop according to the principles of the present invention of

[0019] Figure 6A is a schematic diagram of the generation of a pair of Bennett clocks; and

[0020] Figure 6B is Figure 6B the figure of DETAILED DESCRIPTION

[0021] Disclosed herein are methods and designs for implementing adiabatic computing flip-flops and SRAM cells using SCRL logic. The present disclosure makes use of the disclosure of U.S. Patent No. 11,488,660, the entire content of which is incorporated herein by reference.

[0022] Despite exponential progress over the past five decades, the growth in the performance of computing devices is coming to an end. Modern microprocessors are limited by heat generation, and their speed has been capped at 4 GHz since 2004. Traditional complementary MOSFET logic (CMOS) circuits consume energy in the form of heat every time they switch. CMOS devices use sharp transitions for switching and result in the following energy dissipation equation:

[0023]

[0024] Here, C is the load capacitance of the logic gate, and V DD is the power supply. After each switching event, the energy is discarded and dissipated as heat, thus imposing a limit on the operating speed of modern computing devices. A diagram of a traditional CMOS inverter is presented as Figure 1A and its timing diagram is presented as Figure 1B .

[0025] Adiabatic reversible computing is a viable alternative to traditional circuits as it reduces heat generation by avoiding unnecessary dissipation. Adiabatic reversible computing, or simply adiabatic computing, uses reversible logic and quasi-adiabatic transitions to reduce heat generation by introducing a trade-off between speed and power consumption. This can be achieved by using a slow ramp clock as the voltage source, resulting in the following energy dissipation expression:

[0026]

[0027] Here, C is the load capacitance of the logic gate, and V t is the ramp power supply, RC is the intrinsic time constant of the gate, and T is the ramp time of the power supply. The additional terms including the RC time constant of the gate and the ramp time of the power supply can further reduce the dissipated energy. When T is much lower than the RC time constant of the gate, the dissipation can be significantly reduced.

[0028] Adiabatic computing can be implemented as split-rail charge recovery logic (SCRL) [1]. An inverter using adiabatic SCRL logic is shown as Figure 2A where the power supply and ground terminals have been replaced by positive and negative ramp clocks with a ramp period T. As shown in Figure 2B , the two ramp clocks start from the empty state (0 V), then ramp up in time T until they reach the effective logic state, hold the logic value for a period of time, and then ramp down to the empty state to recover energy. The energy recovery step is crucial for avoiding energy dissipation, but the information or data (logic '1' or '0')) is lost between cycles as the inverter returns to the empty state. During the computing phase, the output of the inverter will follow the positive or negative clock to represent logic '1' or logic '0'.

[0029] Since the logical values of the SCRL gates are invalid when the clock is rising and falling or during the empty state, any subsequent gates will need to have clocks with different phases. Figure 3 shows three SCRL inverter chains, which illustrate how phased clocks are used to power different stages of the circuit. This clocking scheme is called Bennett clock [2], and it consists of clocks that rise only when the previous stage has a valid state. Similarly, during energy recovery, the last stage of the logic first falls, and once the empty state is reached, the previous stage is then executed. The Bennett clock ensures that the logical input values of the circuit are valid during the computation and energy recovery steps, and provides a straightforward implementation for adiabatic computing.

[0030] Although adiabatic microprocessors using SCRL circuits have been successfully implemented as described above [3], the timing elements have not been implemented using adiabatic logic. Modern CMOS circuits have both combinational logic (such as the three-inverter chain described above) and sequential logic consisting of memory elements for storing the results obtained from the combinational logic. Figures 4A and 4B illustrate this, showing that multiple combinational logic blocks can use the same set of Bennett clocks, and their intermediate results are stored in the timing elements. The timing elements are buffers composed of flip-flop gates or memory elements composed of SRAM cells. When all the clocks have risen, these timing elements sample the results of the combinational logic during the active cycle and provide the results as inputs to the next set of combinational logic in the next cycle.

[0031] The adiabatic flip-flop and memory designs are the first of their kind and can be used in any practical implementation of adiabatic computing, such as adiabatic microprocessors. An adiabatic microprocessor using the proposed cells will fully implement adiabatic computing.

[0032] The timing requirements for separating the timing elements of the Bennett clock combinational logic are somewhat complex because when all the Bennett clock phases are valid, the data should be latched, but the data should not appear at the latch output until all the Bennett phases have ramped back down. This can be achieved using a master-slave flip-flop with two power clocks and an additional control signal.

[0033] The design of the adiabatic master-follower flip-flop (AMFFF) is as Figure 5A shown. This design uses an energy-recovery master latch as the master stage and another energy-recovery latch as the slave latch, located on the left and right sides respectively, separated by an energy-recovery isolation stage. Although some energy loss is required to destroy the data at the latch, the energy-recovery phase helps to minimize the loss.

[0034] Figure 5BIllustrated are exemplary Bennett clock signals BClk1+ and BClkN+. These Bennett clock signals BClk1+ and BClkN+ and their inverted (negative) clocks BClk1- and BClkN- (not shown) can be used with a number of unillustrated examples of combinational logic as shown in FIG. 4A. In Figure 5A In an instance of combinational logic (not shown) in series before the AMFFF shown in

[0035] Still referring to Figure 5B , the clock signals BClk1+, BClkN+, MClk+, CTRL+, and FClk+, and their respective inverted signals BClk1-, BClkN-, MClk-, CTRL-, and FClk- (not shown for simplicity) can be generated by an external circuit (not shown for simplicity). However, it is contemplated that this external circuit can include logic circuitry, resonators (such as but not limited to piezoelectric resonators), and / or a microprocessor programmed or configured to provide these clock signals and their respective inverted signals.

[0036] Furthermore, in Figure 5B , the clock signals MClk+, CTRL+, and FClk+ and their respective inverted signals MClk-, CTRL-, and FClk- (not shown for simplicity) are examples of Bennett clock signals such as BClk1+ and BClkN+, and their inverted (negative) clocks BClk1- and BClkN-, but are described and designated herein as clock signals MClk+, CTRL+, and FClk+ and their respective inverted signals MClk-, CTRL-, and FClk- for the purposes of describing the present invention.

[0037] Figure 5A Illustrated is a non-limiting example of an AMFFF including 14 transistors. However, this should not be construed in a limiting sense as it is contemplated that in practice, the AMFFF can include more or fewer transistors as would be deemed appropriate and / or desirable by one of ordinary skill in the art of the present invention for a particular application or function. Thus, Figure 5A the AMFFF shown should not be construed as limiting the present invention.

[0038] In Figure 5AIn the illustrated example of AMFFF, transistors M2, M3, M5, M7, M10, M11, and M13 are p-channel MOSFETs; while transistors M1, M4, M6, M8, M9, M12, and M14 are n-channel MOSFETs. Input “In” passes through a transmission gate, which, in the example, includes transistors M1 and M2, and transistors M1 and M2 are controlled by positive and negative master clock signals (MClk+ and MClk-). In one example, transistors M3, M4, M5, and M6 form a master latch that latches data (logic ‘1’ or ‘0’) present at the input IN of the transmission gate into the master latch under the control of master clock signals MClk+ and MClk- during time period A in Figure 5B while the Bennett clock is at VDD (for BClk1+ and BClkN+) and VSS (not shown, for BClk1- and BClkN-). Additionally, the transistors of the master latch perform partial energy recovery on any data retained in the master latch during time period B after time t2 in Figure 5B while the Bennett clocks BClk1+, BClkN+, BClk1-, and BClkN- are in their idle state, e.g., 0 V.

[0039] In the example, each of BClk1+ and BClkN+ switches between an idle logic state, e.g., represented by 0 V (e.g., don't care), and VDD, while each of BClk1- and BClkN- (the inverse of BClk1+ and BClkN+) switches between an idle logic state, e.g., represented by 0 V (e.g., don't care), and VSS (not specifically shown for BClk1- and BClkN-) (simultaneously with BClk1+ and BClkN+).

[0040] In the example, each of the clock signals MClk+ and FClk+ switches between an idle logic state, e.g., represented by 0 V (e.g., don't care), and VDD, while each of the clock signals MClk- and FClk- (the inverse of the clock signals MClk+ and FClk+) switches between an idle logic state, e.g., represented by 0 V (e.g., don't care), and VSS (not specifically shown for the clock signals MClk- and FClk) (simultaneously with the clock signals MClk+ and FClk+).

[0041] Finally, in one example, the control signal CTRL+ switches between VSS and VDD (as Figure 5Bas shown), and the control signal CTRL- (not specifically shown; the inverse of the control signal CTRL+) switches between VDD and VSS (simultaneously with the control signal CTRL+). In an example, VDD can be a positive voltage and VSS can be a negative voltage. However, this should not be construed in a limiting sense, as it is contemplated that each of VDD and VSS can be any suitable and / or desired voltage that facilitates the operation of the AMFFF described herein.

[0042] Once the Bennett clocks BClk1+, BClkN+, BClk1- and BClkN- are in their empty state (e.g., 0 V) and the master clocks MClk+ and MClk- are at the respective VDD and VSS (not shown) at time t1, in response to the control signals CTRL+ and CTRL- being activated during the time period C between times t1 and t2 in Figure 5B the retained data (logical '1' or '0') in the master latch is moved to the slave latch through the transistors M7, M8, M9 and M10 of the isolation stage. The transistors M7, M8, M9 and M10 act as an isolation stage between the master latch and the slave latch.

[0043] Once the Bennett clocks BClk1+, BClkN+, BClk1- and BClkN- have ramped from VDD (for BClk1+ and BClkN+) and VSS (not shown, for BClk1- and BClkN-) to the empty logic state (e.g., don't care), e.g., 0 V, during the time period E in FIG. 5C starting from time t1 the clock signals FClk+ and FClk- ramp from VDD and VSS respectively to the empty logic state (e.g., don't care), e.g., 0 V, thereby partially recovering the energy of any data retained in the slave latch.

[0044] As described above, during the time period C between times t1 and t2, the isolation stage is activated by ramping the control signal CTRL+ from VSS up to VDD and ramping CTRL- from VDD up to VSS, thereby transferring the data (logical '1' or '0') from the master latch to the isolation stage. Starting from time t2, in Figure 5BDuring the time period F in [ ], in response to the ramps of the latch clocks FClk+ and FClk-, from 0 V or floating to the respective VDD (for FClk+) and VSS (for FClk-), the data is loaded into the slave latch (composed of transistors M11, M12, M13, and M14). The data (logical '1' or '0') is held at the output (OUT) of the slave latch and is ready to be used as an input (IN) in the next Bennett cycle of another circuit (not shown) (e.g., the combinational logic in Figure 4A).

[0045] Thereafter, during Figure 5B the time period D in [ ], the control signals CTRL+ and CTRL- ramp from VDD to VSS (for CTRL+) and from VSS to VDD (for CTRL-), thereby recovering energy in the isolation stage and isolating the master latch from the slave latch from each other.

[0046] In Figure 5B the AMFFF timing diagram shown, when all phases of the Bennett clocks BClk1+, BClkN+, BClk1-, and BClkN- ramp up from 0 V to VDD (for BClk1+ and BClkN+) and from 0 V to VSS (not shown, for BClk1- and BClkN-), starting from time t0 in Figure 5B the time period A, the master clocks MClk+ and MClk- of the master latch ramp up from 0 V to VDD (for MClk+) and from 0 V to VSS (not shown, for MClk-), and thus, the data at the input (IN) is latched into the master latch following the ramps of the master clocks MClk+ and MClk- during the time period A starting from time t0.

[0047] Then, the master clocks MClk+ and MClk- remain active at VDD and VSS, respectively, until all of the Bennett phases BClk1+, BClkN+, BClk1-, and BClkN- ramp back from VDD and VSS to 0 V or floating. At this time (marked as time t1), the slave clocks FClk+ and FClk- Figure 5B ramp from the respective VDD and VSS to 0 V or floating during the time period E in [ ], thereby partially recovering the energy of any data retained in the slave latch.

[0048] Thereafter, between time t1 and t2, during Figure 5B the time period C in [ ], the control signals CTRL+ and CTRL- ramp from VSS to VDD (for CTRL+) and from VDD to VSS (for CTRL-), to move the data (logical '1' or '0') from the master latch to the isolation stage.

[0049] Next, during a period F starting at time t2 in Figure 5B , FClk+ and FClk- slew from 0 V or a floating state to the respective VDD and VSS. As a result, the data in the isolation stage is input to the slave latch, and this data is then stored at the output OUT of the slave latch and can be used as an input (IN) to another circuit (not shown) (e.g., another instance of AMFFF) in its next Bennett cycle.

[0050] The current data at the output OUT remains stored in the slave latch until Figure 5B the next cycle of MClk+, FClk+, CTRL+, MClk-, FClk-, and CTRL- as shown in

[0051] Here, for clarity, the in-phase Bennett clocks BClk1 and BClkN, the master clock MClk, the slave clock FClk, and the inversion of the control signal CTRL are omitted from the timing diagram in Figure 5B .

[0052] In this design, energy is recovered from the master latch during a period B in Figure 5B , from the slave latch during a period F in Figure 5B , and from the isolation stage during a period D in Figure 5B . The amount of energy recovered from each of the master latch, slave latch, and isolation stage is limited by the threshold voltages of the transistors forming the master latch, slave latch, and isolation stage.

[0053] The AMFFF described above can be used in any timing element for adiabatic computing, such as buffers and pipelined microprocessors.

[0054] As can be seen from the above, during the periods A to F in Figure 5B the following occurs:

[0055] ● Period A: The data (logical '1' or '0') at the transmission gate In is moved to and stored in the master latch;

[0056] ● Period B: Part of the energy of the data stored in the master latch is recovered by a circuit including MClk+, MClk-, or both;

[0057] ● Period C: The data in the master latch is moved to and stored in the isolation stage;

[0058] ● Time period D: A portion of the energy of the data retained in the isolation stage is recovered by a circuit including CTRL+, CTRL-, or both;

[0059] ● Time period E: A portion of the energy of the data retained from the latch is recovered by a circuit including FClk+, FClk-, or both; and

[0060] ● Time period F: The data in the isolation stage is moved to and stored in the slave latch, i.e., the data is present at the output (Out) of the slave latch.

[0061] Thus, the operation of the AMFFF shown in Figure 5A has been described. Now, a clock circuit for generating a Bennett clock sequence using a resonator (particularly but not limited to a piezoelectric contour mode resonator) will be described. However, the description of the resonator as a piezoelectric contour mode resonator in this article should not be construed in a limiting sense, as it is contemplated that any suitable and / or desired resonator that helps generate the Bennett clock sequence in the manner described below can be used.

[0062] In a Bennett clock, the power clock sequentially powers on the logic at successive levels during the computation phase and then powers off the logic in the reverse manner during the de-computation phase, thereby recovering the energy in the circuit. The timing of a three-level positive Bennett clock is shown in Figure 3. To generate the Bennett clock waveform of Figure 3, multiple resonators and switching transistors are required. A non-limiting example is to use a set of four resonators, each with a 90° phase difference.

[0063] As shown in Figure 6, each Bennett phase is generated by connecting an appropriate generator or resonator to the output clock line to provide a rising edge or a falling edge at the appropriate time. At the end of the first ramp-up (or ramp-down) time period, e.g., 1 / 4 of the period of each generator or resonator, the switching transistor (labeled Ramp in Figure 6) is turned off to disconnect the clock line. Thus, the circuit is driven by the clock signal on the clock line from the generator or resonator. Depending on the application, an optional second Hold transistor can be added to the clock line to help enforce a subsequent Hold time period.

[0064] If there is no ramp transistor to isolate the generator or resonator from the circuit driven by the clock signal on the clock line, the logic of the circuit driven by the generator or resonator is essentially dynamic logic. Then, the bit energy is held in the load capacitor, i.e., the capacitance of the transistors forming the circuit is driven by each generator or resonator for a time limited by the leakage time.

[0065] After the HOLD period, the circuits driven by each generator or resonator are reconnected to the generator or resonator through the Ramp transistors for a second ramp-down (or ramp-up) time period, and thus some irreversible dissipation will occur because the generator or resonator "makes up" for the energy lost due to leakage. However, if the clock line is connected to a static power supply, this energy loss will be equivalent to the energy lost due to passive leakage during the HOLD time.

[0066] During the second ramp-down (or ramp-up) time period, energy returns to the generator or resonator. Logic techniques for reducing leakage will reduce this energy loss.

[0067] If a Hold transistor is added, as Figure 6A shown. This transistor holds the clock line at VDD or VSS and supplies the energy leaked in the circuit driven by the generator or resonator.

[0068] Four generators or resonators with a 90° phase difference are sufficient to implement any number of Bennett phases. These generators or resonators can provide the positive and negative ramps required for SCRL using the method in Figure 6. The number of Bennett phases in an actual adiabatic reversible system may be 16 or less, so for a reasonable clock speed, the time amount of the dynamically held bits will not be too long. Controlling the switching transistors will require some logic, but a simple state machine implemented using conventional CMOS logic and synchronized with the generator or resonator can be used.

[0069] The features of the present disclosure may be the following clauses:

[0070] Clause 1: A method, comprising: during a time period A, in a computer storage element having first and second power inputs separated by a transistor array of computer storage elements, the computer storage element being configured to store computer data bits, moving an input, a logic value "1" or "0", of the transistor array to a master latch; during a time period B after the time period A, recovering at least a portion of the energy including the input using a first clock; during a time period C after the time period B, moving the input in the master latch to an isolation stage; during a time period D after the time period C, recovering at least a portion of the energy including the input from the isolation stage using a second clock; during a time period E after the time period D, recovering at least a portion of the energy including the input from a slave latch using a third clock; and during a time period F after the time period E, moving at least a portion of the input from the isolation stage to the slave latch.

[0071] Clause 2. The method according to Clause 1, further comprising, during the time period A, ramping the first clock from a null value to VDD.

[0072] Clause 3. The method according to any one of Clauses 1 to 2 further includes ramping the first clock from VSS to VDD during time period B.

[0073] Clause 4. The method according to any one of Clauses 1 to 3 further includes ramping the second clock from VSS to VDD during time period C.

[0074] Clause 5. The method according to any one of Clauses 1 to 4 further includes ramping the second clock from VDD to VSS during time period D.

[0075] Clause 6. The method according to any one of Clauses 1 to 5 further includes ramping the third clock from VDD to empty during time period E.

[0076] Clause 7. The method according to any one of Clauses 1 to 6 further includes ramping the third clock from empty to VDD during time period F.

[0077] Although various examples of the disclosed electrochemical devices and conductive substrates have been shown and described, modifications may occur to those of ordinary skill in the art upon reading the specification. This application includes such modifications and is limited only by the scope of the claims.

Claims

1. A method, the method comprising: During a time period A, in the computer storage element having first and second power inputs separated by a transistor array of computer storage elements, the computer storage element is configured to store the computer data bit, and move the input, a logic value "1" or "0", of the transistor array to a master latch; During a time period B after the time period A, recover at least a portion of the energy including the input using a first clock; During a time period C after the time period B, move the input in the master latch to an isolation stage; During a time period D after the time period C, recover at least a portion of the energy including the input from the isolation stage using a second clock; During a time period E after the time period D, recover at least a portion of the energy including the input from a slave latch using a third clock; And During a time period F after the time period E, move at least a portion of the input from the isolation stage to the slave latch.

2. The method according to claim 1, further comprising, during the time period A, ramping the first clock from a null value to VDD.

3. The method according to claim 1, further comprising, during the time period B, ramping the first clock from VSS to VDD.

4. The method according to claim 1, further comprising, during the time period C, ramping the second clock from VSS to VDD.

5. The method according to claim 1, further comprising, during the time period D, ramping the second clock from VDD to VSS.

6. The method according to claim 1, further comprising, during the time period E, ramping the third clock from VDD to a null value.

7. The method according to claim 1, further comprising, during the time period F, ramping the third clock from a null value to VDD.

Citation Information

Patent Citations

  • Adiabatic flip-flop and memory cell design

    US11488660B2