Power consumption control based on random bus inversion
By applying the random bus inversion device (RBI) in the SoC network, random decisions are made independent of the data bit value, and the peak power consumption problem in the SoC is solved, and power consumption optimization and delay control are achieved.
Patent Information
- Application Number
- CN202280092735.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-01
- Filing Date
- 2022-12-21
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-12-21
AI Technical Summary
The peak power consumption problem in existing SoCs, especially under high switching rate conditions, can lead to voltage drops and physical damage, and conventional data bus inversion methods are costly and undesirable in terms of chip area and delay.
By randomly applying bus inversion operations at selected locations in the SoC network, a random bus inversion device (RBI) is used to reduce power consumption peaks, which make random decisions independent of the data bit value to reduce the bit switching rate.
It effectively reduces the peak power consumption, avoids the increase in chip area and communication delay, reduces the probability of exceeding the target switching rate, and protects the voltage stability of the SoC.
Smart Images

Figure CN118805163B_ABST
Abstract
Description
Technical Field
[0001] Embodiments described herein relate generally to system-on-chip (SoC) integrated circuits (ICs), and particularly to methods and systems for limiting power consumption in a SoC by randomly applying bus inversion operations. Background Art
[0002] Various computer systems, such as system-on-chip (SoC), include multiple agent devices that communicate through a fabric.Power consumption in an SoC depends on various factors, such as the SoC structure, supply voltage, and the amount of traffic passing through the fabric.
[0003] Power consumption in an SoC can be reduced, for example, by reducing the power consumed by each link in the fabric. Methods for reducing power consumption on a bus are known in the art. For example, data bus inversion (DBI) is a technique designed to reduce power consumption caused by bit switching between successive transmissions over the bus. In conventional DBI, a data unit is logically inverted when at least half of the bits in the data unit differ from those in the previously transmitted data unit.
[0004] U.S. patent application Ser. No. 17 / 402,547 describes an electronic device comprising a bus driver and circuitry. The bus driver is coupled to a parallel bus comprising N data lines. The circuitry is configured to receive a data unit for transmission over the N data lines, determine a first count indicating the number of data bits having a predefined value in the data unit and a second count indicating the number of data bits that have been inverted relative to corresponding bits in a previously transmitted data unit, make a decision based on the first and second counts whether to invert the data unit depending on whether such inversion is desired to reduce power consumption in transmitting the data unit over the bus, generate an output data unit by retaining or inverting the data unit based on the decision, and transmit the output data unit over the data lines via the bus driver.
[0005] U.S. Patent Application Publication No. 2016 / 0173134 describes methods and apparatus related to Enhanced Data Bus Inversion (EDBI) encoding for an OR chain bus. In one embodiment, incoming data on a bus is encoded based at least in part on a determination of whether the next data value on the bus will transition from a valid value to a paused state. Summary of the Invention
[0006] Embodiments described herein provide an electronic device comprising circuitry and a plurality of ports. The plurality of ports comprises an input port and an output port, the input port and the output port being configured to transmit a data unit with one or more other devices across a system-on-chip (SoC) configuration, the data unit comprising N data bits, where N is an integer greater than 1. The circuitry is configured to receive an input data unit via the input port, make a random decision of whether to invert the N data bits in the input data unit, generate an output data unit by retaining or inverting the N data bits of the input data unit based on the random decision, and transmit the output data unit via the output port.
[0007] In some embodiments, the circuit is configured to receive the input data unit after the input data unit has been processed in the network device included in the fabric, and to transmit the output data unit to the fabric's link via the output port. In other embodiments, the circuit is configured to receive the input data unit from the fabric's link via the input port, and to transmit the output data unit via the output port for processing in the network device included in the fabric. In still other embodiments, the circuit is configured to make the random decision independent of the value of the data bit in the received input data unit.
[0008] In one embodiment, the circuit is configured to receive a subsequent input data unit via the input port and make another decision, either randomly or based on the input data unit, whether to invert the subsequent input data unit. In another embodiment, the circuit is configured to receive one or more other input data units that passed through the fabric together with the input data unit via the input port and make a corresponding decision, based on the values of the output data unit and the other input data units, whether to invert the other input data units.
[0009] In some embodiments, the circuit is configured to make a first random decision of whether to invert a first subset of the N data bits of the input data unit, and to make a second random decision of whether to invert a second subset of the N data bits of the input data unit, such that making the second random decision is independent of making the first random decision. The circuit is further configured to generate the output data unit by retaining or inverting the first subset of the N data bits based on the first random decision and retaining or inverting the second subset of the N data bits based on the second random decision. In other embodiments, the circuit is configured to generate a first indication and a second indication, respectively, of whether the first subset of the N data bits and the second subset of the N data bits have been inverted, and to output the first indication and the second indication via the output port or via another interface of the electronic device. In yet another embodiment, the electronic device resides in a first location in the configuration, and the second electronic device that makes the random decision of whether to invert the data unit resides in a second, different location in the configuration, and the circuit is configured to make the random decision independently of the random decision made by the second electronic device.
[0010] According to one embodiment described herein, a method for bus inversion is further provided, the method comprising: in an electronic device including a plurality of ports (including input ports and output ports), transmitting a data unit with one or more other devices across a system-on-chip (SoC) configuration, the data unit comprising N data bits, where N is an integer greater than 1. An input data unit is received via the input port. A random decision is made as to whether to invert the N data bits in the input data unit. An output data unit is generated by retaining or inverting the N data bits of the input data unit based on the random decision. The output data unit is sent via the output port.
[0011] According to one embodiment described herein, an electronic system is further provided, comprising a configuration, circuitry, a plurality of proxy devices, and a plurality of bus inversion devices. The configuration comprises a plurality of network devices interconnected by links, each link comprising a plurality of lines for transmitting data units comprising a plurality of data bits. The proxy devices are coupled to communicate via the configuration. The bus inversion devices are incorporated at selected locations within the configuration, each bus inversion device comprising circuitry and a plurality of ports. The plurality of ports comprise input ports and output ports, the input ports and the output ports being configured to transmit the data units via the configuration. The circuitry is configured to receive an input data unit via the input port, make a random decision as to whether to invert at least some of the data bits in the input data unit, generate an output data unit by retaining or inverting at least some of the data bits in the input data unit based on the random decision, and transmit the output data unit via the output port.
[0012] In some embodiments, the architecture includes a plurality of sub-architectures for providing separate communications between respective subsets of the proxy devices.
[0013] According to one embodiment described herein, a method is further provided that includes, for a system on a chip (SoC) having a plurality of elements interconnected by a fabric including a plurality of links, calculating a number of electronic devices in the SoC required to meet a peak power requirement for transmitting a data unit through the fabric, the electronic devices making corresponding random decisions on whether to invert data units passing through the electronic devices. At least the calculated number of electronic devices is assigned to selected links or network devices in the SoC.
[0014] In some embodiments, calculating the number includes calculating the number based on a specified target switching rate across the configuration. In other embodiments, calculating the number includes calculating a minimum number that satisfies a probability of failure exceeding the target switching rate. In still other embodiments, calculating the minimum number includes calculating the probability of failure based on a cumulative binomial distribution function.
[0015] In one embodiment, the configuration has a plurality of available locations for performing bus reversal, and assigning the electronic device includes assigning the electronic device to at least some of the available locations. In another embodiment, assigning the electronic device includes assigning a plurality of bus reversal devices to a plurality of corresponding subsets of lines of at least some of the links in the available locations in response to identifying that the number of required electronic devices is greater than the number of available locations. In yet another embodiment, assigning the electronic device includes assigning the electronic device to a location at an input of a network device in the configuration. In yet another embodiment, assigning the electronic device includes assigning the electronic device to a location at an output of a network device in the configuration.
[0016] These and other embodiments will be more fully understood from the following detailed description of embodiments of the invention taken in conjunction with the accompanying drawings, in which: BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1A is a block diagram schematically illustrating a system on a chip (SoC) including a CPU network supporting random bus turnaround according to one embodiment described herein;
[0018] Figure 1B is a block diagram schematically illustrating a SoC including an IO network supporting random bus inversion according to another embodiment described herein;
[0019] Figure 2A is a block diagram schematically illustrating an electronic device for making random bus reversal decisions according to one embodiment described herein;
[0020] Figure 2B is a block diagram schematically illustrating an electronic device for making random bus inversion decisions and data-driven bus inversion decisions according to one embodiment described herein;
[0021] Figure 3 is a block diagram schematically illustrating an electronic device used as a bus receiver according to one embodiment described herein;
[0022] Figure 4 is a flow chart schematically illustrating a method for random bus reversal according to one embodiment described herein;
[0023] Figure 5 is a flow chart schematically illustrating a bus inversion method applied to packets according to one embodiment described herein;
[0024] Figure 6 is a flow chart schematically illustrating a method for determining the number of random bus inversion (RBI) devices required in a SoC to meet specified power consumption requirements according to one embodiment described herein;
[0025] Figure 7 is a block diagram schematically illustrating a SoC including multiple internal networks supporting random bus reversal according to one embodiment described herein; and
[0026] Figure 8 is a block diagram schematically illustrating a multi-die system supporting random bus reversal according to one embodiment described herein. DETAILED DESCRIPTION
[0027] Overview
[0028] Embodiments described herein provide methods and systems for mitigating power consumption spikes in a system-on-chip (SoC) by randomly applying bus inversion operations across the SoC fabric.
[0029] SoCs typically include proxy devices that communicate through one or more internal structures. The proxy devices may include, for example, a central processing unit (CPU), a graphics processing unit (GPU), a memory controller (MC) coupled to a memory device, and input / output (IO) peripheral devices. SoC structures are also referred to herein as "SoC networks." Example SoC networks include a CPU network in which the CPU can communicate with the MC, an IO network in which peripheral devices can communicate with the CPU and MC, and a loosely ordered network in which the GPU can communicate with the MC.
[0030] During SoC operation, peaks in power consumption or current may overwhelm the SoC's power delivery system, causing the voltage to drop below an acceptable level. In addition, power consumption peaks may cause physical damage to the SoC, for example, due to overheating.
[0031] In principle, power consumption in an SoC can be reduced by performing conventional data bus inversion (DBI) on each link of the SoC network. However, when replicated across many links, conventional DBI circuitry is expensive in terms of chip area and power consumption. Furthermore, DBI circuitry typically introduces a delay of one or more transmission cycles, which can accumulate to significant delays along a network path with multiple DBI circuits.
[0032] In the disclosed embodiment, power consumption peaks are mitigated by randomly applying bus inversions in selected locations of the SoC network. In the following description, the electronic device that makes the random decision of whether to invert a data unit is referred to as a random bus inversion (RBI) device. A single RBI device requires very little chip area and does not introduce significant delay. By combining a sufficiently large number of RBI devices across the SoC network, peak power consumption events are essentially eliminated.
[0033] A SoC network typically comprises network switches (NS) interconnected by links in a suitable topology. A link comprises multiple lines for transmitting multi-bit data units. An agent device may be coupled to the NS directly or via a suitable network interface (NI).
[0034] When a data unit passes through a link or network device, a "bit switch" occurs when the corresponding bits in two consecutive data units have opposite bit values. The average number of bit switches across the SoC network within a specified time window relative to the maximum number of bits transmitted across the SoC network in a specified time window is referred to as the "switching rate" in this article. The product of the switching rate and the construction utilization factor is generally highly correlated with the amount of power consumed in the SoC. Generally speaking, a high switching rate generally leads to high power consumption, and vice versa. Therefore, power consumption peaks can be mitigated by controlling the switching rate in the SoC.
[0035] Certain traffic patterns through the fabric can result in high switching rates. For example, a power virus program may impose a 100% switching rate. As will be described below, using the disclosed random bus reversal technique, the probability of exceeding the desired target switching rate is reduced, substantially eliminating power consumption peaks.
[0036] Consider an embodiment of an electronic device comprising circuitry and a plurality of ports (including input ports and output ports). These ports transmit data units to one or more other devices across a system-on-chip (SoC) configuration. The data unit comprises N data bits, where N is an integer greater than 1. The circuitry is configured to receive an input data unit via the input port and make a random decision on whether to invert the N data bits in the input data unit. The circuitry is further configured to generate the output data unit by retaining or inverting the N data bits of the input data unit based on the random decision, and to send the output data unit via the output port.
[0037] In some embodiments, the device resides at an output of a network device (e.g., a network switch or a network interface) in the fabric, in which case the circuitry receives input data units after the input data units have been processed in the network device included in the fabric and transmits output data units to a link of the fabric via an output port. In other embodiments, the device resides at an input of a network device in the fabric, and the circuitry receives input data units from a link of the fabric via the input port and transmits output data units via the output port for processing in the network device included in the fabric.
[0038] In one embodiment, the circuit makes a random decision independent of the values of the data bits in the received input data unit (and independent of the values of other data units received or transmitted by the network device). In another embodiment, the circuit receives a subsequent input data unit via the input port and makes another decision, either randomly or based on the input data unit, whether to invert the subsequent input data unit. In yet another embodiment, the circuit receives one or more other input data units that pass through the fabric together with the input data unit via the input port and makes a corresponding decision whether to invert the other input data units based on the values of the output data unit and the other input data units.
[0039] In some embodiments, the circuit makes a first random decision as to whether to invert a first subset of N data bits of an input data unit, and makes a second random decision as to whether to invert a second subset of the N data bits of the input data unit, wherein making the second random decision is independent of making the first random decision. The circuit generates an output data unit by retaining or inverting the first subset of N data bits based on the first random decision and retaining or inverting the second subset of N data bits based on the second random decision. The circuit also generates respective first and second indications of whether the first and second subsets of N data bits have been inverted, and outputs the first and second indications via an output port or via another interface of the electronic device.
[0040] In some embodiments, multiple electronic devices that make random decisions about whether to invert data cells reside in various locations in the architecture. In such embodiments, each electronic device makes a local random inversion decision independent of the random inversion decisions made by other electronic devices. These locations are selected so that the switching rate (using RBI devices) is expected to be forced toward 50%. It should be noted that the primary goal is to reduce power consumption and prevent power consumption spikes under challenging conditions, such as under attacks from power viruses that cause a 100% switching rate. When the switching rate (without bus inversion) is below 50%, using RBI devices can even increase power consumption, but not to dangerous levels.
[0041] As noted above, random bus inversion can be used to mitigate power consumption peaks. Consider an embodiment of an electronic system (e.g., a SoC) that includes a fabric, a plurality of proxy devices, and a plurality of bus inversion devices. The fabric includes a plurality of network devices interconnected by links, each link including a plurality of lines for transmitting data units including a plurality of data bits. The proxy devices are coupled to communicate via the fabric. The bus inversion devices are incorporated at selected locations in the fabric, wherein each bus inversion device includes circuitry and a plurality of ports for transmitting data units through the fabric, the ports including input ports and output ports. The circuitry is configured to receive an input data unit via the input port, make a random decision whether to invert at least some of the data bits in the input data unit, generate an output data unit by retaining or inverting at least some of the data bits of the input data unit based on the random decision, and send the output data unit via the output port.
[0042] In some embodiments, the architecture includes a plurality of sub-architectures for providing separate communications between respective subsets of the proxy devices.
[0043] Further described below is a method for determining the number and location of RBI devices in a system-on-chip (SoC), the method comprising: for a system-on-chip (SoC) having a plurality of elements interconnected by a fabric including a plurality of links, calculating a number of electronic devices in the SoC required to meet a peak power requirement for transmitting a data unit through the fabric, the electronic devices making corresponding random decisions on whether to reverse data units passing through the electronic devices. At least the calculated number of electronic devices is assigned to selected links or network devices in the SoC.
[0044] In some embodiments, calculating the number includes calculating the number based on a specified target switching rate across the configuration, such as by calculating a minimum number that satisfies a probability of failure to exceed the target switching rate. In some embodiments, calculating the minimum number includes calculating the probability of failure based on a cumulative binomial distribution function.
[0045] Distributing electronic devices can be implemented in various ways. For example, the SoC network has multiple available locations for performing bus inversion, and electronic devices are distributed to at least some of the available locations. Generally speaking, electronic devices can be distributed to locations at the input of network devices in the configuration, or to locations at the output of network devices in the configuration.
[0046] In one embodiment, in response to identifying that the number of required electronic devices is greater than the number of available positions, a plurality of bus turn devices are assigned to respective subsets of the wires of at least some of the links in the available positions.
[0047] In the disclosed technology, an electronic device that randomly applies bus reversal operations is incorporated into the fabric of a system-on-chip (SoC). The electronic device consumes very little chip area and has no significant impact on communication latency. Using the disclosed embodiments, the probability of exceeding a specified target switching rate corresponding to a peak power consumption event can be reduced to an acceptable level.
[0048] System Description
[0049] Figure 1A is a block diagram schematically illustrating a system on a chip (SoC) 20 including a CPU network 22 supporting random bus turnarounds according to one embodiment described herein.
[0050] In SoC 20, the proxy devices communicating over the CPU network include a CPU cluster 24, which includes one or more processors 26, and a memory controller (MC) 30, each of which is coupled to one or more external memory devices 31 (depicted in dashed lines). Links 32 are used for point-to-point connections within SoC 20.
[0051] In this example, the CPU network 22 includes network switches (NS) 28A, 28B, and 28C interconnected in a ring topology. Alternatively, other suitable topologies may be used. The NS is also coupled to an agent device in the SoC via a suitable network interface, as described herein.
[0052] The CPU cluster 24 is coupled to the CPU network via a network interface (NI) 36, designated CP-NI, and NS 28A. Each of the MCs 30 is coupled to the CPU network via a NI 40, designated MC-NI, and NS 28B. In some embodiments, the processors 26 can store data in and read data from the memory devices 31 by transmitting appropriate transactions via the CPU network 22.
[0053] In some embodiments, SoC 20 supports communication with one or more other SoCs. In this example, NI 44, denoted as C2C-NI, is coupled to the CPU network via NS28C. C2C-NI extends the CPU network through one or more other SoCs.
[0054] The CPU network 22 includes a plurality of random bus inversion (RBI) devices 60. The RBI devices are depicted as arrows indicating the direction of traffic output by the RBI devices. The purpose of incorporating RBI devices 60 into SoC networks such as the CPU network and / or the IO network is to increase the randomization of traffic traversing the SoC network in order to shift the switching rate in the SoC toward 50%.
[0055] The CPU network 22 also includes a plurality of bus receiver devices 64 that are incorporated at edge locations (not shown) of the CPU network. The bus receiver devices terminate network paths that include one or more RBI devices 60.
[0056] Generally speaking, RBI devices can be incorporated into locations where power consumption increases with the number of bit switches between consecutive data units.Such locations include the outputs of network devices such as NS and NI and the inputs of network devices such as NI.
[0057] In some embodiments, at least some of the NSs in the fabric support a fast path for transactions (data units) that need to pass through the NS within a single cycle. In such embodiments, transactions that may have a delay exceeding a single cycle period are subject to RBI processing, while single cycle transactions are not subject to RBI processing.
[0058] The RBI device can be incorporated at the selected location in a variety of ways. For network device outputs, the RBI device can be incorporated into the network device before the output port, as part of the output port circuitry, or between the output port and the link. For network device inputs, the RBI device can be incorporated between the link and the input port, as part of the input port circuitry, or after the input port.
[0059] In some embodiments, the RBI device 60 can operate in a random mode or in a mixed mode. In the random mode, the RBI device 60 makes a random decision whether to invert a data unit. In this context and in the claims, the term "random inversion decision" or "random decision" means that, on average, the decision to invert a data unit is made with a probability of approximately 50%. The actual probability distribution and the probability of any given decision may be different from 50%, for example, depending on the technology used to generate the decision. In addition, the term "random decision" also refers to bus inversion decisions that are made "pseudo-randomly", so that the sequence of pseudo-random decisions appears to be "random" even though it is generally generated by a deterministic and repeatable process. Figure 2A and Figure 2B Describes a mechanism for making pseudorandom decisions using a pseudorandom number generator (PRNG).
[0060] In hybrid mode, the RBI device 60 makes random inversion decisions for selected data units and makes data-driven decisions (e.g., as conventional DBI does) for other data units. Figure 2A and Figure 2B The structures of the RBI device operating in random mode and in mixed mode are described in detail respectively.
[0061] Figure 1B is a block diagram schematically illustrating a SoC 80 including an IO network 82 supporting random bus inversion according to another embodiment described herein.
[0062] In SoC 80, the proxy devices communicating via IO network 82 include: CPU cluster 24, which includes one or more processors 26; MC 30, each coupled to one or more external memory devices 31 (depicted in dashed lines); and IO cluster 84, which includes one or more peripheral devices 88. Links 32 are used for point-to-point connections within SoC 80. In addition to the CPU network also providing communication between processors 26 and MC 30, IO network 82 also supports communication between peripheral devices 88 and MC 30.
[0063] In SoC 80, IO cluster 84 is coupled to IO network 82 via IO interface 90, denoted as IO-NI, and NS 28A. In some embodiments, peripheral devices 88 can store data in and read data from memory device 31 by transmitting appropriate transactions over IO network 82.
[0064] The IO network 82 includes RBI devices 60 that are incorporated into various locations in the IO network (such as the output and input ends to network devices), as described above. The IO network 82 also includes a plurality of bus receiver devices 64 that are incorporated into edge locations (not shown) of the IO network for terminating a network path that includes one or more RBI devices 60.
[0065] In some embodiments, a SoC (e.g., such as SoCs 20 and 80) supports activation and deactivation of RBI devices. For example, the RBI device can be activated under heavy load and deactivated when no excessive power consumption and / or overheating is expected (e.g., even at 100% switching rate). For example, the RBI device can be deactivated when the utilization factor is below 50%, in which case the use of random bus inversion can increase average power consumption.
[0066] Figure 1A and Figure 1B The SoC and SoC network configurations shown are given as examples, and other suitable SoC and SoC network configurations may also be used. For example, in alternative embodiments, proxy devices such as CPU clusters, MCs, and IO clusters may be instantiated as multiple discrete instances with their own corresponding NIs. Similarly, in alternative embodiments, NSs may be replicated as necessary. In addition, CPU clusters and / or IO clusters such as Figure 1A and Figure 1B Those shown may be increased or subdivided as required.
[0067] Structure of RBI and receiver equipment
[0068] Figure 2A is a block diagram schematically illustrating an electronic device 100 that makes random bus inversion decisions according to one embodiment described herein.
[0069] The electronic device 100 acts as an RBI device operating in a random mode and can be used to implement the above Figure 1A and Figure 1B The RBI device 60 in the SoC 20 and 80.
[0070] RBI device 100 includes an input port 102A for receiving input data units 104 and an output port 106 for outputting output data units 108. The input data units and the output data units include N-bit data units, where N is a positive integer.
[0071] The RBI device 100 includes a bus inversion (BI) decider 114 that is configured to make a random decision 116 on whether to invert an input data unit 104. The decider 114 makes the random decision 116 for a given data unit independently of the values of the given data unit and other data units received or output by the RBI device. In the RBI device 100, a multiplexer 120 generates an output data unit 108 by selecting an input data unit 104 or an inverted version of the input data unit based on the random decision 116.
[0072] In some embodiments, the decision maker 114 pseudo-randomly makes the random decision using a pseudo-random number generator (PRNG) 118. The PRNG generates a randomly occurring, cyclical, and deterministic sequence of numbers. In some embodiments, the decision maker 114 makes the random inversion decision (116) by comparing the corresponding number generated by the PRNG to a predefined threshold number. In one embodiment, the numbers generated by the PRNG are uniformly distributed within the predefined range of numbers. By setting the threshold number to a mid-range value, the decision maker 114 makes a decision to invert the input data unit with a probability of 50% (or approximately 50%).
[0073] In some embodiments, the RBI device 100 receives an input polarity signal 124 via the input interface 102B (e.g., along with the input data unit 104). The RBI 100 generates a corresponding output polarity signal 126 and transmits it along with the output data unit 108 via the output interface 106B. The output polarity signal serves as an input polarity signal for the next RBI in the network path or as a final polarity signal for a bus receiver (e.g., bus receiver 64). The RBI device 100 includes a multiplexer 130 that selects between the input polarity signal and an inverted input polarity signal based on the random decision 116.
[0074] Figure 2B is a block diagram schematically illustrating an electronic device 150 that makes random bus inversion decisions and data-driven bus inversion decisions according to one embodiment described herein.
[0075] The electronic device 150 acts as an RBI device operating in a hybrid mode and can be used to implement the above Figure 1A and Figure 1B The RBI device 60 in the SoC 20 and 80.
[0076] Similar to RBI device 100, RBI device 150 receives data units 104 via input port 102A, generates corresponding output data units 108, and transmits the output data units via output port 106A. For each input data unit, BI decision 154 generates BI decision 158, which controls multiplexer 120 to select the input data unit or an inverted version of the input data unit for the output data unit. RBI device 150 receives input polarity signal 124 via interface 102B and generates output polarity signal 126 based on BI decision 158 using multiplexer 130.
[0077] For each input data unit 104, the BI decision maker generates a BI decision using either a random decision maker 162 or a data driven decision maker 166. The random decision maker 162 is substantially similar to the above Figure 2A The random decision maker 114 may be configured to make random inversion decisions, for example, using the PRNG 118, as described above. The data-driven decision maker 166 makes BI decisions based on the values in one or more data cells. In this example, the RBI device 150 includes a latch 168 that latches the value of the output data cell 108. The data-driven decision maker makes data-driven decisions based on the current input data cell and the previously transmitted output data cell.
[0078] In some embodiments, the RBI device 150 is incorporated into a SoC network that supports packet-based communication, where each packet includes multiple data units that pass through the fabric together. For multiple input data units received via an input port and belonging to the same packet, the BI decider 154 makes a random decision for the first input data unit in the packet (using the random decider 162) and makes a data-driven decision for one or more other data units in the packet (using the data-driven decider 166).
[0079] Figure 3 is a block diagram schematically illustrating an electronic device 180 acting as a bus receiver according to one embodiment described herein.
[0080] The bus receiver 180 can be used to Figure 1A and Figure 1B The SoCs 20 and 80 are used to terminate communications via a network path that includes one or more RBI devices 60. However, the processing in the bus receiver 180 is independent of the actual bus inversion method used and, for example, the bus receiver may be adapted for use in configurations employing any suitable bus inversion technology, such as conventional DBI or a combination of DBI and RBI.
[0081] Bus receiver 180 receives data unit 182 via input port 184 and polarity signal 186 via input interface 188. Alternatively, both the data unit and the polarity signal may be received via the input port.
[0082] Data unit 182 and polarity signal 186 are generated by an RBI device (e.g., RBI device 100 or 150) that is the last RBI device along the network path from the source agent device to the destination agent device. Bus receiver 180 includes a multiplexer 190 that selects between data unit 182 and an inverted version of data unit 182 based on the polarity signal to produce a recovered data unit 194. The recovered data unit is equal to the data unit sent by the source agent.
[0083] Despite Figure 2A 、 Figure 2B and Figure 3 In an example embodiment of the present invention, the polarity signal is received and transmitted via a dedicated interface, but in an alternative embodiment, the polarity signal is transmitted via the input port and the output port used for communication of the data unit.
[0084] Method for bus inversion
[0085] Figure 4 is a flow chart schematically illustrating a method for random bus reversal according to one embodiment described herein.
[0086] The method will be described as being performed by an RBI device operating in a random mode (e.g., Figure 2A The RBI device resides on a SoC (such as, for example, Figure 1A and Figure 1B In the construction of SoC 20 or Soc 80).
[0087] The method begins at a receive phase 200, where the RBI device receives an input data unit and an input polarity signal. At a decision phase 204, the RBI device makes a random decision whether to invert the received data unit.
[0088] At decision application stage 208, the RBI device generates an output data unit and an output polarity signal based on the random decision of stage 204. Based on the random decision, the output data unit is equal to the received data unit or an inverted version of the received data unit, and the output polarity signal is equal to the input polarity signal or an inverted version of the input polarity signal. At output stage 212, the RBI device transmits the output data unit along with the output polarity signal (e.g., to the constructed link or constructed network device). After stage 212, the method terminates.
[0089] Figure 5 is a flow chart schematically illustrating a bus inversion method applied to packets according to one embodiment described herein;
[0090] The method will be described as being performed by an RBI device operating in a hybrid mode (e.g., Figure 2B The RBI device resides on a SoC (such as, for example, Figure 1A and Figure 1B In the construction of SoC 20 or SoC 80).
[0091] The method begins at a receiving phase 250, where the RBI device receives an input data unit belonging to a packet including two or more data units. At a query phase 254, the RBI device checks whether the received data unit is the first data unit in the packet, and if so, proceeds to a random decision-making phase 258. Otherwise, the RBI device proceeds to a data-driven decision-making phase 262.
[0092] At stage 258, the RBI device makes a random decision of whether to invert the input data unit for generating the corresponding output data unit. At stage 262, the RBI device makes a non-random data-driven decision of whether to invert the input data unit for generating the corresponding output data unit based on the value of one or more data units of the packet (e.g., the current input data unit and the previously transmitted data unit).
[0093] At decision application stage 266, the RBI device generates an output data unit based on the inverted decision of stage 258 or 262. At output stage 270, the RBI device sends the output data unit (e.g., to the constructed link or constructed network device). After stage 266, the method terminates.
[0094] Design Methodology for Determining the Number of RBI Devices Required in an SoC
[0095] Figure 6 is a flow chart schematically illustrating a method for determining the number of random bus inversion (RBI) devices required in a SoC to meet specified power consumption requirements, according to one embodiment described herein.
[0096] The method is described as being performed by a processor as an off-line process.
[0097] The method begins with the processor receiving design requirements at a specification reception stage 300. In this example, the processor receives at stage 300: (i) a SoC structure specifying various components and interconnections in the underlying SoC, (ii) a target switching rate required to meet peak power requirements in the SoC, and (iii) a failure probability for failing to meet the target switching rate.
[0098] The peak power requirement specifies the maximum power consumption allowed within a predefined time window. The target switching rate specifies an upper limit on the ratio between the number of bit switches in the time window and the maximum number of bits transmitted across the SoC network within the time window. To meet the power consumption requirement, the switching rate in the SoC needs to be kept below the target switching rate.
[0099] The probability of failing to meet the target switching rate (denoted as "PF") is given by the following expression:
[0100] Equation 1:
[0101] PF = Pr (switching rate > target switching rate)
[0102] At the RBI device number determination phase 304, the processor calculates the (minimum) number of RBI devices required in the SoC to meet the peak power requirements. To this end, the processor calculates the number of RBI devices required to meet the target switching rate within a specified failure probability (making corresponding random reversal decisions).
[0103] At RBI allocation stage 306, the processor allocates the number of RBI devices determined at stage 304 for incorporation into the selected locations in the SoC network. After stage 306, the method terminates.
[0104] In some embodiments, at stage 304, the processor determines the number of RBI devices based on a binomial distribution function given by:
[0105] Equation 2:
[0106]
[0107] in
[0108]
[0109] The binomial distribution function depends on the parameters (n, k, p), where "n" represents the number of independent trials, each of which can succeed or fail, "k" represents the number of successful trials out of n trials, and "p" represents the probability of success in a single trial. The binomial distribution function in Equation 2 calculates the probability of k successful trials out of n trials.
[0110] The cumulative binomial distribution function calculates the probability of up to and including k successful trials out of n trials, as given by:
[0111] Equation 3:
[0112]
[0113] In this context, the term "test" refers to a single random bus reversal decision made by the RBI device. A test is considered successful when the random decision results in a flip of at most half a bit between an output data unit and the previous output data unit.
[0114] Let "FR" denote the bus frequency. FR indicates the number of data units traversing the link per time unit. The number of trials performed by a single RBI device within a time window denoted as "TW" is given by FR·TW. Furthermore, let "NR" denote the number of RBI devices coupled across the fabric. The total number of trials within the time window TW is given by:
[0115] Equation 4:
[0116] n=FR·TW·NR
[0117] In some embodiments, the processor estimates the probability of failure using the cumulative binomial distribution function of Equation 3. For example, given a target switching rate (expressed as TTR), failure occurs when the number of successful trials (k) does not exceed (1-TTR) of the total number of trials (n). The probability of failure is given by:
[0118] Equation 5:
[0119] PF=Pr(X≤(1-TTR)·n)=F(n,(1-TTR)·n,p=0.5)
[0120] For example, for a TTR of 60%, the PF is given by:
[0121] Equation 6:
[0122] PF=Pr(X≤(0.4·n)=F(n,0.4·n,p=0.5)
[0123] In some embodiments, a desired PF value is given, and the processor solves Equation 5 to determine the value of "n" that satisfies the desired PF. The processor then uses Equation 4 to determine the number of RBI devices. In practice, the number of RBI devices can be set large enough so that the probability of failure becomes negligible. Generally speaking, the number of RBI devices increases as the target switching rate decreases (limited to 50%). The inventors have found that with a target switching rate of 60%, a failure rate of one failure event in several years can be achieved in practical SoCs.
[0124] In some embodiments, the SoC network is configured with multiple locations that can be used to perform bus inversion, and the processor assigns the number of RBI devices (as calculated above) to at least some of the available locations. Generally speaking, the processor can assign RBI devices to inputs and / or outputs of network devices in the SoC network, as described above. Note that when one or more of the available locations remain free of RBI devices, the target switching rate can be adjusted accordingly.
[0125] When only a portion of a fabric is protected with an RBI device, the savings in power consumption are typically small compared to the savings from protecting the entire fabric. For example, for a target switching rate of 60%, the power savings with full protection is 40%, but for protection of 80% of the fabric, the power savings decreases to 32% (40% of 80%).
[0126] In some embodiments, the number of RBI devices required to achieve a target switching rate with a desired PF is greater than the number of locations available in the SoC for incorporating RBI devices. In such embodiments, the number of available locations can be increased by dividing at least some of the links in the SoC into two or more sectors, each of which includes a partial subset of the wires of the link. Sectors of the same link are assigned corresponding RBI devices, which make inversion decisions independently of each other (and independently of other RBI devices in the SoC).
[0127] Example SoC configuration incorporating random bus inversion devices
[0128] The random bus inversion implementation scheme described above can be extended to buses and structures of arbitrary complexity. Figure 7 and Figure 8 Describes example SoCs and configurations of medium / high complexity. U.S. Patent Application No. 17 / 337,805, filed June 3, 2021, “Multiple Independent On-chip Interconnects” describes Figure 7 and Figure 8 The disclosure of this U.S. patent application is incorporated herein by reference in its entirety for various aspects of the SoC in FIG. 1 (excluding the RBI device 350 and the bus receiver 354). To the extent of any inconsistency between this document and any document incorporated herein by reference, this document is intended to control.
[0129] Figure 7 is a block diagram schematically illustrating a SoC 320 including multiple internal networks supporting random bus reversal according to one embodiment described herein.
[0130] In this example, SoC 320 includes a CPU network, an IO network, and a relaxed sequential network. The CPU network provides communication between CPU clusters (such as 322A and 322B) and MCs (such as 326A...326D). The CPU network includes an interconnected NS 332 to which the CPU clusters and MCs are coupled. In some embodiments, the CPU clusters and / or MCs are coupled to the NS of the CPU network via appropriate NIs, which are omitted in the figure for clarity.
[0131] The IO network provides communication between IO clusters such as 324A...324D and CPU clusters, as well as between IO clusters and MCs. The IO network includes interconnected NSs 324 to which the IO clusters, CPU clusters, and MCs are coupled. In some embodiments, the CPU clusters, IO clusters, and MCs are coupled to the IO network NSs via appropriate NIs, which are omitted in the figure for clarity.
[0132] The loose order network provides communication between GPUs such as 328A ... 328D and MCs. The loose order network includes interconnected NSs 336 to which the GPUs and MCs are coupled. In some embodiments, the GPUs and MCs are coupled to the NSs of the loose order network via appropriate NIs, which are omitted in the figure for clarity.
[0133] In some embodiments, power consumption in SoC 320 is controlled based on the random bus inversion technique described above. In such embodiments, SoC 320 includes RBI device 350 and bus receiver 354. RBI device 350 may be used, for example, Figure 2A RBI device 100 or Figure 2B The bus receiver 354 can be implemented using, for example, the RBI device 150. Figure 3 This is achieved by the bus receiver 180.
[0134] The RBI device 350 and bus receiver 354 are incorporated into selected locations within one or more of the CPU network, IO network, and / or loosely ordered network. For example, as described above, the RBI device 350 can be incorporated at the output ports of the NS and NI and at the input ports of the NI. The bus receiver 354 can be incorporated at the edge of various SoC networks, as described above.
[0135] Figure 8 is a block diagram schematically illustrating a multi-die system 310 supporting random bus rollover according to one embodiment described herein.
[0136] In this example, multi-die system 310 includes SoCs 320A and 320B implemented on two separate semiconductor dies. Each of SoCs 320A and 320B may include, for example, Figure 7 3. Multi-die system 310 includes a CPU network 346, an IO network 344, and a relaxed sequential network 348. For clarity, the NS and NI of these networks are omitted. In multi-die system 310, each of the SoC networks extends across two SoC dies, forming a logically identical network even though the networks extend across two dies.
[0137] In some embodiments, power consumption in multi-die system 310 is controlled based on a random bus inversion technique. In such embodiments, multi-die system 310 includes RBI device 350 and bus receiver 354 incorporated into selected locations of the SoC network in SoCs 320A and 320B, as described above.
[0138] The configurations of SoCs 20, 80, and 320, multi-die system 310, RBI devices 100 and 150, and bus receiver 180 are example configurations, which are selected solely for conceptual clarity. In alternative embodiments, other suitable SoCs, multi-die systems, RBI devices, and bus receiver configurations may also be used. For clarity, elements not necessary for understanding the principles of the present invention, such as various interfaces, addressing circuits, timing and sequencing circuits, and debug circuits, have been omitted from the drawings.
[0139] Some elements of the SoCs 20, 80, and 320 and the multi-die system 310 (such as the CPU clusters 24 and 322A...322B, the IO clusters 84 and 324A...324D), and some elements of the RBI devices 100 and 150 (such as the decision makers 114 and 154) may be implemented in hardware, for example, in one or more application-specific integrated circuits (ASICs) or FPGAs. Additionally or alternatively, the CPU clusters 24 and 322A...322B and the GPUs 328A...328D may be implemented using software or a combination of hardware and software elements.
[0140] In some embodiments, some functions in CPU clusters 24 and 322A...322B and GPUs 328A...328D may be implemented by a general-purpose processor that is programmed with software to implement the functions described herein. The software may be downloaded to the relevant processor in electronic form, for example, over a network, or alternatively or additionally, it may be provided and / or stored on non-transitory tangible media such as magnetic, optical, or electronic memory.
[0141] The above-described embodiments are given by way of example, and other suitable embodiments may also be used.
[0142] It should be understood that the above embodiments are cited by way of example, and the following claims are not limited to what has been particularly shown and described above. On the contrary, the scope includes both combinations and subcombinations of the various features described above, as well as variations and modifications of the various features that will occur to a person skilled in the art upon reading the foregoing description and that are not disclosed in the prior art. The documents incorporated by reference in this patent application are considered an integral part of this application, but if any term is defined in these incorporated documents to conflict with a definition explicitly or implicitly made in this specification, only the definition in this specification shall be considered.
Claims
1. An electronic device, comprising: a plurality of ports, the plurality of ports comprising input ports and output ports, the plurality of ports configured to communicate a data unit with one or more other devices across a network in a system on a chip (SoC), the data unit comprising N data bits, where N is an integer greater than 1; and A circuit, the circuit being configured to: receiving an input data unit via the input port; making a random decision of whether to invert the N data bits in the input data unit, wherein the random decision is made independent of the values of the data bits in the received input data unit; generating an output data unit by retaining or inverting the N data bits of the input data unit based on the random decision; as well as The output data unit is sent via the output port.
2. The electronic device of claim 1 , wherein the circuit is configured to receive the input data unit after the input data unit has been processed in a network device included in the network, and is configured to transmit the output data unit to a link of the network via the output port.
3. The electronic device of claim 1 , wherein the circuit is configured to receive the input data unit from a link of the network via the input port and to send the output data unit via the output port for processing in a network device included in the network. 4 . The electronic device of claim 1 , wherein the circuit is configured to receive a subsequent input data unit via the input port and to make another decision whether to invert the subsequent input data unit, either randomly or based on the input data unit.
5. An electronic device according to claim 1, wherein the circuit is configured to receive one or more other input data units that pass through the network together with the input data unit via the input port, and is configured to make a corresponding decision on whether to invert the other input data units based on the values of the output data unit and the other input data units.
6. The electronic device of claim 1 , wherein the circuitry is configured to make a first random decision whether to invert a first subset of the N data bits of the input data unit, and is configured to make a second random decision whether to invert a second subset of the N data bits of the input data unit, and is configured to generate the output data unit by retaining or inverting the first subset of the N data bits based on the first random decision and retaining or inverting the second subset of the N data bits based on the second random decision, wherein making the second random decision is independent of making the first random decision.
7. An electronic device according to claim 6, wherein the circuit is configured to generate respective first and second indications of whether the first subset of the N data bits and the second subset of the N data bits have been inverted, and is configured to output the first and second indications via the output port or via another interface of the electronic device.
8. The electronic device of claim 1 , wherein the electronic device resides in a first location in the network and a second electronic device that makes the random decision of whether to invert a data unit resides in a second, different location in the network, and wherein the circuitry is configured to make the random decision independently of the random decision made by the second electronic device.
9. A method for bus inversion, the method comprising: In an electronic device including a plurality of ports, a data unit is communicated to one or more other devices across a network in a system on a chip (SoC), the plurality of ports including an input port and an output port, the data unit including N data bits, where N is an integer greater than 1. receiving an input data unit via the input port; making a random decision of whether to invert the N data bits in the input data unit, wherein the random decision is made independent of the values of the data bits in the received input data unit; generating an output data unit by retaining or inverting the N data bits of the input data unit based on the random decision; and The output data unit is sent via the output port.
10. The method of claim 9, wherein receiving the input data unit comprises receiving the input data unit after the input data unit has been processed in a network device included in the network, and wherein sending the output data unit comprises transmitting the output data unit to a link of the network via the output port.
11. The method of claim 9, wherein receiving the input data unit comprises receiving the input data unit from a link of the network via the input port, and wherein sending the output data unit comprises sending the output data unit via the output port for processing in a network device included in the network.
12. A method according to claim 9 and comprising receiving a subsequent input data unit via the input port and making a further decision whether to invert the subsequent input data unit, either randomly or based on the input data unit.
13. A method according to claim 9, and the method includes receiving one or more other input data units that pass through the network together with the input data unit via the input port, and making a corresponding decision whether to invert the other input data units based on the values of the output data unit and the other input data units.
14. The method of claim 9 , further comprising making a first random decision whether to invert a first subset of the N data bits of the input data unit, making a second random decision whether to invert a second subset of the N data bits of the input data unit, wherein making the second random decision is independent of making the first random decision, and wherein generating the output data unit comprises retaining or inverting the first subset of the N data bits based on the first random decision and retaining or inverting the second subset of the N data bits based on the second random decision.
15. A method according to claim 14, and the method includes generating respective first and second indications of whether the first subset of the N data bits and the second subset of the N data bits have been inverted, and outputting the first and second indications via the output port or via another interface of the electronic device.
16. The method of claim 9, wherein the electronic device resides in a first location in the network and a second electronic device that makes a random decision of whether to invert a data unit resides in a second, different location in the network, and wherein making the random decision comprises being independent of a random decision made by the second electronic device.
17. An electronic system, comprising: a network comprising a plurality of network devices interconnected by links, each link comprising a plurality of lines for transmitting a data unit comprising a plurality of data bits; a plurality of proxy devices coupled to communicate via the network; a plurality of bus inversion devices incorporated at selected locations in the network, wherein each bus inversion device comprises: a plurality of ports, the plurality of ports comprising input ports and output ports, the plurality of ports configured to transmit the data unit over the network; and A circuit, the circuit being configured to: receiving an input data unit via the input port; making a random decision of whether to invert at least some of the data bits in the input data unit, wherein the random decision is made independent of values of the data bits in the received input data unit; generating an output data unit by retaining or inverting the at least some of the data bits of the input data unit based on the random decision; and The output data unit is sent via the output port.
18. The electronic system of claim 17, wherein the network comprises a plurality of sub-networks for providing separate communications between respective subsets of the proxy devices.
Citation Information
Patent Citations
Enhanced Data Bus Invert Encoding for OR Chained Buses
US20160173134A1
Multiple Independent On-chip Interconnect
US20220334997A1
Methods for data bus inversion
US20230052170A1
Bus management unit and high safety system on chip
CN106383790A
Minimum input / output toggling rate for interfaces
CN113508370A