Burr-free multiplexer and prevention of burr propagation
By inserting storage elements into the sequential circuit and using a ready signal with delay matching, the problem of high power consumption caused by signal glitches is solved, the generation of glitch-free signals is achieved, and the energy consumption of the circuit is reduced.
Patent Information
- Application Number
- CN202110619839.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-02
- Filing Date
- 2021-06-03
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2041-06-03
AI Technical Summary
The propagation of signal glitches in existing sequential circuits causes frequent changes in the output signals of the combinational logic, which increases the number of node charging and discharging times and thus increases power consumption.
By inserting storage elements and using a delay-matched ready signal, the control signal is sampled after reaching its final value to generate a glitch-free output signal, avoiding the output signal changing multiple times within a clock cycle.
The number of node charging and discharging times in the signal receiving logic is reduced, power consumption is reduced, and energy efficiency of the circuit is improved.
Smart Images

Figure CN113890515B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to preventing glitches from propagating in a circuit. In particular, the present disclosure relates to eliminating glitches in a signal by inserting a storage element that samples the signal to produce a glitch-free output signal. BACKGROUND
[0002] Conventional sequential circuits include combinational logic with inputs driven by synchronous registers or flip-flops. "Combinational logic" refers to logic that receives one or more inputs, combines those inputs to produce an output, without storing a state of the inputs, outputs, or any intermediate values. In other words, combinational logic is "stateless" and can be asynchronous (not driven by a clock signal). In contrast, for sequential circuits (logic), registers store a state.
[0003] At the rising edge of the clock, the output of the register changes exactly once. However, multiple paths through the combinational logic can cause the signal of the combinational logic output to change multiple times before reaching its final level. The signal can be data input to a multiplexer or a select signal that causes the multiplexer to select one of the data inputs for output. The output of the multiplexer can change several times in response to changing data and select an input before settling on a final state. The changes in the multiplexer output are considered glitches, and additional combinational logic receiving the output can respond by charging and / or discharging nodes and dissipating power. Providing a glitch-free output signal can reduce the number of times nodes are charged and / or discharged, and thus the power dissipated by the additional combinational logic. These problems and / or other problems associated with the prior art need to be addressed. SUMMARY
[0004] When a signal glitches, the logic receiving the signal can change response, charging and / or discharging nodes within the logic and dissipating power. In the context of the following description, a glitch is at least one high pulse or low pulse of at least one bit of a signal within a clock cycle. In particular, the pulse is a high-low-high transition or a low-high-low transition. Providing a glitch-free signal can reduce the number of times nodes are charged and / or discharged, and thus the power dissipation. A technique to eliminate glitches in a signal is to insert a storage element that samples the signal after the signal change is complete to produce a glitch-free output signal. The storage element is enabled by a "ready" signal that has a delay that matches the delay of the circuit generating the signal. The ready storage element can prevent the output signal from changing until the final value of the signal is reached. The output signal changes only once, transitioning from low to high or high to low, typically reducing the number of times nodes in the logic receiving the signal are charged and / or discharged, and thus the power dissipation.
[0005] A method, computer readable medium, and system for preventing glitch propagation are disclosed. In one embodiment, a decoder circuit is configured to receive a select ready signal that is negated until a select signal generated by combinational logic is invariant and is assigned after the select signal is invariant. The decoder circuit generates at least one sample enable signal corresponding to a set of data input signals from the select signal, wherein the at least one sample enable signal is negated while the select ready signal is negated and is assigned in response to the assignment of the select ready signal. The decoder circuit generates a hold signal that is assigned while the at least one sample enable signal is negated and is negated in response to the assignment of the at least one sample enable signal. A sample circuit is configured to hold an output signal invariant while the hold signal is assigned and to sample one of the data input signals to transfer a level of the sampled data input signal to the output signal from the at least one sample enable signal in response to the hold signal being negated.
[0006] A method, computer readable medium, and system for preventing glitch propagation are disclosed. In one embodiment, a delay circuit is configured to generate a ready signal that is negated at a first transition of a clock signal and is assigned after a first delay relative to the first transition, wherein the first delay is at least as long as a second delay. A sample circuit is configured to receive an input signal generated by combinational logic, wherein a change in a first signal received at an input of the combinational logic after the second delay after a first transition of the clock signal causes a corresponding change in the input signal at an output of the combinational logic. The sample circuit is further configured to sample the input signal to transfer a level of the input signal to an output signal of the sample circuit while the ready signal is assigned, wherein the input signal is invariant from the second delay until the input signal is sampled. BRIEF DESCRIPTION OF DRAWINGS
[0007] Figure 1A A block diagram of a glitchless sample circuit and combinational logic is shown in accordance with one embodiment.
[0008] Figure 1B A block diagram of a glitchless sample circuit and combinational logic is shown in accordance with one embodiment. Figure 1A A timing diagram of the circuit shown in
[0009] Figure 1C A flowchart of a method for generating a glitchless signal is shown in accordance with one embodiment.
[0010] Figure 1D A glitchless sample circuit is shown in accordance with one embodiment.
[0011] Figure 1E An overlapping inverter circuit is shown in accordance with one embodiment.
[0012] Figure 1F Another glitch-free sampling circuit is shown in accordance with one embodiment.
[0013] Figure 1G An asymmetric ready signal generation circuit is shown in accordance with one embodiment.
[0014] Figure 1H A timing diagram of the asymmetric ready signal generation circuit of Figure 1G is shown in accordance with one embodiment.
[0015] Figure 2A A block diagram of a glitch-free N-to-l multiplexer is shown in accordance with one embodiment.
[0016] Figure 2B A timing diagram of the glitch-free N-to-l multiplexer of Figure 2A is shown in accordance with one embodiment.
[0017] Figure 2C A block diagram of a timing decoder is shown in accordance with one embodiment. Figure 2A
[0018] Figure 2D A flowchart of a method for generating a glitch-free multiplexer output signal is shown in accordance with one embodiment.
[0019] Figure 2E An extended ready signal generation circuit is shown in accordance with one embodiment.
[0020] Figure 2F A timing diagram of the extended ready signal generation circuit of Figure 2E is shown in accordance with one embodiment.
[0021] Figure 2G A fast return circuit is shown in accordance with one embodiment.
[0022] Figure 2H A timing diagram of the fast return circuit of Figure 2G is shown in accordance with one embodiment.
[0023] Figure 3 A parallel processing unit is shown in accordance with one embodiment.
[0024] Figure 4A A general processing cluster within the parallel processing unit of Figure 3 is shown in accordance with one embodiment.
[0025] Figure 4B A memory partition unit of the parallel processing unit of Figure 3 is shown in accordance with one embodiment.
[0026] Figure 5A A flow multiprocessor according to one embodiment is shown. Figure 4A of the flow multiprocessor.
[0027] Figure 5B A conceptual diagram of a processing system implemented according to an embodiment using a PPU. Figure 3
[0028] Figure 5C An exemplary system in which various previous embodiments can be implemented is shown. DETAILED DESCRIPTION
[0029] In response to changing data and select inputs, the output of a multiplexer can change multiple times before settling to a final state. In the context of the following description, a change is a transition in voltage level that is identified as a different state. For example, a change is identified as an assigned state compared to an unassigned or negated state. In another example, a change is a transition from a high level to a low level or from a low level to a high level. In the context of the following description, a stable or constant level can vary while remaining within a range of voltage values that are identified as the same state (e.g., logic true or logic false).
[0030] Multiple changes in the output of a multiplexer are considered glitches, and the combinational logic that receives the output can respond by charging and / or discharging nodes within the combinational logic and dissipating power. Glitches can be prevented by inserting additional registers to reduce power consumption. For example, the data and select inputs can be registered by inserting flip-flops at the inputs to the multiplexer. The registers add pipeline stages and prevent glitches at the inputs to the multiplexer because each input only changes (or remains constant) once per clock edge. However, registers are very expensive in terms of power consumption, delay, and chip area. Alternatively, the output of the multiplexer can be registered using a delayed clock that ensures that all inputs to the multiplexer reach their final state and that the output of the multiplexer no longer changes before the output of the multiplexer is registered. Inserting a register at the output of the multiplexer can reduce power consumption of the multiplexer and additional combinational logic. However, a register is an expensive solution.
[0031] Another technique to eliminate glitches in a signal is to insert a storage element that samples the signal after a change is complete to produce a glitch-free output signal within a clock cycle. The storage element is enabled by a "ready" signal that has a delay that matches the delay of the circuit that generates the signal. This technique can prevent the output signal from changing until the final value of the signal is reached. Providing a glitch-free signal can reduce the number of times that nodes in the logic that receives the signal are charged and / or discharged, thereby reducing the power consumed by the logic.
[0032] For example, a multiplexer can be used to select only non-zero activations and / or weights for convolution operations in a convolutional neural network. If the select signal to the multiplexer glitches, the inputs to the multiplexer can glitch, causing the multiplexer to evaluate several different products, which can be several times more power consuming than evaluating a single product. As described further herein, providing a glitch-free select signal to select glitch-free non-zero values can reduce power consumption. In the context described below, a varying signal is ignored until the signal becomes stable, at which point the signal can be sampled so that the combinational logic receiving the signal evaluates once.
[0033] Figure 1A A block diagram 100 of a glitch-free sampling circuit 110 and combinational logic 103 and 113 is shown, according to one embodiment. A register outputs a signal A in response to a rising edge of a clock signal. In response to an input (not explicitly shown) to the register 101, the level of the input signal A can change or remain the same at each rising edge of the clock signal, and is stable (constant) between rising edges of the clock signal. Combinational logic 103 receives the input signal A and generates an output signal B. Combinational logic 103 can also receive one or more additional inputs and generate one or more additional outputs (not explicitly shown). Due to different timing paths in the combinational logic 103, the output signal B can transition through several intermediate states before stabilizing to a final value. If the glitched output signal B is directly input to combinational logic 113, such as a complex arithmetic circuit, the glitching can result in high power consumption.
[0034] A delay circuit 105 is configured to match the propagation delay for a transition of signal A to the effective output for signal B. In one embodiment, the propagation delay is equal to or greater than the worst case propagation delay. In one embodiment, the propagation delay is greater than the worst case propagation delay. In one embodiment, the delay circuit 105 includes an even number of inverters coupled in series. The delay circuit 105 receives a clock as an input and outputs a signal BR (B ready), which indicates when the signal B is stable (glitch-free) and should be sampled.
[0035] Sampling circuit 110 is a latch storage element configured to sample input D when enable input E is asserted, propagating the sampled level of input signal B to output Q to generate output signal BX. When signal BR is asserted and output signal BX "tracks" input signal B, sampling circuit 110 is transparent. Output signal BX is the input signal to combinatorial logic 113. When signal BR, coupled to enable input E, is negated, sampling circuit 110 becomes opaque, input signal B is not sampled, and signal BX remains stable. However, because input signal B is glitch-free when signal BR is asserted, output signal BX changes at most once—in response to the transition from negated signal BR to asserted signal. In other words, the level of B is sampled after reaching steady state and in response to signal BR being asserted.
[0036] Figure 1B According to one embodiment, Figure 1A The timing diagram of sampling circuit 110 is shown. Shortly after the rising edge of the clock signal, signal A is valid and stable. Signal A propagates through combinatorial logic 103, and signal B glitches, switching between high and low levels before finally reaching a stable level. The matched delay Δ provided by delay circuit 105 delays the clock signal to produce the rising edge of signal BR. The matched delay used to generate signal BR is long enough to ensure that signal B has completed its change and is stable. When signal BR is asserted, signal B is sampled and propagated to signal BX. When signal BR is negated, sampling of signal B stops, and the level of signal BX remains unchanged. Importantly, for each clock cycle, the glitch-free signal BX transitions once or only remains constant (when signal B is unchanged compared to the previous clock cycle). Sampling circuit 110 effectively filters out glitches, allowing combinatorial logic 113 to be evaluated only once, thereby minimizing power consumption.
[0037] In another embodiment, signal A is valid and stable on the falling edge of the clock signal rather than on the rising edge of the clock signal. In this alternative embodiment, when signal BR is negated, signal B is sampled and propagated to signal BX, and when signal BR is asserted, sampling of signal B stops and the level of signal BX is maintained. One of ordinary skill in the art will appreciate that executing block diagram 100 and Figure 1B Any logic of operation of the corresponding waveforms shown in is within the scope and spirit of the embodiments of the present disclosure.
[0038] Now, more illustrative information will be provided regarding various optional architectures and features that can be used to implement the aforementioned framework, depending on the user's needs. It should be noted that the following information is provided for illustrative purposes only and should not be construed as limiting in any way. Any of the following features may be incorporated into or without excluding the other features described, as needed.
[0039] Figure 1C A flowchart illustrating a method 115 for generating a glitch-free signal according to one embodiment is shown. The method 115 is described in the context of a logic or circuit and can also be executed within a processor. For example, the method 115 can be executed by a GPU (graphics processing unit), a CPU (central processing unit), or any processor capable of generating a glitch-free signal. Moreover, one of ordinary skill in the art will appreciate that any system executing the method 115 is within the scope and spirit of embodiments of the present disclosure.
[0040] At step 120, a ready signal is generated that is inverted at a first transition of the clock signal and is assigned a value after a first delay relative to the first transition, where the first delay is produced by a delay circuit configured to match the second delay. In one embodiment, the first delay matches the second delay when the first delay is equal to the second delay. In one embodiment, the first delay matches the second delay when the first delay is equal to or greater than the second delay. In one embodiment, the first delay is greater than the second delay and is within the same clock cycle. In one embodiment, the ready signal is BR, the first delay is produced by the delay circuit 105, and the second delay is the propagation delay through the combinational logic 103. In one embodiment, the delay circuit 105 inverts the clock signal to generate a delayed signal. In one embodiment, the delay circuit 105 includes a chain of inverters coupled in series.
[0041] At step 125, the input signal generated by the combinational logic is received and a change in the first signal causes a corresponding change in the input signal at the second delay after the first transition of the clock signal. In one embodiment, the input signal B generated by the combinational logic 103 is received at the sampling circuit 110 and a change in the signal A causes a corresponding change in the input signal B to occur at the matched delay. In one embodiment, the change in the first signal is a falling transition from a high level to a low level. In another embodiment, the change in the first signal is a rising transition from a low level to a high level.
[0042] At step 130, the input signal is sampled to transfer the level of the input signal to the output signal when the ready signal is assigned, where the input signal remains unchanged from the second delay until the sampling. In one embodiment, the input signal B is sampled by the sampling circuit 110 to transfer the level of the input signal B to the output signal BX. In one embodiment, the sampling circuit 110 is configured to hold the output signal BX at a constant level from the first transition of the clock signal until the input signal B is sampled. In one embodiment, the input signal B changes (glitches) at least once after the first transition of the clock signal and before the second delay.
[0043] In an embodiment, the delay circuit is further configured to invert the ready signal after a third delay relative to the first delay, wherein the ready signal is inverted prior to the next clock cycle. In one embodiment, the sampling circuit is further configured to generate an output ready signal that is inverted for the first delay and is assigned a value once the level of the input signal is transferred to the output signal. In one embodiment, the output ready signal is inverted in response to the assignment of the value.
[0044] Figure 1D A glitchless sampling circuit 110 is shown in accordance with one embodiment. The sampling circuit 110 includes two "AND" gates 122 and 124, an "OR" gate 126, and an enable inverter 118. When both the enable E and D inputs are high (assigned), the output of the AND gate 122 is high, and thus, the output Q of the OR gate 126 is assigned a value. Otherwise, the output of the gate 122 is low (inverted). The output Q of the OR gate 126 is fed back within the sampling circuit 110 and input to the AND gate 124. The enable signal E is inverted by the enable inverter 118 to produce a signal NE (not enable) that is input to the AND gate 124. Thus, when the enable signal is high, the AND gate 124 effectively disables the feedback path. When the enable signal is low (NE is high), the AND gate 124 enables the feedback path, propagating the level of the output Q to the OR gate 126 to hold the level of the output Q until the enable signal is assigned a value. In other words, the enable inverter 118 is configured to enable the feedback path to hold the output signal Q constant until the D input signal is sampled. The enable inverter 118 ensures that when the enable signal E transitions from low to high, the NE signal remains high such that the feedback path remains enabled through the AND gate 124 until the enable signal rises and the input D propagates through to the output Q, ensuring that Q does not glitch as the enable E is assigned a value. For example, if when E transitions high, before the high level D propagates through the AND gate 122, if the NE signal transitions low and drives the output of the AND gate 124 low, then both inputs to the OR gate 126 will be low at the same time, causing the output Q to go to a glitched low before the output of the "AND" gate 122 drives the output Q high.
[0045] The output Q of the OR gate 126 is fed back within the sampling circuit 110 and input to the AND gate 124. The enable signal E is inverted by the enable inverter 118 to produce a signal NE (not enable) that is input to the AND gate 124. Thus, when the enable signal is high, the AND gate 124 effectively disables the feedback path. When the enable signal is low (NE is high), the AND gate 124 enables the feedback path, propagating the level of the output Q to the OR gate 126 to hold the level of the output Q until the enable signal is assigned a value. In other words, the enable inverter 118 is configured to enable the feedback path to hold the output signal Q constant until the D input signal is sampled. The enable inverter 118 ensures that when the enable signal E transitions from low to high, the NE signal remains high such that the feedback path remains enabled through the AND gate 124 until the enable signal rises and the input D propagates through to the output Q, ensuring that Q does not glitch as the enable E is assigned a value. For example, if when E transitions high, before the high level D propagates through the AND gate 122, if the NE signal transitions low and drives the output of the AND gate 124 low, then both inputs to the OR gate 126 will be low at the same time, causing the output Q to go to a glitched low before the output of the "AND" gate 122 drives the output Q high.
[0046] In one embodiment, an output ready signal is generated that indicates when the Q output is ready to be sampled by the receive logic. In one embodiment, the output ready signal is inverted during the delay provided by delay circuit 105 and is assigned a value once the level of input signal D is transferred to output signal Q. In one embodiment, an enable signal E is inverted in response to the assignment of the output ready signal. In one embodiment, a delay circuit that matches the propagation delay from the enable signal E to the output signal Q is used to delay the enable signal E to generate the output ready signal. In one embodiment, an even number of inverters coupled in series delay the enable signal E to produce the output ready signal.
[0047] Figure 1E An inverter circuit 119 is shown in accordance with one embodiment. Inverter circuit 119 includes two cross-coupled logic gates that generate EO (enable output) and NE. The cross-coupled logic gates are an OR gate 132 with one inverting input to generate EO, and a NAND gate 134 to generate NE. The enable input E is input to both logic gates. When enable E transitions from low to high, OR gate 132 drives EO from low to high. When the rising edge of EO is received at NAND gate 134, the NE output transitions low. The propagation delay through NAND gate 134 ensures that EO and NE are high at the same time, thereby providing overlap time. Similarly, when enable E transitions from high to low, NAND gate 134 drives NE from low to high. When the rising edge of NE is received at OR gate 132, EO is driven from high to low, and the propagation delay through OR gate 132 ensures that EO and NE are high at the same time, thereby providing overlap time. It should be understood that inverter circuit 119 can be replaced with enable inverters 118 in sample circuit 110 with EO routed to AND gate 122 and NE routed to AND gate 124.
[0048] Figure 1F Another glitchless sample circuit 140 is shown in accordance with one embodiment. Glitchless sample circuit 110 can be replaced by glitchless sample circuit 140. OR gate 138 receives intermediate signals generated by AND gates 135, 136, and 137, with AND gate 137 having one inverting input. When both input D and enable E are high, AND gate 135 propagates input D to OR gate 138 to drive output Q high. When enable E is low, AND gate 137 provides a feedback path to maintain output Q high. AND gate 136 provides a feedback path to maintain output Q high when input D is high regardless of the level of enable E, thereby preventing output Q from glitching low during the rising edge of enable E when input D is high. When input D is low and enable E is high, or when enable E is low and output Q is low, output Q is driven low.
[0049] Returning to Figure 1Aof the block diagram 100, it can be desirable to have an "asymmetric" ready signal BRA that is high for longer than it is low in the case of a large delay (more than half a clock) of signal B. In one embodiment, the asymmetric ready signal BRA is low at the end of the clock cycle.
[0050] Figure 1G An asymmetric ready signal generation circuit 145 is shown according to one embodiment. The asymmetric ready signal generation circuit 145 can replace the delay circuit 105 in the block diagram 100. The ready signal BRA is delayed Dl from the rising edge of the clock signal and has a width equal to D2. A first delay circuit 147 provides the delay Dl and a second delay circuit 146 provides the delay D2. An AND gate 148 with one inverting input generates the output BRA. In one embodiment, the first and second delay circuits 147 and 146 are implemented using chains of inverters coupled in series, with the intermediate output in the chain providing the delay Dl and the further delayed output of the chain providing the delay D2.
[0051] Figure 1H An asymmetric ready signal generation circuit is shown according to one embodiment. Figure 1G A timing diagram of the asymmetric ready signal generation circuit of the block diagram 100. The first delay circuit 147 causes the rising edge of BRA to occur Dl time after the rising edge of the clock. The second delay circuit 146 sets the pulse width of the signal BRA to D2. The asymmetric ready signal BRA rises while the clock signal is low and falls before the beginning of the next clock cycle.
[0052] The glitchless sampling circuit 110 is a storage element that can be used to eliminate glitches in a signal by sampling the signal after it changes level and stabilizes. The delay of the ready signal matches the delay of the circuit that generates the signal and enables the sampling to transfer the stable level of the signal to the output in a glitchless manner. The sampling based on the ready signal will produce a glitchless output signal. Any changes in the signal before the signal is sampled are filtered out by the glitchless sampling circuit 110, thereby reducing the number of times a node within the logic that receives the output signal is charged and discharged. As a result, the power consumed by the logic that receives the output signal is reduced compared to directly receiving the changing signal.
[0053] Glitchless multiplexer
[0054] Ideally, to minimize charging and discharging nodes within the combinatorial logic that receives the multiplexer's output, the multiplexer's output will remain constant until all inputs, including the select signal, have reached their final state, and then the output changes exactly once (or remains constant). Glitchless (few-glitch) multiplexing can be achieved by using "bundled" self-timed signals. Specifically, a "ready" signal can be associated with each possible multi-bit logic signal (e.g., data inputs and select). The ready signal is only assigned a value when the associated logic signal has reached a final state and is constant (e.g., glitchless) for a clock cycle.
[0055] For example, the signal "BR" indicates when the signal "B" is ready. As previously described, the ready signal can be generated from the clock signal by a delay circuit having a delay that matches the delay of the combinatorial logic block that generates the signal B. In one embodiment, as shown in Figure 1G The NOR gate is used to provide asymmetric rise and fall delays. In one embodiment, an inverted version of the clock signal is used directly as the ready signal.
[0056] Figure 2A A block diagram of a glitchless 4-to-l multiplexer 200 is shown, according to one embodiment. At a high level, a multiplexer is a logic component that receives a plurality (N) of data input signals and a select signal, and then transmits one of the data input signals to an output signal based on the value of the select signal. For example, a multiplexer can be configured with inputs a, b, and select, and output c, such that c = a when select = 0, and c = b when select = 1. However, as input a, b, and select glitch, the glitches are propagated to the output c. Reducing the glitches of the output c can reduce the power consumed by combinatorial logic that receives the output c as an input.
[0057] The glitchless techniques previously described can be used to implement a glitchless 4-to-l multiplexer 200, where the data input signals are A,..., D. The input ready signals AR,..., DR are delayed to match the respective input signals. The select signal "Sel" selects one of the inputs to sample, and the sampled input is transmitted to the output signal X. Feedback from X prevents glitches from appearing at the output by maintaining the previous level on X when it is assigned a value.
[0058] The multiplexer 200 includes a timing decoder 210 and an output stage 215, each having signal terminals that carry inputs and outputs to and from the circuit. Glitch-free multiplexing can be achieved by associating a ready signal with each input signal (Sel and the data input signals), which can each be multi-bit. The ready signal goes high when the associated input signal reaches a stable state (e.g., a constant or glitch-free value) within the current clock cycle. For example, the signal AR indicates when the input signal A is stable. The ready signals can be generated from the circuit clock signal by a chain of inverters that is delay-matched to the combinational logic block that generates each input signal. In some cases, as previously described in connection with Figure 1G and 1H In other cases, the positive or negative version of the clock can be used directly as the ready signal.
[0059] The timing decoder 210 receives as inputs the select signal (Sel), the select ready signal (SelR), and a plurality of input ready signals (AR,... DR). The timing decoder 210 generates sample enable signals (AS,... DS), each of which is assigned to cause the output stage 215 to sample a corresponding data input signal. The timing decoder 210 also generates a hold signal. In the context of the following description, the sample enable signals can be considered to be multi-bit sample (e.g., one hot) enable signals, where each bit corresponds to a different one of the data input signals and the corresponding ready signal. In one embodiment, the timing of the different sample enable signals can vary based on SelR and the input ready signals.
[0060] In general, each data input signal can be a multi-bit signal, and the corresponding ready signal indicates that all bits of the data input signal are glitch-free. In one embodiment, the sample enable signal for each data input can be used to sample all bits of the multi-bit data input signal, and the output stage 215 is replicated for each bit to generate a multi-bit output signal X.
[0061] The output stage 215 accepts as inputs the data input signals, the sample enable signals, and the hold signal. The output stage 215 generates at least one output signal. The output signal can be fed back to the output stage 215 for use as a hold feedback input. The timing decoder 210 is coupled to the output stage 215 so as to provide the sample enable signals to the output stage 215. The sample enable signals are timed so that the output stage 215 is able to sample one of the data input signals according to the select input after each input signal has reached a final value in each clock cycle. The sample signal is propagated to the at least one output signal, as further shown. Figure 2B
[0062] The combination of the select ready and input ready signals keeps the multiplexer 200 stable (unchanging) until all the select and data input signals are ready, meaning at the final stable value for the clock interval, as indicated by the respective ready signals. The different logic paths and combinational logic that generate each data input and select signal can cause timing differences between the individual signals to reach their own ready state. The individual ready signals can be adjusted based on the logic path that their respective data input signal traverses to provide the appropriate delay. Once the select ready signal and the ready signal corresponding to the selected data input indicate that the data input and select signals have reached a glitch-free and stable state, the multiplexer 200 propagates the selected data input to the output signal.
[0063] Figure 2A The illustrated output stage 215 includes a select gate 206 for each data input signal (A through D), as well as a select gate 205 for the hold signal. The output stage 215 also includes an output gate 212. A "gate" refers to any logic configured to combine one or more inputs according to a logic equation or truth table. For example, an AND gate is logic that combines multiple inputs according to a Boolean AND operation. A gate does not imply a particular arrangement of transistors, and in some cases gates can be implemented using different Boolean operations and combinations appropriate to the implementation.
[0064] Each select gate 205 and 206 can logically act as a two-input AND gate. Each select gate 206 receives a data input signal and a corresponding sample enable signal (A and AS, B and BS,...) as inputs. When the sample enable signal is a logic "1", the output of the select gate matches the input. When the select signal is a logic "0", the output of the select gate is a logic "0". In one embodiment, the timing decoder 210 is configured as a one-hot output generator, such that no more than one sample enable signal is high (logic "1") at any given time.
[0065] The output of the select gate 206 is coupled to an input of the output gate 212. The output gate 212 can act as a multiple-input OR gate. In this way, the sampled signals as previously described at the select gate 206 can be propagated through the output gate 212 to the output signal X. In one embodiment, the output signal X is also fed back as an input to the select gate 205. The select gate 206 is configured to input the feedback to the output gate 212 to hold the output signal X high while the hold signal is high. In this way, the output signal is selected to be the next input to be propagated through the output gate 212 while the hold signal is high, effectively holding that output signal stable for the clock interval. In other words, when the hold signal is asserted (all sample enable signals are low), the output signal X holds its current state.
[0066] The hold signal must remain asserted until one of the sample enable signals (AS...DS) is asserted to prevent glitches at the output when the selected input signal is sampled. Specifically, when the output signal X is asserted (high) and the selected input signal (A...D) is high, the overlap of the hold and sample enable signals can ensure that the output signal X does not glitch low (e.g., when the output signal X is high, it does not transition low and back high within one clock cycle). If the overlap is not used, a low glitch can result due to a race condition when the hold transitions low. A low glitch occurs when the input signal has a high value and the output of the select gate 205 transitions low before the sample enable signal transitions high at the input of the select gate 206, thus causing all of the inputs to the output gate 212 to go low at the same time.
[0067] In terms of transistors and power (the cost of a single NAND gate is typically 4 to 6 transistors), the feedback path is inexpensive. The added resource cost is similar to adding an additional input to the multiplexer 200 and is less expensive than adding a flip-flop. The feedback path, along with the proper sequence of the hold and sample enable signals, can prevent glitches in the output signal X of the multiplexer 200.
[0068] One of ordinary skill in the art will readily recognize that the output stage 215 can include additional or slightly different elements not shown and unnecessary to this description. AND gates are used to describe the select gates 205 and 206, and OR gates are used to describe the output gate 212 because these symbols represent the function of the underlying logic used to generate the output signal based on the inputs. However, other combinations of logic can also be implemented to achieve the effect of the indicated functionality. Five two-input NAND gates can feed into, for example, a single five-input NAND gate. Alternatively, three AND-OR-INVERT (AOI) can feed into a three-input NAND gate. Many circuits can be used to achieve the same logic result when implementing the desired logic result, while affecting the number of transistors used, the delay incurred, and the power consumed. One of ordinary skill in the art will readily recognize that the multiplexer 200 can include additional or slightly different elements not shown and unnecessary to this description.
[0069] The multiplexer 200 provides a reliably stable output signal value even when the data inputs and / or Sel signals change value multiple times and / or at different times. Thus, for the multiplexer 200, power consumption and noise associated with glitches and unnecessary output signal toggling are reduced, as well as the additional work performed by a circuit receiving potentially unstable signals as inputs.
[0070] Figure 2B An example of a multiplexer 200 according to one embodiment is shown in FIG. 2. The multiplexer 200 includes a select gate 205, a select gate 206, and an output gate 212. The select gate 205 has an input coupled to a first data input signal A, a second data input signal B, a third data input signal C, and a fourth data input signal D. The select gate 206 has an input coupled to a first sample enable signal AS, a second sample enable signal BS, a third sample enable signal CS, and a fourth sample enable signal DS. The output gate 212 has an input coupled to the output of the select gate 205 and an input coupled to the output of the select gate 206. The output gate 212 provides an output signal X. Figure 2AA timing diagram for a glitch-free 4-to-l multiplexer 200. The timing diagram depicts the timing of the sample enable signals generated by the timing decoder 210 to achieve glitch-free multiplexing, with arrows indicating causality. The signal SelR is asserted to indicate that the select signal Sel is ready. Prior to the assertion of the signal SelR, one or more glitches can occur to the Sel, changing between values (as shown by the transitions labeled "glitch") until a stable state is reached, after which the SelR is asserted. In this example, the select signal selects the data input signal A, and the input ready signal AR is asserted prior to the assertion of SelR, indicating that the data input signal A is ready. In response to both AR and SelR having been asserted, the timing decoder 210 asserts the sample enable signal AS to sample the data input signal A and propagate the sample value to the output signal X of the multiplexer 200.
[0071] If the AR signal has not been asserted at the time SelR is asserted, the timing decoder 210 waits for the AR signal to be asserted before asserting the AS signal. After a short overlap delay (to) after the assertion of the AS signal, the hold signal is de-asserted (negated). The overlap of the assertion of AS and the hold signal during time to is for the case where the bit of the output signal X and the data input signal A are both asserted. The overlap ensures that the bit of the output signal X does not glitch to the de-asserted state between the time the output of the hold select gate 205 in the output stage 215 becomes de-asserted and the time the output of the A select gate 206 in the output stage 215 is asserted. To reset the multiplexer 200 for the next clock cycle, after both the SelR signal and AR are de-asserted, the hold signal is asserted, and after an overlap delay (tl), the AS signal is de-asserted. The time durations of the delays to and tl can be equal or different. In one embodiment, the time durations of the delays to and tl are approximately the delay of a logic gate. In some implementations, an inverted version (possibly with a delay) of the hold signal can be used as the "output ready" signal XR.
[0072] Figure 2C A block diagram of the timing decoder 210 is shown, according to one embodiment. Figure 2A A block diagram of the timing decoder 210 is shown, according to one embodiment.
[0073] In the ready signal logic 221, each one-hot signal ad,..., dd is defined by selecting the ready signal SelR and the corresponding input ready signal (e.g., AR for input signal A, etc.). Depending on the implementation, the AND gates performing the ready signal logic 221 definition can be implemented within the one-hot decoder 220 or by extending the AND gates in the sample enable signal logic 216 and the hold signal logic 218.
[0074] The resulting decoded ready signals adr,..., ddr output from the ready signal logic 221 are applied to drive a set of reset / set (RS) flip-flops or latches, implemented as AND gates in the depicted embodiment, to generate the input select signals AS,..., DS and hold. The combined logic including the sample enable logic 216 block and the hold signal logic 218 is gated by the ready signal logic 221 gates to generate the sample enable and hold signals.
[0075] When a decoded input ready signal (e.g., adr) is asserted, it sets the output of the latch implemented by the corresponding sample enable signal logic 216 (e.g., signal AS). When any sample enable signal (e.g., AS) is asserted, it resets the flip-flop implemented in the hold signal logic 218, causing the hold signal to fall. If additional overlap is needed to ensure reliable operation, a delay circuit can be inserted in the reset path. However, in many cases, the delay of the RS flip-flop in the sample signal logic 216 itself can be sufficient to avoid the de-assertion glitch when both the held output signal X and the newly selected input signal are asserted.
[0076] For a multi-bit multiplexer (one multiplexer enabled to select a multi-bit input signal as a multi-bit output signal), each sample enable logic 216 block (e.g., AND gate configuration) performs the operation of RS flip-flops or latches for all bits of one multi-bit input signal. Other low gate count configurations for RS / SR flip-flop behavior can also be utilized.
[0077] To reset the timing decoder 210 and prepare it for the next clock cycle, the hold RS flip-flop implemented in the hold signal logic 218 is set when the enable signal such as adr goes low. Since one enable signal at a time is asserted, the condition is detected in the hold signal logic 218 by an OR NOT gate (e.g., AND gate with inverted inputs) configured to detect all de-asserted enable signals. The hold signal becomes asserted and then resets the selected RS flip-flop in the sample enable logic circuit 216. If an extended overlap period is needed, a delay can be inserted in the hold sample enable reset path.
[0078] Figure 2DA flowchart illustrating a method 225 for generating glitch-free multiplexer output signals according to one embodiment is shown. The method 225 can be performed by logic or custom circuitry. For example, the method 225 can be performed by a GPU, CPU, or any processor capable of generating glitch-free signals. Moreover, one of ordinary skill in the art will appreciate that any system performing the method 225 is within the scope and spirit of embodiments of the disclosure.
[0079] At step 230, a selection ready signal is received at the timing decoder 210. The selection ready signal is negated until the selection signal generated by the combinational logic is invariant (glitch-free) and is assigned after the selection signal is invariant. At step 235, the timing decoder 210 generates at least one sample enable signal from the selection signal, where the at least one sample enable signal corresponds to a set of data input signals. The at least one sample enable signal is negated while the selection ready signal is negated and is assigned when both the selection ready signal and the corresponding data ready signal are assigned.
[0080] In one embodiment, the selection signal comprises a multi-bit signal, and each bit in the selection signal is associated with a different one of the set of data input signals, and only one bit is assigned at a time. For example, the selection signal can be one-hot encoded by the one-hot decoder 220 to produce adr,..., ddr. In one embodiment, each bit in the selection signal is used to sample the associated data input signal. For example, the multi-bit selection signal is used to generate sample enable signals each associated with a different one of the data input signals. In one embodiment, the selection signal can be binary encoded, where each binary pattern selects one of the data input signals.
[0081] In one embodiment, the timing decoder 210 receives a set of ready signals, each ready signal in the set associated with a different one of the set of data input signals. Moreover, each ready signal is negated until the associated data input signal is invariant, and is assigned after the associated data input signal is invariant.
[0082] In one embodiment, the timing decoder 210 is further configured to generate a set of enable signals, e.g., adr,..., ddr, where each enable signal in the set is associated with a different one of the data input signals in the set. Moreover, each of the enable signals is negated when the associated ready signal is negated, and is assigned in response to the assignment of the associated ready signal when the associated data input signal is assigned.
[0083] At step 240, the timing decoder 210 generates a hold signal that is valued when the at least one sample enable signal is negated and is negated in response to the at least one sample enable signal being valued. In one embodiment, the timing decoder 210 is configured to value the hold signal when the hold signal is valued and the at least one sample enable signal is negated. In one embodiment, both the hold signal and the at least one sample enable signal are valued for a first duration (e.g., tO and / or ti). In one embodiment, the hold signal is inverted to produce an output ready signal associated with the output signal.
[0084] At step 245, the sample circuit, e.g., the output stage 215, holds the output signal constant when the hold signal is valued. At step 250, one of the data input signals is sampled in accordance with the at least one sample enable signal while the hold signal is negated to transfer the value of the sampled data input signal to the output signal.
[0085] In one embodiment, the set of input data signals includes three input data signals and the select signal is configured to select one of the three input data signals to produce the output signal. In one embodiment, the set of input data signals includes four input data signals and the select signal is configured to select two of the four input data signals to produce the output signal and an additional output signal.
[0086] A particular embodiment of a glitchless multiplexer is a 4-2 multiplexer that simultaneously selects two of four inputs aO to a3 to propagate to two output signals qO and q1. In the case of a hold order, as shown in Table 1, there are six possibilities encoded by a three-bit select signal for this multiplexer:
[0087] se1 q0 q1 0 a0 a1 1 a0 a2 2 a0 a3 3 a1 a2 4 a1 a3 5 a2 a3
[0088] Table 1
[0089] The data path of the 4-2 multiplexer can be implemented as two 3-1 multiplexers with corresponding decoded input ready signals derived from the 3-6 decoding scheme as shown in the following equations:
[0090] For the 3-1 multiplexer producing qO:
[0091]
[0092]
[0093]
[0094] For the 3-1 multiplexer producing q1:
[0095]
[0096] a12dr = (s1 V s3) A selr A a2r
[0097] a13dr = (s2 V s4 V s5) A selr A a3r
[0098] Here, sx, where x takes on values 0-5, indicates that the select signal is named x; auvdr is the decoded input ready signal that gates the input signal av to the output signal qu. These signals are used to drive the four RS latches of each 3-1 multiplexer of the 4-2 multiplexer. For the q0 multiplexer, the three RS latches produce signals a00s, a01s, and a02s, thus enabling paths from inputs a0, a1, and a2 to q0, respectively. The final RS latch holds q0. A similar set of four RS latches controls the q1 multiplexer. For example, the a00s RS latch is formulated as:
[0099]
[0100] Figure 2E An expanded ready signal generation circuit 255 is shown in accordance with one embodiment. The select ready signal is expanded using RS flip-flops including AND gates 256 and OR gates 257 to ensure overlap with the input ready signal. In order for the glitchless multiplexer to function properly, the select ready (SelR) signal and the input ready signal (XR) overlap for sufficient time to set the selected RS flip-flop in the sample enable signal logic circuit 216 of the timing decoder 210. The delay of the input ready signal is much greater than the delay of the select ready signal, and the overlap can not occur. This problem can be solved by expanding the select ready signal using additional RS flip-flops. The expanded signal (SelXR) is set by setting SelR high and is cleared by setting the corresponding sample enable signal (XS) high. The timing decoder 210 can be modified to include additional RS flip-flops for each input data signal. Specifically, the expanded ready signal generation circuit 255 can be inserted to replace the SelR input to each AND gate within the input ready signal logic 221.
[0101] Figure 2F A timing diagram of the expanded ready signal generation circuit 255 of Figure 2E is shown in accordance with one embodiment. The XS signal is generated using the expanded SelR signal SelXR because there is insufficient overlap in time when SelR and XR are asserted simultaneously. SelXR is reset in response to XS being asserted. SelXR is set in response to SelR being asserted.
[0102] In some cases, the duration of the input ready signal is longer than required and can unnecessarily delay the reset of the glitchless N-l multiplexer 200 to prepare it for the next set of inputs. A fast return circuit can be implemented on the input ready signal. The fast return circuit returns the input ready signal to the de-asserted state immediately upon detection of the output ready signal.
[0103] Figure 2G A fast return circuit 260 according to one embodiment is shown. The logical formula for this circuit is:
[0104] ARq = AR Λ S'
[0105] S = (AR Λ S) V (ARq Λ OR)
[0106] Here AR is the input data ready signal, ARq is the fast return version of the input data ready signal, OR is the delayed version of the output ready signal - AS, S is the state variable, and S' is the complement of S.
[0107] Figure 2H A timing diagram for the fast return circuit 260 of Figure 2G is shown according to one embodiment. When the input ready signal AR rises, the signal ARq rises, but as soon as the output ready signal OR goes high, the signal ARq falls. The signal AR goes high and stays high for a period of time. To enable the glitchless N-l multiplexer 200 to reset faster, the fast return circuit 260 generates the signal ARq, which returns to zero with a low delay after OR goes high. The fast return behavior can be enabled by setting the state signal S, which is generated by an RS flip-flop. The RS flip-flop is set by OR going high when ARq is high. The RS flip-flop keeps S high until AR goes low to reset it.
[0108] In other cases, the output ready signal (OR) can be extended until all of the outputs it is combined with are also ready, and the subsequent circuit accepts the combination. This event can be signaled by an ack signal, which can also be the output ready signal of the subsequent circuit. In these cases, using the complement hold signal as OR can not be long enough to meet the timing constraints of the implementation. An extended output ready signal can be generated using an RS flip-flop, which is set when held de-asserted and reset when ack is asserted:
[0109] OR = hold' V (ack' Λ OR)
[0110] In this case, the fast return circuit 260 can be modified to: Figure 2CThe hold signal logic 218 is to prevent the hold from becoming de-asserted again until the ack is asserted.
[0111] The glitchless technique can be used to reduce power consumption of multiplexer logic. Changes to the output signal are prevented until propagation of the glitchless value to the output signal results in the output signal transitioning only once (or remaining constant) per clock cycle. When a glitch occurs in the output signal, the logic receiving the signal can respond by charging and / or discharging nodes within the logic and consuming power. Providing a glitchless signal can reduce the time for which nodes are charged and / or discharged, thereby reducing the power consumed by the logic.
[0112] The glitchless technique associates a ready signal with each input of a glitchless multiplexer to control sampling of the selected input after the input is stable and glitchless. Embodiments of the sampling element include a transparent latch or a multiplexer with feedback. Each delay of the ready signal is matched to the associated input and can be extended, appear asymmetric, or quickly reset ready. Within the multiplexer, a hold signal is generated to hold the output signal stable through a feedback path until the sampled signal value is propagated. The hold signal can also be used to generate an output ready signal.
[0113] The glitchless technique can be implemented within any circuitry including combinational logic and sequential logic. For example, the glitchless technique can be used within one or more logic blocks within a processor and used for inputs or outputs of the processor. In particular, a glitchless multiplexer can be used to select only non-zero activations and / or weights for a convolution operation. In one embodiment, the glitchless technique can be implemented within a parallel processing architecture, as further described herein.
[0114] Parallel processing architecture
[0115] Figure 3 A parallel processing unit (PPU) 300 according to one embodiment is shown. In one embodiment, the PPU 300 is a multi-threaded processor implemented on one or more integrated circuit devices. The PPU 300 is a latency hiding architecture designed to process multiple threads in parallel. A thread (e.g., an execution thread) is an instance of a set of instructions configured to be executed by the PPU 300. In one embodiment, the PPU 300 is a graphics processing unit (GPU) configured to implement a graphics rendering pipeline for processing three-dimensional (3D) graphics data in order to generate two-dimensional (2D) image data for display on a display device such as a liquid crystal display (LCD). In other embodiments, the PPU 300 can be used to perform general purpose computations. Although one exemplary parallel processor is provided herein for illustrative purposes, it is strongly
[0116] One or more PPUs 300 are configured to accelerate high-performance computing ("HPC"), data center, and machine learning applications. In at least one embodiment, the PPU 300 is configured to accelerate deep learning systems and applications, including autonomous vehicle platforms, deep learning, high-precision speech, image, and text recognition systems, intelligent video analysis, molecular simulation, drug discovery, disease diagnosis, weather forecasting, big data analysis, astronomy, molecular dynamics simulation, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations.
[0117] like Figure 3 As shown, the PPU 300 includes, but is not limited to, an input / output ("I / O") unit 305, a front-end unit 315, a scheduler unit 320, a work distribution unit 325, a hub 330, a crossbar switch ("Xbar") 370, one or more general purpose processing clusters (GPCs) 350, and one or more memory partitioning units 380. In at least one embodiment, the PPU 300 is interconnected to a host processor or other PPUs 300 via one or more NV Links 310. The PPU 300 is connected to a host processor or other peripheral devices via an interconnect 302. The PPU 300 is connected to a local memory 304 comprising a plurality of memory devices. In at least one embodiment, the local memory 304 comprises a plurality of dynamic random access memory (DRAM) devices. The DRAM devices are configured as a high-bandwidth memory (HBM) subsystem, with multiple DRAM dies stacked within each device.
[0118] The NV Link 310 enables the system to be scaled and includes one or more PPUs 300 combined with one or more CPUs, supporting cache coherence between the PPU 3300 and the CPU and CPU mastering. The NV Link 310 transmits data and / or commands to other units of the PPU 300, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown), through the hub 330. Figure 5B The NV link 310 is described in more detail.
[0119] The I / O unit 305 is configured to send and receive communications (e.g., commands, data) from a host processor (not shown) over the interconnect 302. The I / O unit 305 communicates with the host processor directly over the interconnect 302, or through one or more intermediary devices such as a memory hub. In one embodiment, the I / O unit 305 can communicate with one or more other processors (such as one or more PPUs 300) via the interconnect 302. In at least one embodiment, the I / O unit 305 implements a Peripheral Component Interconnect Express (“PCIe”) interface for communications over a PCIe bus, and the interconnect 302 is a PCIe bus. In at least one embodiment, the I / O unit 305 implements another known interface for communications with external devices.
[0120] The I / O unit 305 decodes packets received via the interconnect 302. In one embodiment, the packets represent commands that configure the PPU 300 to perform various operations. The I / O unit 305 sends the decoded commands to various other units of the PPU 3300 as specified by the commands. For example, some commands are sent to the hub 330 or other units of the PPU 3300 such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). In other words, the I / O unit 305 is configured to route communications between various logical units of the PPU 300.
[0121] In one embodiment, a program executed by the host processor encodes a command stream in a buffer that provides a workload to the PPU 300 for processing. The workload includes a plurality of instructions and data to be processed by those instructions. The buffer is a region in memory that is accessible (e.g., read / write) by both the host processor and the PPU 300. The I / O unit 305 may, for example, be configured to access the buffer in system memory connected to the interconnect 302 via memory requests transmitted over the interconnect 302. In one embodiment, the host processor writes the command stream to the buffer and then sends a pointer to the beginning of the command stream to the PPU 300. The front-end unit 315 receives the one or more command stream pointers. The front-end unit 315 manages the one or more streams, reads commands from the stream, and forwards the commands to various units of the PPU 300.
[0122] The front-end unit 315 is coupled to a scheduler unit 320, which configures the various GPCs 350 to process tasks defined by one or more workgroups. The scheduler unit 320 is configured to track state information related to various tasks managed by the scheduler unit 320. The state information can indicate which task is assigned to which GPC 350, whether the task is active or inactive, priority of the task associated with the task, and so on. In at least one embodiment, the scheduler unit 320 manages a plurality of tasks for execution on the one or more GPCs 350.
[0123] In at least one embodiment, the scheduler unit 320 is coupled to a work distribution unit 325, which is configured to dispatch tasks for execution on GPCs 350. In at least one embodiment, the work distribution unit 325 tracks a number of tasks to be processed from the scheduler unit 320. In one embodiment, the work distribution unit 325 manages a pending task pool and an active task pool for each GPC 350. The pending task pool can include a number of slots (e.g., 32 slots) that contain tasks that are assigned to be processed by a particular GPC 350. The active task pool can include a number of slots (e.g., 4 slots) for tasks that are actively being processed by a GPC 350. When a task is completed in a GPC 350, the task is evicted from the active task pool for the GPC 350 and another task is selected from the pending task pool and scheduled for execution on the GPC 350. In at least one embodiment, if an active task is idle, for example, while waiting for a data dependency to resolve, the active task is evicted from the GPC 350 and returned to the pending task pool while another task is selected from the pending task pool and scheduled for execution on the GPC 350.
[0124] The work distribution unit 325 communicates with one or more GPCs 350 via an XBar 370. The XBar 370 is an interconnect network that couples many units of the PPU 300 to other units of the PPU 300. For example, the XBar 370 can be configured to couple the work distribution unit 325 to a particular GPC 350. Although not explicitly shown, other units of the PPU 300 can also be connected to the XBar 370 via the hub 330.
[0125] Tasks are managed by the scheduler unit 320 and assigned to GPCs 350 by the work distribution unit 325. The GPCs 350 are configured to process tasks and generate results. The results can be consumed by other tasks in the GPC 350, routed to different GPCs 350 via the XBar 370, or stored in the memory 304. The results can be written to the memory 304 via the memory partition unit 380, which implements a memory interface for writing data to or reading data from the memory 304. The results can be transferred to another PPU 300 or CPU via the NV link 310. In one embodiment, the PPU 300 includes, but is not limited to, U memory partition units 380, which is equal to the number of separate and distinct storage devices 304 coupled to the PPU 300. Figure 4B The memory partition unit 380 is described in more detail.
[0126] In one embodiment, the host processor executes a driver core that implements an application programming interface (API) that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 300. In one embodiment, multiple computing applications are executed simultaneously by the PPU 300, and the PPU 300 provides isolation, quality of service ("QoS"), and independent address spaces for the multiple computing applications. The application generates instructions (e.g., in the form of API calls) that cause the driver core to generate one or more tasks for execution by the PPU 300. The driver core outputs the tasks to one or more streams processed by the PPU 300. Each task includes one or more related groups of threads, which may be referred to as warps. In one embodiment, a warp includes 32 related threads that may be executed in parallel. Collaborating threads may refer to multiple threads that include instructions for performing tasks and exchanging data through shared memory. In combination Figure 5A Describes threads and cooperative threads in more detail.
[0127] Figure 4A According to one embodiment, Figure 3 PPU300 GPC350. Figure 4A As shown, each GPC 350 includes multiple hardware units for processing tasks. In one embodiment, each GPC 350 includes a pipeline manager 410, a pre-raster operation unit PROP415, a raster engine 410, a work distribution crossbar switch WDX480, a memory management unit MMU490, and one or more data processing clusters DPC420. It should be understood that Figure 4A The GPC 350 may include other hardware units instead of or in lieu of Figure 4A The unit shown in .
[0128] In one embodiment, the operation of GPC 350 is controlled by pipeline manager 410. Pipeline manager 410 manages configuration of one or more DPCs 420 to process tasks allocated to GPC 350. In one embodiment, pipeline manager 410 configures at least one of one or more DPCs 420 to implement at least a portion of a graphics rendering pipeline. For example, a DPC 420 is configured to execute vertex shader programs as in a programmable stream processor SM 440. Pipeline manager 410 can also configure DPCs 420 to process data packets received from a work distribution unit packed to an appropriate logic unit within GPC 350. For example, some data packets can be routed to fixed function hardware units in PROP 415 and / or raster engine 425, while other data packets can be routed to DPCs 420 for processing by work item execution resources in MIMD array 435 or SM 440. In at least one embodiment, pipeline manager 410 can configure at least one of one or more DPCs 420 to implement a neural network model and / or compute pipeline.
[0129] PROP unit 415 is configured to route data generated by raster engine 425 and DPCs 420 to Figure 4B ROP units, which are described in greater detail below, are configured to perform raster operations and / or to perform other non-shadow computation operations. In some embodiments, PROP unit 415 is configured to perform optimizations for color blending, organize pixel data, perform address translations, and so forth. Figure 4B ROP units, which are described in greater detail below, are configured to perform raster operations and / or to perform other non-shadow computation operations. In some embodiments, PROP unit 415 is configured to perform optimizations for color blending, organize pixel data, perform address translations, and so forth.
[0130] Raster engine 425 includes a number of fixed function hardware units to perform various raster operations. In one embodiment, raster engine 425 includes a setup engine, a coarse raster engine, a cull engine, a clip engine, a fine raster engine, and a tile aggregation engine. The setup engine receives transformed vertices and generates a plane equation associated with the geometric primitive defined by the vertices. The plane equation is transmitted to the coarse raster engine to generate coverage information (e.g., x, y coverage masks for tiles) of the primitive. The output of the coarse raster engine is transmitted to the cull engine where fragments associated with primitives that fail a z-test are culled. The output of the cull engine is transmitted to the clip engine where fragments that fall outside the view volume are clipped. The output of the clip engine is transmitted to the fine raster engine to generate attributes for the pixel fragments based on the plane equation generated by the setup engine. The output of raster engine 425 includes fragments that are processed, e.g., by a fragment shader implemented within DPCs 420.
[0131] Each DPC 420 included in the GPC 350 includes an M-pipeline controller MPC 430, a primitive engine 435, and one or more SMs 440. In at least one embodiment, the MPC 430 controls the operation of the DPC 420, routing packets received from the pipeline manager 410 to appropriate units within the DPC 420. For example, packets associated with vertices are routed to the primitive engine 435, which is configured to retrieve vertex attributes associated with the vertices from memory; conversely, packets associated with shader programs may be sent to the SM 440.
[0132] SM 440 includes a programmable streaming processor configured to process tasks represented by multiple threads. Each SM 440 is multi-threaded and configured to execute multiple threads (e.g., 32 threads) from a particular thread group simultaneously. In one embodiment, SM 440 implements a SIMD ("single instruction, multiple data") architecture, in which each thread in a group of threads (e.g., a warp) is configured to process a different data set based on a common instruction set. All threads in a thread group execute the same instructions. In one embodiment, SM 440 implements a SIMT ("single instruction, multiple thread") architecture, in which each thread in a group of threads is configured to process a different data set based on a common instruction set, but in which individual threads in a thread group are allowed to diverge during execution. In at least one embodiment, a program counter, call stack, and execution state are maintained for each warp, thereby enabling concurrency between warps and serial execution within a warp when threads in the warp diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, thereby enabling equal concurrency between all threads within a warp and between warps. When maintaining execution state for each individual thread, threads that execute common instructions can be converged and executed in parallel to improve efficiency. Figure 5A SM 440 is described in more detail.
[0133] MMU 490 provides an interface between GPC 350 and memory partition unit 380, and MMU 490 provides virtual to physical address translation, memory protection, and arbitration of memory requests. In one embodiment, MMU 490 provides one or more translation lookaside buffers ("TLBs") for performing translation of virtual addresses to physical addresses in memory 304.
[0134] Figure 4B According to one embodiment, Figure 3 The memory partition unit 380 of the PPU 300 is as follows. Figure 4BAs shown, memory partition unit 380 includes, without limitation, a raster operations (ROP) unit 450; a level two (L2) cache 460; a memory interface 470. Memory interface 470 is coupled to memory. Memory interface 470 can implement a 32, 64, 128, 1024-bit data bus, etc. for high-speed data transfer. In one embodiment, PPU 300 includes U memory interfaces 470. One memory interface 470 per pair of partition units 350, with each pair of memory partition units 380 connected to a respective memory device. For example, PPU 300 can be connected to Y memory devices such as high bandwidth memory stacks or graphics double data rate, version 5, synchronous dynamic random access memory or other introspective memory.
[0135] In one embodiment, memory interface 470 implements an HBM2 memory interface and Y is equal to half of U. In one embodiment, HBM2 memory stacks are located on the same physical package as PPU 300, saving significant power and area compared to a traditional GDDR5 SDRAM system. In one embodiment, each HBM2 stack includes four memory dies and Y = 4, while the HBM2 stack includes two 128-bit channels per die for a total of 8 channels and a data bus width of 1024 bits.
[0136] In one embodiment, memory 304 supports single error correction double error detection (SECDED) error correcting code (ECC) to protect data. ECC can provide higher reliability for compute applications that are sensitive to data corruption. Reliability is especially important in large scale cluster computing environments where PPU 300 processes very large data sets and / or long running applications.
[0137] In one embodiment, PPU 300 implements a multi-level memory hierarchy. In one embodiment, memory partition unit 380 supports a unified memory to provide a single unified virtual address space for CPU and PPU 300 memory, enabling data sharing between virtual memory systems. In one embodiment, PPU 300 tracks frequency of accesses to memory locations, and moves frequently accessed pages of memory to higher tiers of the memory hierarchy. In one embodiment, NV link 310 supports address translation services, which allow PPU 300 to directly access CPU's page tables and provide full access to CPU memory by PPU 300.
[0138] In one embodiment, the copy engine transfers data between multiple PPUs 300 or between a PPU 300 and a CPU. The copy engine can generate a page fault for an address that is not mapped into a page table, and the memory partition unit 380 then services the page fault, mapping the address into a page table, after which the copy engine performs the transfer. In a conventional system, fixed (i.e., non-paged) memory is allocated for multiple copy engine operations between multiple processors, substantially reducing the available memory. Because of the hardware page fault, an address can be passed to the copy engine without worrying about whether the memory page is resident, and the copy process is transparent.
[0139] Data from memory 304 or other system memory is fetched by the memory partition unit 380 and stored in an L2 cache 460, which is on-chip and shared among various GPCs. As shown, each memory partition unit 380 includes at least a portion of the L2 cache 460 that is associated with a corresponding memory 304. Lower level caches are then implemented within individual units within the GPC 350. For example, each SM 440 can implement a level one (LI) cache, where the LI cache is private to the particular SM 440. Data from the L2 cache 460 can be fetched and stored in each LI cache for processing in the functional units of the SM 440. The L2 cache 460 is coupled to the memory interface 470 and the XBar 370.
[0140] The ROP unit 450 performs graphics raster operations, such as color compression, pixel blending, and the like. The ROP unit 450 performs depth tests in conjunction with the raster engine 425, receiving depth for sample positions associated with pixel fragments from the culling engine of the raster engine 425. For a sample position associated with a fragment, a depth test is performed in the depth buffer for the respective depth. If the fragment passes the depth test for the sample position, the ROP unit 450 updates the depth buffer and sends the results of the depth test to the raster engine 425. It will be appreciated that the number of memory partition units 380 can be different from the number of GPCs 350, and thus each ROP unit 450 can be coupled to each GPC. The ROP unit 450 tracks the packets received from different GPCs and determines to which GPC 350 to route the results generated by the ROP unit 450 through the XBar 370. Although the ROP unit 450 is included within the memory partition unit 380 in Figure 4B In other embodiments, the ROP unit 450 can be external to the memory partition unit 380. For example, the ROP unit 450 can reside in the GPC 350 or another unit.
[0141] Figure 5A FIG. 1 illustrates a system 100, in accordance with one embodiment, in which a streaming multiprocessor (“SM”) 440 is shown in accordance with one embodiment. Figure 4A Figure 5A As shown, the SM 440 includes, without limitation, an instruction cache 505; one or more scheduler units 510; a register file 520; one or more processing cores 550; one or more special function units SFU 552; one or more load / store units LSU 554; an interconnect network 580; shared memory / L1 cache 570.
[0142] As described above, the work distribution unit 325 allocates tasks to be performed with the GPCs 350 of the PPU 300. The tasks are allocated to a particular DPC 420 within a GPC 350, which can be assigned to the SM 440 if the task is associated with a shader program. The scheduler unit(s) 510 receive tasks from the work distribution unit 325 and manage dispatch of instructions to the one or more thread blocks assigned to the SM 440. The scheduler unit(s) 510 can schedule the thread blocks to be executed to be implemented as parallel thread of execution, where each thread block is allocated at least one thread of execution. In one embodiment, there are 32 threads of execution for each thread block. The scheduler unit(s) 510 can manage a plurality of different thread blocks, allocating thread blocks to different ones of the cores 550, and scheduling instructions of a thread block to be executed on different ones of the functional units (e.g., the cores 550, the SFUs 552, and the LSUs 554) in each clock cycle.
[0143] Cooperative groups are a programming model for organizing groups of communicating threads that allow developers to express the granularity at which threads communicate, enabling richer, more efficient parallel decomposition. Cooperative launch APIs support synchronization between thread blocks to execute parallel algorithms. Conventional programming models provide a single, simple construct for synchronizing cooperating threads: a barrier across all threads of a thread block (e.g., the syncthreads() function). However, programmers often wish to define thread groups at a smaller than thread block granularity, and synchronize within defined groups, to achieve higher performance, design flexibility, and software reuse in the form of collective group-level functional interfaces.
[0144] Cooperative groups enable programmers to explicitly define thread groups at sub-block (e.g., down to individual threads) and multi-block granularities, and perform collective operations, such as synchronizing threads in a cooperative group. The programming model supports clean composition across software boundaries, so library and utility functions can safely synchronize within their local context without having to make assumptions about convergence. Cooperative group primitives enable new patterns of cooperative parallelism, including producer-consumer parallelism, opportunistic parallelism, and global synchronization across a grid of thread blocks.
[0145] Dispatch units 515 are configured to send instructions to one or more functional units. In this embodiment, scheduler units 510 include two dispatch units 515, which enable two different instructions from the same thread bundle to be dispatched within each clock cycle. In alternative embodiments, each scheduler unit 510 can include a single dispatch unit 515 or additional dispatch units 515.
[0146] Each SM 440 includes a register file 520 that provides a set of registers for the functional units of the SM 440. In one embodiment, the register file 520 is divided into a number of separate register files, one for each functional unit. In another embodiment, the register file 520 is time-multiplexed so that each functional unit is allocated a portion of the register file 520 at various intervals. The register file 520 provides temporary storage for operands that are used by the data paths connected to the functional units.
[0147] Each SM 440 includes L processing cores 550. In one embodiment the SM 440 includes a large number (e.g., 128, etc.) of different processing cores 550. Each core 550 can include a full pipe lined single precision, double precision, and / or mixed precision processing unit including floating point arithmetic logic and integer arithmetic logic. In one embodiment the floating point arithmetic logic implements IEEE 754-2008 standard for floating point arithmetic. In one embodiment, the cores 550 include 64 single precision (32-bit) floating point cores, 64 integer cores, 32 double precision (64-bit) floating point cores, and 8 tensor cores.
[0148] Tensor cores are configured to perform matrix operations and in one embodiment one or more tensor cores are included in the cores 550. In particular, the tensor cores are configured to perform deep learning matrix arithmetic, such as for convolutional neural network training and inferencing. In one embodiment, each tensor core operates on a 4x4 matrix and performs a matrix multiply and accumulate operation D = AxB + C, where A, B, C, and D are 4x4 matrices.
[0149] In one embodiment, the matrix multiply inputs A and B are 16-bit floating point matrices, while the accumulation matrices C and D can be 16-bit floating point or 32-bit floating point matrices. The tensor core operates on 16-bit floating point input data with 32-bit floating point accumulation. The 16-bit floating point multiplication requires 64 operations and produces a full precision product, which is then accumulated with other intermediate products using 32-bit floating point addition for a 4x4x4 matrix multiplication. In effect, the tensor core is used to perform larger two-dimensional or higher dimensional matrix operations that are composed of these smaller elements. APIs such as CUDA 9 C++ API expose specialized matrix load, matrix multiply and accumulate, and matrix store operations to efficiently use the tensor core in CUDA-C++ programs. At the CUDA level, the warp level interface assumes 16x16 sized matrices across all 32 threads of a warp.
[0150] Each SM 440 also includes M SFUs 552 that perform special functions, such as certain mathematical functions (e.g., reciprocal square root). In one embodiment, SFU 552 can include a tree traversal unit configured to traverse a hierarchical tree data structure. In one embodiment, SFU 552 can include a texture unit configured to perform texture lookups. In one embodiment, the texture unit is configured to fetch texture data (e.g., a 2D array of texture pixels) from memory 304 and perform sampling operations to produce sampled texture values for use in a shader program executed by SM 440. The texture values can be fetched from a texture map stored in shared memory / L1 cache 470. The texture unit uses texture maps (e.g., texture maps having different levels of detail) to perform texture operations such as filtering operations. In one embodiment, each SM 340 includes two texture units.
[0151] Each SM 440 also includes N LSUs 554 that implement load and store operations between shared memory / L1 cache 570 and register file 520. Each SM 440 includes an interconnect network 580 of conductors that connect each of the functional units to register file 520 and LSUs 554 to register file 520, shared memory / L1 cache 570. In one embodiment, interconnect network 580 is a cross-bar switch that is configured to connect any of the functional units to any of the registers in register file 520. Also, LSUs 554 are connected to register files and memory locations in shared memory / L1 cache 570 via interconnect network 580.
[0152] Shared memory / L1 cache 570 is an array of on-chip memory that allows data storage and communication between SM 440 and primitive engines 435, as well as between threads within SM 440. L1 cache 570 includes 128KB of storage capacity and is located in the path from SM 440 to memory partition unit 380. Shared memory / L1 cache 570 can be used to cache reads and writes. One or more of shared memory / L1 cache 570, L2 cache 460, and memory 304 are backing stores.
[0153] Combining data cache and shared memory functionality into a single memory block provides the best overall performance for both types of memory access. This capacity can be used as a cache for programs that do not use shared memory. For example, if shared memory is configured to use half of its capacity, texture and load / store operations can use the remaining capacity. This integration within shared memory / L1 cache 570 enables shared memory / L1 cache 570 to serve as a high-throughput pipeline for streaming data, while providing high-bandwidth and low-latency access to frequently reused data.
[0154] When configuring for general parallel computing, a simpler configuration can be used compared to graphics processing. Specifically, bypassing Figure 3 , creating a simpler programming model. In a general-purpose parallel computing configuration, the work distribution unit 325 allocates and dispatches thread blocks directly to the DPC 420. Threads in a block execute the same program using unique thread IDs in computations to ensure that each thread produces unique results. SMs 440 are used to execute the program and perform computations. Shared memory / L1 cache 570 is used for communication between threads. LSUs 554 are used to read and write global memory via shared memory / L1 cache 570 and memory partitioning unit 380. SMs 440 configured for general-purpose parallel computing can also write commands that the scheduler unit 320 can use to initiate new work on the DPC 420.
[0155] The PPU 300 may be included in a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smartphone (e.g., a wireless, handheld device), a personal digital assistant (PDA), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, etc. In one embodiment, the PPU 300 is embodied on a single semiconductor substrate. In another embodiment, the PPU 300 is included in a system-on-chip (SoC) along with one or more other devices (such as an additional PPU 300, memory 304, a reduced instruction set computer (RISC) CPU, memory, a memory management unit (MMU), a digital-to-analog converter (DAC), etc.).
[0156] In one embodiment, PPU 300 can be included on a graphics card that includes one or more memory devices. The graphics card can be configured to interface with a PCIe slot on a motherboard of a desktop computer. In another embodiment, PPU 300 can be an integrated graphics processing unit (iGPU) or parallel processor that is integrated in a chipset of a motherboard.
[0157] Exemplary computing system
[0158] As developers expose and utilize more parallelism in applications such as artificial intelligence computing, systems with multiple GPUs and CPUs are used in various industries. High performance GPU accelerated systems with tens of thousands of compute nodes have been deployed in data centers, research facilities, and supercomputers to solve increasingly larger problems. As the number of processing devices in high performance systems increases, communication and data transfer mechanisms need to scale to support the increased bandwidth.
[0159] Figure 5B is a conceptual diagram of a processing system 500 implemented using Figure 3 PPU 300 according to one embodiment. Exemplary system 565 can be configured to implement methods 115 and / or 225 as shown in Figure 1C Figure 2D Processing system 500 includes CPU 530, switch 510, and multiple PPUs 300, along with respective memories 304. NVLinks 310 provide high speed communication links between each PPU 300. Although a particular number of NVLinks 310 and interconnects 302 connections are shown as in Figure 5B
[0160] In another embodiment (not shown), NVLink 310 provides one or more high-speed communication links between each PPU 300 and CPU 530, and switch 510 interfaces between interconnect 302 and each PPU 300. PPU 300, memory 304, and interconnect 302 can be located on a single semiconductor platform to form a parallel processing module 525. In yet another embodiment (not shown), interconnect 302 provides one or more communication links between each PPU 300 and 300, and CPU 530 and switch 510 interface between each PPU 300 using NVLink 310 to provide one or more high-speed communication links between the PPUs 300. In another embodiment (not shown), NVLink 310 provides one or more high-speed communication links between PPU 300 and CPU 530 through switch 510. In still another embodiment (not shown), interconnect 302 provides one or more communication links directly between each PPU 300. One or more of the NVLink 310 high-speed communication links can be implemented as a physical NVLink interconnect or an on-chip or off-chip interconnect using the same protocol as NVLink 310.
[0161] In the context of this specification, a single semiconductor platform can refer to a sole unitary semiconductor-based integrated circuit that is fabricated in a single chip or die. It should be noted that the term single semiconductor platform can also refer to multi-chip modules with increased connectivity which simulate on-chip operation and make substantial improvements over utilizing a conventional bus implementation to implement the multiple individual components.
[0162] In one embodiment, the signaling rate of each NVLink 310 is 20 to 25 gigabytes per second, and each PPU 300 includes six NVLink 310 interfaces (as shown in FIG. 1, each PPU 300 includes five NVLink 310 interfaces). Each NVLink 310 provides a data transfer rate of 25 gigabytes per second in each direction, with six links providing 300 gigabytes per second. Figure 5B As shown in FIG. 1, each PPU 300 includes five NVLink 310 interfaces). Each NVLink 310 provides a data transfer rate of 25 gigabytes per second in each direction, with six links providing 300 gigabytes per second. Figure 5B As shown in FIG. 1, each PPU 300 includes five NVLink 310 interfaces). Each NVLink 310 provides a data transfer rate of 25 gigabytes per second in each direction, with six links providing 300 gigabytes per second.
[0163] In one embodiment, the NVLink 310 allows direct load / store / atomic access from the CPU 530 to each PPU 300 memory 304. In one embodiment, the NVLink 310 supports coherency operations allowing data read from the memory 304 to be stored in the cache hierarchy of the CPU 530, which reduces cache access latency for the CPU 530. In one embodiment, the NVLink 310 includes support for address translation services (ATS) allowing the PPU 300 to directly access page tables within the CPU 530. One or more of the NVLink 310 can also be configured to operate in a low power mode.
[0164] Figure 5C An exemplary system 565 in which various previous embodiments can be implemented is shown. The exemplary system 565 can be configured to implement the Figure 1C the method 115 shown and / or Figure 2D the method 225 shown.
[0165] As shown, a system 565 is provided that includes at least one central processing unit 530 coupled to a communication bus 575. The communication bus 575 can be implemented using any suitable protocol, such as PCI (Peripheral Component Interconnect), PCI-Express, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol. The system 565 also includes main memory 540. Control logic (software) and data are stored in the main memory 540, which can take the form of random access memory (RAM).
[0166] The system 565 also includes an input device 560, parallel processing system 525, and a display device 545. A conventional CRT (cathode ray tube), LCD (liquid crystal display), LED (light emitting diode), plasma display, or the like can be used for the display device 545. User input can be received from the input device 560, e.g., a keyboard, mouse, touchpad, microphone, etc. Each of the aforementioned modules and / or devices can even be located on a single semiconductor platform, as is true of the system 565. Alternatively, individual modules can be located in different semiconductor platforms, which can be configured to operate as a
[0167] In addition, the system 565 can be coupled to a network (e.g., a telecommunications network, local area network (LAN), wireless network, wide area network (WAN) such as the Internet, peer-to-peer network, cable broadcast network, etc.) for communication with one or more other systems or devices.
[0168] System 565 may also include auxiliary storage (not shown). Auxiliary storage device 610 includes, for example, a hard disk drive and / or a removable storage drive, which represents a floppy disk drive, a tape drive, an optical disk drive, a digital versatile disk (DVD) drive, a recording device, or a universal serial bus (USB) flash memory. The removable storage drive reads from and / or writes to a removable storage unit in a well-known manner.
[0169] Computer programs or computer control logic algorithms may be stored in the main memory 540 and / or the secondary memory. When such computer programs are executed, the system 565 is enabled to perform various functions. The memory 540, memory and / or any other memory are possible examples of computer-readable media.
[0170] The architecture and / or functionality of each of the previous figures can be implemented in a general purpose computer system, a circuit board system, a game console system dedicated for entertainment purposes, a dedicated system, and / or any other desired system in any environment. For example, the system 565 can take the form of a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smartphone (e.g., a wireless, handheld device), a personal digital assistant (PDA), a digital camera, a vehicle, a head mounted display, a handheld electronic device, a mobile phone device, a television, a workstation, a game console, an embedded system, and / or any other type of logic.
[0171] Although various embodiments have been described above, it should be understood that they have been presented by way of example only and not limitation. Therefore, the breadth and scope of the preferred embodiment should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the appended claims and their equivalents.
[0172] Graphics processing pipeline
[0173] In one embodiment, the PPU 300 comprises a graphics processing unit (GPU). The PPU 300 is configured to receive commands specifying a shader program for processing graphics data. Graphics data can be defined as a set of primitives, such as points, lines, triangles, quadrilaterals, triangle strips, and the like. Typically, a primitive includes data specifying a plurality of vertices for the primitive (e.g., in a model space coordinate system) and attributes associated with each vertex of the primitive. The PPU 300 can be configured to process the graphics primitives to generate a frame buffer (e.g., pixel data for each pixel of a display).
[0174] An application writes model data (e.g., a collection of vertices and attributes) for a scene into memory, such as system memory or memory 304. The model data defines each object that is visible on the display. The application then makes an API call to the driver kernel to request that the model data be rendered and displayed. The driver kernel reads the model data and writes commands to one or more streams to perform operations to process the model data. The commands can reference different shader programs to be implemented on the SMs 440 of the PPU 300, including one or more of a vertex shader, a hull shader, a domain shader, a geometry shader, and a pixel shader. For example, one or more of the SMs 440 can be configured to execute a vertex shader program that processes a number of vertices defined by the model data. In one embodiment, different ones of the SMs 440 can be configured to execute different shader programs at the same time. For example, a first subset of the SMs 440 can be configured to execute a vertex shader program while a second subset of the SMs 440 can be configured to execute a pixel shader program. The first subset of the SMs 440 processes the vertex data to produce processed vertex data and writes the processed vertex data to L2 cache 460 and / or memory 304. After the processed vertex data is rasterized (e.g., converted from three-dimensional data to two-dimensional data in screen space) to generate fragment data, the second subset of the SMs 440 executes a pixel shader to generate processed fragment data, which is then blended with other processed fragment data and written to a frame buffer in memory 304. The vertex shader program and the pixel shader program can execute at the same time, processing different data from the same scene in a pipelined fashion until all of the model data for the scene has been rendered to the frame buffer. The contents of the frame buffer are then sent to a display controller to be displayed on a display device.
[0175] Images produced using one or more of the techniques disclosed herein can be displayed on a monitor or other display device. In some embodiments, the display device can be directly coupled to the system or processor that generates or renders the images. In other embodiments, the display device can be indirectly coupled to the system or processor, e.g., via a network. Examples of such networks include the Internet, mobile telecommunications networks, WIFI networks, and any other wired and / or wireless networking systems. When the display device is indirectly coupled, the images generated by the system or processor can be streamed over the network to the display device. Such streaming allows, for example, video games or other applications that render images to be executed on a server or in a data center, and the rendered images can be streamed and displayed on one or more user devices (e.g., computers, video game consoles, smartphones, other mobile devices, etc.) that are physically separate from the server or data center. Thus, the techniques disclosed herein can be applied to enhance streamed images and to enhance services that stream images, such as NVIDIA GeForce Now (GFN), Google Stadia, etc.
[0176] Machine Learning
[0177] Deep neural networks (DNNs) developed on processors such as PPU 300 have been used for a variety of use cases, from self-driving cars to faster drug development, from automatic image captioning in online image databases to intelligent real-time language translation in video chat applications. Deep learning is a technology that models the neural learning process of the human brain, learns continuously, gets smarter and provides more accurate results faster over time. Initially, an adult teaches a child how to correctly identify and classify various shapes, eventually without any guidance. Likewise, a deep learning or neural learning system needs to be trained in object identification and classification as it gets smarter and more efficient in identifying basic objects, occluded objects, etc., while also assigning context to the objects.
[0178] At the simplest level, a neuron in the human brain looks at various inputs received, assigns a level of importance to each of these inputs, and passes an output to other neurons to operate on it. Artificial neurons are the most basic model of a neural network. In one example, a neuron can receive one or more inputs representing various features of an object that the neuron is being trained to identify and classify, and assign a certain weight to each of these features based on how important that feature is in defining the shape of the object.
[0179] Deep neural network (DNN) models include multiple layers of connected nodes (e.g., neurons, Boltzmann machines, radial basis functions, convolutional layers, etc.) that can be trained with large amounts of input data to quickly solve complex problems with high accuracy. In one example, the first layer of a DNN model breaks down an input image of a car into individual parts and looks for basic patterns such as lines and angles. The second layer assembles the lines to look for higher-level patterns, such as wheels, windshields, and rearview mirrors. The next layer identifies the type of vehicle, and the last few layers generate a label for the input image to identify a specific model of a car brand.
[0180] Once a DNN is trained, it can be deployed and used to recognize and classify objects or patterns in a process called inference. Examples of inference (the process by which a DNN extracts useful information from a given input) include recognizing handwritten numbers on checks deposited into an ATM machine, recognizing friend images in a photograph, providing movie recommendations to over 50 million users, recognizing and classifying cars, pedestrians, and road hazards in different types of self-driving cars, or translating human speech in real-time.
[0181] During training, data flows through the DNN in a forward propagation phase until a prediction is made that indicates the label corresponding to the input. If the neural network did not label the input correctly, the error between the correct label and the predicted label is analyzed and the weights of each feature are adjusted in a backward propagation phase until the DNN correctly labels the input and other inputs in the training dataset. Training complex neural networks requires a large amount of parallel computing performance, including floating point multiplication and addition supported by PPU 300. Inference is less computationally intensive than training, which is a latency-sensitive process in which a trained neural network is applied to new inputs that have never been seen before to classify images, translate speech, and infer new information.
[0182] Neural networks rely heavily on matrix math operations, and complex multi-layer networks require a large amount of floating point performance and bandwidth for efficiency and speed. PPU 300 has thousands of processing cores optimized for matrix math operations and provides tens to hundreds of TFLOPS of performance, making it a computing platform that can provide the performance needed for deep neural network-based artificial intelligence and machine learning applications.
[0183] Further, images generated using one or more techniques disclosed herein can be used to train, test, or certify DNNs for recognizing objects and environments in the real world. Such images can include scenes of roads, factories, buildings, urban environments, rural environments, humans, animals, and any other physical objects or real environments. Such images can be used to train, test, or certify DNNs used in machines or robots to manipulate, process, or modify physical objects in the real world. Further, such images can be used to train, test, or certify DNNs used in autonomous vehicles to navigate and move the vehicle in the real world. Additionally, images generated using one or more techniques disclosed herein can be used to convey information to users of these machines, robots, and vehicles.
[0184] Note that the techniques described herein can be embodied in executable instructions stored in a computer-readable medium for use by or in connection with an instruction execution machine, system, apparatus, or device. As used herein, a "computer- readable medium" includes one or more of any suitable media for storing the executable instructions for use by or in connection with the instruction execution machine, system, apparatus, or device. Suitable storage formats include one or more of electronic, magnetic, optical, and electromagnetic formats. A non-exhaustive list of conventional exemplary computer-readable media includes: portable computer disks, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM), flash memory devices, optical storage devices, including portable compact disc (CD), portable digital video disc (DVD), and the like.
[0185] It should be understood that the arrangement of components shown in the drawings is for purposes of illustration, and other arrangements are possible. For example, one or more of the elements described herein can be implemented in whole or in part as electronic hardware components. Other elements can be implemented in software, hardware, or a combination of software and hardware. Also, some or all of these other elements can be combined, some can be omitted entirely, and additional components can be added, while still achieving the functionality described herein. Therefore, the subject matter described herein can be embodied in many different variations and all such variations are considered within the scope of the claims.
[0186] To facilitate an understanding of the subject matter described herein, a number of aspects are described in the context of an action sequence. Those skilled in the art will realize that the various actions can be performed by specific circuits or circuitry, by program instructions being executed by one or more processors, or by a combination of both. The descriptions of any of the action sequences herein are not intended to imply a particular order of execution unless otherwise specifically noted or clearly contradicted by context. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context.
[0187] The use of the terms “a” and “one” and “the” and similar referents in the context of describing the subject matter (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The use of the term “at least one” followed by a list of one or more items (for example, “at least one of A and B”) is to be construed to mean one item from the list of items (A or B) or any combination of two or more of the items in the list (A and B), unless otherwise indicated herein or clearly contradicted by context. Furthermore, unless otherwise required by context, singular terms used herein shall include pluralities and the plural terms shall include the singular. The term “another” is used herein to mean “one or more.” The term “or” is used herein to mean, and is used interchangeably with, the term “and / or,” unless otherwise indicated herein or unless contradicted by context. The term “including” is used herein to mean, and is used interchangeably with, the term “comprising.” The term “based on” is used herein to mean “based, at least in part, on,” unless otherwise indicated herein or unless contradicted by context. Furthermore, any and all examples or exemplary language (e.g., “such as”) used herein are intended to better illuminate the subject matter and do not pose a limitation to the scope of the subject matter unless otherwise claimed. The use of the term “based on” and other like phrases indicating a condition for bringing about a result is not intended to foreclose any other condition that could bring about the result unless otherwise indicated. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.
Claims
1. A circuit comprising: A decoder circuit configured to: receiving a selection ready signal, wherein the selection ready signal is negated until a selection signal generated by the combinational logic does not change, and the selection ready signal is asserted after the selection signal does not change; generating at least one sampling enable signal corresponding to a set of data input signals based on the selection signal, wherein the at least one sampling enable signal is negated at the same time as the selection ready signal is negated, and the at least one sampling enable signal is asserted in response to the assertion of the selection ready signal; as well as generating a hold signal, the hold signal being asserted while the at least one sampling enable signal is negated, and the hold signal being negated in response to the assertion of the at least one sampling enable signal; and A sampling circuit configured to: holding an output signal while the hold signal is being asserted, wherein the sampling circuit includes a feedback loop configured to assert the output signal while the hold signal is being asserted and the output signal is being asserted; and One of the data input signals is sampled according to the at least one sampling enable signal while the hold signal is negated to transfer a level of the sampled data input signal to the output signal.
2. The circuit of claim 1 , wherein the select signal comprises a multi-bit signal, and each bit in the select signal is associated with a different one of the set of data input signals, and only one bit is assigned a value at a time.
3. The circuit of claim 2, wherein each of the bits in the select signal is used to sample an associated data input signal.
4. The circuit of claim 2 , wherein the decoder circuit is further configured to: receiving a set of ready signals, each ready signal in the set of ready signals being associated with a different one of the set of data input signals, wherein Each ready signal is negated until the associated data input signal is unchanged, and Each ready signal is asserted after the associated data input signal is unchanged.
5. The circuit of claim 4 , wherein the decoder is further configured to: generate a set of enable signals, each enable signal in the set of enable signals being associated with a different one of the set of data input signals, wherein each of said enable signals being negated at the same time as the associated said ready signal is negated; and Each of the enable signals is asserted in response to the assertion of the associated ready signal when the associated data input signal is asserted.
6. The circuit of claim 1 , wherein the decoder is further configured to: receiving a set of ready signals, each ready signal in the set of ready signals being associated with a different one of the set of data input signals; Each ready signal is negated until the associated data input signal is unchanged, and Each ready signal is asserted after the associated data input signal is unchanged. 7 . The circuit of claim 1 , wherein the decoder circuit is further configured to assert the hold signal while the hold signal is asserted and the at least one sample enable signal is negated. 8 . The circuit of claim 1 , wherein the hold signal and the at least one sample enable signal are both asserted for a first duration.
9. The circuit of claim 1, wherein the hold signal is inverted to generate an output ready signal. 10 . The circuit of claim 1 , wherein the set of input data signals includes three input data signals, and the select signal is configured to select one of the three input data signals to generate the output signal.
11. The circuit of claim 1, wherein the set of input data signals includes four input data signals, and the select signal is configured to select two of the four input data signals to generate the output signal and the additional output signal.
12. The circuit of claim 1, wherein the circuit is included within a processor configured to generate an image, and the processor is part of a server or a data center, and streams the image to a user device.
13. The circuit of claim 1 , wherein the circuit is included within a processor configured to train, test, or validate a neural network employed in a machine, robot, or autonomous vehicle.
14. The circuit of claim 1 , wherein the circuit is included within a processor configured to implement a neural network model.
15. A computer-implemented method comprising: receiving a selection ready signal, wherein the selection ready signal is negated until a selection signal generated by the combinational logic does not change, and the selection ready signal is asserted after the selection signal does not change; generating at least one sampling enable signal corresponding to a set of data input signals based on the selection signal, wherein the at least one sampling enable signal is negated at the same time as the selection ready signal is negated, and the at least one sampling enable signal is asserted in response to the assertion of the selection ready signal; generating a hold signal, the hold signal being asserted while the at least one sampling enable signal is negated, and the hold signal being negated in response to the assertion of the at least one sampling enable signal; holding an output signal while the hold signal is being asserted, wherein the feedback loop is configured to assert the output signal while the hold signal is being asserted and the output signal is being asserted; as well as When the hold signal is negated, one of the data input signals is sampled according to the at least one sampling enable signal to transfer a value of the sampled data input signal to the output signal.
16. The computer-implemented method of claim 15, wherein the select signal comprises a multi-bit signal, and each bit in the select signal is associated with a different one of the set of data input signals, and only one bit is assigned a value at a time.
17. The computer-implemented method of claim 16, further comprising: receiving a set of ready signals, each ready signal in the set of ready signals being associated with a different one of the set of data input signals, wherein Each ready signal is negated until the associated data input signal is unchanged, and Each ready signal is asserted after the associated data input signal is unchanged.
18. The computer-implemented method of claim 15, wherein the generating, receiving, and sampling steps are performed within a processor configured to implement a neural network model.
19. A circuit comprising: a delay circuit configured to generate a ready signal that is negated at a first transition of a clock signal that begins a cycle of the clock signal and is asserted after a first delay relative to the first transition, wherein the first delay is at least as long as a second delay; as well as A sampling circuit configured to: receiving an input signal generated by combinatorial logic, wherein within the second delay after the first transition of a clock signal, a change in a first signal received at an input of the combinatorial logic causes a glitch in the input signal at an output of the combinatorial logic; as well as During the period, the input signal is sampled while the ready signal is asserted to transfer the level of the input signal to the output signal of the sampling circuit, wherein the ready signal is negated during the glitch and asserted after the glitch, and the input signal remains at the transferred level for a remainder of the period.
20. The circuit of claim 19, wherein the sampling circuit is further configured to hold the output signal at a constant level from a first transition of the clock signal until the input signal is sampled.
21. The circuit of claim 19, wherein the sampling circuit comprises a latching storage element.
22. The circuit of claim 21, wherein the latching storage element comprises cross-coupled logic gates configured to cause a feedback path to hold the output signal stable until after the input signal is sampled.
23. The circuit of claim 19, wherein the input signal changes at least once while the ready signal is negated.
24. The circuit of claim 19, wherein the delay circuit delays the clock signal by the first delay to generate the ready signal.
25. The circuit of claim 19, wherein the delay circuit inverts the clock signal to generate the ready signal.
26. The circuit of claim 19, wherein the second delay is a propagation delay of the combinatorial logic, and the first delay is equal to or greater than the second delay.
27. The circuit of claim 19, wherein the delay circuit is further configured to negate the ready signal after a third delay relative to the first delay.
28. The circuit of claim 27, wherein the first delay and the third delay occur within a period of the clock signal.
29. The circuit of claim 19 , wherein the sampling circuit is further configured to generate an output ready signal, the output ready signal being associated with the output signal and negated for the first delay, and the output ready signal being asserted once the level of the input signal is transferred to the output signal.
30. The circuit of claim 29, further comprising negating the ready signal in response to the assertion of the output ready signal.
31. The circuit of claim 19, wherein the circuit is included within a processor configured to generate an image, and the processor is part of a server or data center and streams the image to a user device.
32. The circuit of claim 19, wherein the circuit is included within a processor configured to train, test, or validate a neural network employed in a machine, robot, or autonomous vehicle.
33. The circuit of claim 19, wherein the circuit is included within a processor configured to implement a neural network model.
34. A computer-implemented method comprising: generating a ready signal that is negated at a first transition of a clock signal that begins a cycle of the clock signal and is asserted after a first delay relative to the first transition, wherein the first delay is introduced by a delay circuit and is at least as long as a second delay; receiving an input signal generated by combinatorial logic, wherein a change in a first signal received at an input of the combinatorial logic causes a glitch in the input signal at an output of the combinatorial logic within the second delay after the first transition of the clock signal; as well as During the period, the input signal is sampled while the ready signal is asserted to transfer the level of the input signal to the output signal, wherein the ready signal is negated during the glitch and asserted after the glitch, and the input signal remains at the transferred level for the remainder of the period.
35. The computer-implemented method of claim 34, further comprising holding the output signal at a constant level from a first transition of the clock signal until the input signal is sampled.
36. The computer-implemented method of claim 34, wherein the input signal changes at least once while the ready signal is negated.
37. The computer-implemented method of claim 34, wherein the second delay is a propagation delay of the combinatorial logic, and the first delay is equal to or greater than the second delay.
38. The computer-implemented method of claim 34, wherein the generating, receiving, and sampling steps are performed within a processor configured to implement a neural network model.
Citation Information
Patent Citations
Superconducting circuits and methods for latching data
US10547314B1
Method for changing CPU frequence under control of neural network
US20020171603A1
Apparatus and method for verifying glitch-free operation of a multiplexer
US20070057697A1