Pulse Generation Circuit With End-of-Computation Flag for Idle-Time Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing in-memory computing systems, particularly in matrix-vector multipliers, suffer from suboptimal processing times due to idle states when few pulses are generated, as they are synchronized with the worst-case processing time window, leading to inefficient utilization of computational resources.
Innovation Solution
The proposed method generates an end-of-computation flag that adapts to the actual magnitude of incoming neuron activations, dynamically adjusting computational latency and minimizing idle states without additional circuitry, and optionally incorporates a duty-cycle-controlled clock sprinting scheme to further reduce idle time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the pulse-generating circuit is synchronised with the worst-case processing time window (maximum number of pulses), then the circuit can handle all possible input magnitudes, but the processing time increases and the circuit stays idle when few pulses are generated
Solution Approach 1:
The patent applies dynamics by making the processing time window adaptive rather than fixed. The circuit dynamically adjusts the time window duration based on the actual number of pulses generated for each input magnitude. When few pulses are generated, the time window is shortened accordingly, eliminating idle waiting time. When maximum pulses are generated, the time window extends to handle the worst case, ensuring all input magnitudes are properly processed.
Solution Approach 2:
The patent changes the parameter of processing time window duration from a constant worst-case value to a variable that adapts to the actual computation needs. By monitoring the number of pulses generated and adjusting the time window parameter accordingly, the system optimizes processing time for each specific input while maintaining the ability to handle all possible input magnitudes.
2Reliability
If the circuit waits for the maximum processing time window to complete, then all computation cases are covered, but the circuit remains in idle state for low-magnitude inputs
Solution Approach 1:
The patent implements feedback by monitoring the actual number of pulses generated during computation and using this information to adjust the processing time window. The circuit receives feedback about the input magnitude through pulse counting and dynamically responds by shortening or extending the time window, ensuring computation completeness while maximizing active processing utilization by eliminating unnecessary idle time.
Solution Approach 2:
The system transitions from a static, fixed time window approach to a dynamic, adaptive approach where the processing time window changes based on real-time computation needs. This dynamic adjustment ensures that the circuit maintains computation completeness for all cases while significantly improving productivity by reducing idle states for low-magnitude inputs.
3Adaptability or versatility
If additional circuitry such as sparsity index encoder/decoder is added to manage sparsity, then sparsity management is improved, but the device complexity increases
Solution Approach 1:
The patent applies self-service by enabling the circuit to automatically adapt to sparsity and varying input magnitudes using its existing pulse generation and counting mechanisms. No additional sparsity index encoder/decoder circuitry is needed because the system self-adjusts the processing time window based on the actual number of pulses generated, inherently managing sparsity without increasing device complexity.
Solution Approach 2:
The existing pulse-generating circuit is made multi-functional by enabling it to handle both dense and sparse inputs, as well as all magnitude ranges, through dynamic time window adjustment. This universal approach allows the same circuit to manage sparsity effectively without requiring dedicated sparsity management hardware, thereby avoiding increased device complexity.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present invention proposes a novel integrated circuit architecture for in-memory computing matrix-vector multipliers such that the computational latency is inversely proportional to the incoming magnitude of neuron activations. The main contribution of the present invention is that the proposed circuit is self-aware of the computational latency. At the end of the generated data pulses in which the number of pulses is proportional to the magnitude of incoming neuron activations, the circuit generates an end-of-computation flag such that the computing circuit can shorten the processing time of matrix-vector multiplications. The present invention can be integrated with any kind of analogue readout circuit, and the proposed circuit can be integrated with any kind of memory elements.