Pulse Generation Circuit With End-of-Computation Flag Timing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing in-memory computing systems, such as those described in U.S. Pat. No. 11,322,195B2, suffer from suboptimal processing times for matrix-vector multiplications due to idle states caused by synchronizing with the worst-case processing time window, which is not self-aware of the actual computational latency, especially when few pulses are generated.
Innovation Solution
The proposed method generates an end-of-computation flag that is self-aware of the computational latency, adjusting the time windows based on the maximum magnitude of input neuron activations, and optionally uses a duty-cycle-controlled clock sprinting scheme to minimize idle time, thereby optimizing processing time without additional circuitry.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the pulse generation circuit is synchronised with the worst-case processing time window, then the circuit can handle the maximum number of pulses, but the processing time increases due to idle states when fewer pulses are generated
Solution Approach 1:
The patent applies dynamics by making the time window duration adaptive rather than fixed. The circuit dynamically adjusts the processing time window based on the actual number of pulses generated, allowing it to handle variable workloads efficiently. This resolves the contradiction by enabling the circuit to maintain reliability for maximum pulse handling while minimizing idle time when fewer pulses are generated.
Solution Approach 2:
The patent changes the parameter of time window duration from a static worst-case value to a dynamic value that adapts to the actual computational latency. By monitoring the number of pulses generated and adjusting the time window accordingly, the system maintains the capability to handle maximum pulses while reducing processing time for smaller workloads, thus resolving the time loss issue.
2Stability of the object's composition
If the circuit stays idle to synchronize with the worst-case processing time window, then all digital counters can be synchronized, but the processing time is not optimal especially when few pulses are generated per time window
Solution Approach 1:
The patent implements dynamics by making the synchronization mechanism adaptive. Instead of forcing all counters to wait for the worst-case scenario, the system dynamically synchronizes counters based on their actual completion times. This allows the circuit to maintain counter synchronization stability while significantly improving processing throughput by eliminating unnecessary idle waiting periods.
Solution Approach 2:
The patent applies continuity of useful action by ensuring that the circuit remains active and productive throughout the processing period. By adjusting the time window to match the actual computational latency rather than the worst-case scenario, the system eliminates idle states and maintains continuous useful action, thereby improving productivity while preserving counter synchronization through adaptive control mechanisms.
3Loss of time
If the circuit is made self-aware of computational latency, then processing time can be optimized, but additional circuitry such as sparsity index encoder/decoder would be required
Solution Approach 1:
The patent applies self-service by enabling the circuit to automatically monitor and determine its own computational latency without external control. The pulse generation circuit itself tracks the number of pulses generated and uses this information to adjust the time window duration. This self-awareness mechanism reduces computational latency while avoiding the need for complex additional circuitry like sparsity index encoders or decoders, as the system serves itself by using its own operational data for optimization.
Data Source
AI summary
The present invention proposes a novel integrated circuit architecture for in-memory computing matrix-vector multipliers such that the computational latency is inversely proportional to the incoming magnitude of neuron activations. The main contribution of the present invention is that the proposed circuit is self-aware of the computational latency. At the end of the generated data pulses in which the number of pulses is proportional to the magnitude of incoming neuron activations, the circuit generates an end-of-computation flag such that the computing circuit can shorten the processing time of matrix-vector multiplications. The present invention can be integrated with any kind of analogue readout circuit, and the proposed circuit can be integrated with any kind of memory elements.


