Sparse Recurrent Neural Network Equilibrium Computation Acceleration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Sparse recurrent neural networks face high power consumption and performance fluctuations due to the lack of consideration for sparsity in weight matrices during computation, leading to inefficient voltage and clock frequency usage.
Innovation Solution
An equilibrium computation acceleration method and system that dynamically adjusts the operating voltage and frequency based on the sparsity of the weight matrix, selecting appropriate computation submodules to perform zero-hop and multiply-add operations, thereby reducing power consumption and improving computation speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a same voltage and clock frequency are supplied to the computation array at all time points, then the computation array operates stably, but the power consumption is high and causes performance fluctuation
Solution Approach 1:
The patent implements dynamic voltage and frequency adjustment by dividing the computation array into multiple submodules with different operating voltages and clock frequencies. The system dynamically selects and switches between submodules based on the sparsity of weight matrices, transforming the static computation array into a dynamic system that adapts its operating parameters to match computational requirements, thereby reducing power consumption while maintaining stability.
Solution Approach 2:
The patent changes the operating parameters (voltage and clock frequency) of computation submodules based on the sparsity degree of weight matrices. By arbitrating sparsity and selecting submodules with appropriate voltage-frequency configurations, the system optimizes power consumption according to actual computational needs, directly applying parameter changes to resolve the contradiction between stable operation and energy efficiency.
2Productivity
If a computation array with a larger order of magnitude is designed to handle large quantity of multiplication operations, then the computation capability is improved, but the hardware resource occupation increases
Solution Approach 1:
The patent segments the computation array into multiple computation submodules, each with different operating voltages and clock frequencies. This segmentation allows the system to handle large quantities of multiplication operations by distributing tasks across multiple smaller submodules rather than requiring one large computation array, thereby reducing overall hardware resource occupation while maintaining high computation capability.
Solution Approach 2:
The computation submodules are designed to be multi-functional, capable of handling different types of operations (zero-hop operation and multiply-add operation) with varying computational intensities. This universality allows a single set of submodules to serve multiple purposes, reducing the need for dedicated hardware resources for each operation type and improving resource utilization efficiency.
3Use of energy by moving object
If the sparsity of the weight matrix is considered to dynamically adjust voltage and clock frequency, then the power consumption is reduced, but the system complexity increases
Solution Approach 1:
The patent performs preliminary arbitration of weight matrix sparsity before computation to determine the appropriate computation submodule. By pre-calculating the sparsity degree and selecting the suitable submodule in advance, the system avoids complex real-time adjustments during computation, thereby reducing overall system complexity while achieving power consumption optimization through dynamic voltage and frequency adjustment.
Data Source
AI summary
An equilibrium computation acceleration method and system for a sparse recurrent neural network determine scheduling information based on an arbitration result of sparsity of a weight matrix, select a computation submodule having operating voltage and operating frequency that match the scheduling information or a computation submodule having operating voltage and operating frequency that are adjusted to match the scheduling information, and use the selected computation submodule to perform a zero-hop operation and a multiply-add operation in sequence, to accelerate equilibrium computation. The equilibrium computation acceleration system includes a data transmission module, an equilibrium computation scheduling module with a plurality of independent built-in computation submodules, and a voltage-adjustable equilibrium computation module. An error monitor is configured to achieve dynamic voltage adjustment. The data transmission module quickly exchanges a configuration of a read/write memory to reduce additional data transmission.


