Reconfigurable Arithmetic Circuit for Low-Latency Parallel Computing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing systems face limitations in computation speed and energy efficiency for mathematically intensive applications such as artificial intelligence, neural networks, digital currencies, and blockchain, with inadequate scalability and high latency.
Innovation Solution
A reconfigurable processor architecture featuring an array of fractal cores with a reconfigurable arithmetic engine that includes input reordering queues, a multiplier shifter and combiner network, an accumulator circuit, and control logic, allowing for scalable, low-latency, energy-efficient processing of streaming data in real-time, capable of massively parallel operations and optimized for specific applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If existing computing systems are used for mathematically intensive applications, then general-purpose computing is maintained, but computation speed and energy efficiency are insufficient
Solution Approach 1:
The patent implements reconfigurable arithmetic engines that can dynamically change their operational configuration based on the specific computational task. The arithmetic engine includes reconfigurable multipliers, adders, and data paths that can be adjusted in real-time to optimize for different mathematical operations (e.g., matrix multiplication, convolution, dot products), thereby achieving both high computation speed and energy efficiency for mathematically intensive applications
Solution Approach 2:
The system changes operational parameters such as precision (e.g., 8-bit, 16-bit, 32-bit operations), parallelism degree, and data flow configuration to match the requirements of different applications. This allows the same hardware to efficiently handle diverse workloads from neural network inference to cryptocurrency mining without the energy waste of fixed-architecture systems
2Productivity
If computing systems are scaled up to improve performance, then computation capability increases, but heat dissipation and power consumption increase excessively
Solution Approach 1:
The computing system is divided into multiple independent reconfigurable arithmetic engines that can operate in parallel. Each engine is a self-contained unit with its own multipliers, adders, and control logic, allowing the system to scale computation capability by activating only the necessary number of engines for each task, thereby reducing overall power consumption and heat generation compared to scaling a monolithic system
Solution Approach 2:
The patent uses replicated arithmetic engine units that can be instantiated in different quantities based on computational needs. Rather than increasing the size of a single processor, the system creates multiple copies of optimized arithmetic units, each consuming minimal power, and activates the appropriate number of copies to achieve the desired computation capability while maintaining efficient power-to-performance ratio
3Adaptability or versatility
If fixed-architecture processors are used, then hardware simplicity is maintained, but adaptability to different applications is limited
Solution Approach 1:
The reconfigurable arithmetic engine is designed as a universal computing unit that can perform multiple functions through configuration changes. The same physical hardware can be configured for different operations including but not limited to matrix multiplication, vector operations, polynomial evaluation, and cryptographic functions, eliminating the need for separate specialized hardware for each application type
Solution Approach 2:
The arithmetic engine incorporates dynamic reconfiguration capabilities where control logic can modify the operational mode, data path width, and computational algorithm in real-time based on the input task requirements. This dynamic adaptability allows a single hardware design to replace multiple fixed-architecture processors while managing complexity through systematic control mechanisms
4Loss of time
If processing latency is reduced for real-time applications, then response speed improves, but system complexity increases
Solution Approach 1:
The system pre-loads configuration data and operational parameters into on-chip memory before execution begins. Input data is buffered and pre-processed through reordering queues that prepare data in the optimal format for the specific computational task, eliminating runtime reconfiguration overhead and reducing processing latency without requiring overly complex real-time adaptation mechanisms
Data Source
AI summary
A representative reconfigurable processing circuit and a reconfigurable arithmetic circuit are disclosed, each of which may include input reordering queues; a multiplier shifter and combiner network coupled to the input reordering queues; an accumulator circuit; and a control logic circuit, along with a processor and various interconnection networks. A representative reconfigurable arithmetic circuit has a plurality of operating modes, such as floating point and integer arithmetic modes, logical manipulation modes, Boolean logic, shift, rotate, conditional operations, and format conversion, and is configurable for a wide variety of multiplication modes. Dedicated routing connecting multiplier adder trees allows multiple reconfigurable arithmetic circuits to be reconfigurably combined, in pair or quad configurations, for larger adders, complex multiplies and general sum of products use, for example.


