FPGA Multiplier LUT6 Sharing for Lower Area and Power
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing FPGA multiplier designs are inefficient in terms of area usage and power consumption, limiting the number of multipliers that can be implemented on a single chip and affecting the performance of neural networks that require numerous multiplication operations.
Innovation Solution
The use of radix-4 modified Booth encoding with novel six-input lookup tables (LUT6s) that allow for efficient sharing of inputs between pairs of LUT6 blocks, reducing the number of logic blocks required for multiplication and enabling more compact implementations, such as supporting both signed and unsigned multiplications without additional logic.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If traditional FPGA multiplier designs are used, then multiplication functionality is achieved, but area usage is excessive and power consumption is high
Solution Approach 1:
The patent merges multiple LUT6 blocks to share common inputs, combining their functionality into a unified multiplier structure. This merging reduces the total number of logic blocks required while maintaining multiplication capability, directly addressing the area usage problem.
Solution Approach 2:
The LUT6-based multiplier design implements universal functionality that can perform both signed and unsigned multiplications without requiring additional dedicated logic blocks. This multi-functionality increases the effective number of multipliers that can be packed on a single chip.
2Use of energy by stationary object
If traditional FPGA multiplier designs are used, then multiplication operations are performed, but power consumption is excessive
Solution Approach 1:
By merging LUT6 blocks to share inputs, the patent reduces the total number of active logic elements performing multiplication operations. Fewer logic blocks mean lower overall power consumption, enabling higher training rates for neural networks that require numerous multiplication operations.
Solution Approach 2:
The patent changes the fundamental parameter of logic block configuration by using six-input LUTs instead of traditional smaller LUTs. This parameter change optimizes the power-area efficiency of each multiplication operation, reducing power consumption per operation while maintaining computational capability.
3Productivity
If more multipliers are packed on a single chip, then neural network performance improves, but area constraints are violated
Solution Approach 1:
The patent merges multiple LUT6 blocks into shared input structures, effectively packing more multiplier functionality into the same physical chip area. This merging approach increases the density of multipliers per chip without violating area constraints.
Solution Approach 2:
By changing to six-input LUTs with optimized input sharing, the patent increases the functional density of logic blocks on the chip. This parameter change allows more multipliers to be implemented within the same area budget, directly improving neural network performance.
Data Source
AI summary
In some example embodiments a logical block comprising twelve inputs and two six-input lookup tables (LUTs) is provided, wherein four of the twelve inputs are provided as inputs to both of the six-input lookup tables. This configuration supports efficient field programmable gate array (FPGA) implementation of multipliers. Each six-input LUT comprises two five-input lookup tables (LUT5s) that are used to form Booth encoding multiplier building blocks. The five inputs to each LUT5 are two bits from a multiplier and three Booth-encoded bits from a multiplicand. By assembling building blocks, multipliers of arbitrary size may be formed.


