Compressed Routing Tables for NTT Tile Arrays
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing solutions for data movement between tiles in parallel processing devices for NTT/iNTT operations require significant silicon area and either result in high resource usage or reduced throughput, limiting the performance of fully homomorphic encryption workloads.
Innovation Solution
A scalable and reconfigurable parallel processing device with programmable contention-free routing schedules and compressed routing tables that allow efficient data routing between tiles, using a 2D mesh interconnect architecture and reduced-entry multiplexers for look-up table operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional routing tables are used for data movement between tiles, then routing coverage is complete, but silicon area and resource usage increase significantly
Solution Approach 1:
The patent extracts and removes redundant routing entries from traditional routing tables. By analyzing the specific data movement patterns required for NTT/iNTT operations, the invention identifies and eliminates unnecessary routing paths, keeping only the essential entries needed for correct operation. This extraction principle directly reduces routing table size and silicon area while preserving complete routing coverage for required operations.
Solution Approach 2:
The patent applies local quality by optimizing routing tables specifically for NTT/iNTT workload characteristics rather than providing general-purpose routing. The routing tables are tailored to the local requirements of number-theoretic transform operations, including specific data permutation patterns and tile communication needs. This localized optimization reduces overall routing table size while maintaining complete coverage for the target workload.
2Area of stationary object
If routing tables are compressed to reduce silicon area, then resource usage decreases, but throughput may be reduced
Solution Approach 1:
The patent applies preliminary action by pre-computing and pre-organizing compressed routing tables offline before runtime execution. The compression algorithms and routing path optimizations are prepared in advance, allowing the hardware to use pre-computed lookup tables during actual NTT/iNTT operations. This eliminates runtime compression overhead and ensures that throughput is not reduced, while still achieving significant silicon area savings from the compressed table structure.
3Productivity
If larger routing tables are used to maintain peak performance, then throughput is preserved, but device complexity and silicon area grow
Solution Approach 1:
The patent applies parameter changes by transforming the routing table representation from traditional full-entry format to a compressed format with reduced entries. The invention changes parameters such as routing table entry count, entry width, and organization structure to optimize for both area and performance. These parameter transformations maintain peak throughput for NTT/iNTT operations while significantly reducing device complexity and silicon area requirements.
Data Source
AI summary
Examples include techniques for contention-free routing for number-theoretic-transform (NTT) or inverse-NTT (iNTT) computations routed through a parallel processing device. Examples include a tile array that includes a plurality of tiles arranged in a 2-dimensional mesh interconnect-based architecture. Each tile includes a plurality of compute elements configured to execute NTT or iNTT computations associated with a fully homomorphic encryption workload. Contention-free routing to include use of grouped or compressed source addresses to be used in routing tables maintained at tiles of the tile array.


