Near-Memory Processing Units for N-Point Butterfly Network Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in efficiently processing large N-point butterfly networks due to the Von Neumann bottleneck, which limits processing speed and energy efficiency.
Innovation Solution
A computer program product and method that map data from N input nodes of a first transform network to n input nodes of near-memory processing units implementing a smaller transform network, using a mapping generator system that performs iterative folding and permutation operations to reduce the network size and optimize data processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If an N-point butterfly network is implemented using conventional processing architectures, then the transform computation can be performed, but the Von Neumann bottleneck limits processing speed and energy efficiency
Solution Approach 1:
The patent transitions from conventional von Neumann architecture to a near-memory processing architecture, effectively adding a spatial dimension to data access. Processing units are placed adjacent to memory blocks, allowing data to be processed in-place without traditional memory access bottlenecks. This dimensional change in architecture enables simultaneous data storage and processing, dramatically improving energy efficiency and processing speed for butterfly network computations.
2Device complexity
If a smaller n-point transform network is used instead of an N-point network, then hardware resources are reduced, but the transform accuracy and completeness deteriorate
Solution Approach 1:
The patent segments the large N-point butterfly network into multiple smaller n-point butterfly networks through a systematic mapping approach. The N input nodes are divided and mapped to n input nodes of the smaller network, with each node in the smaller network receiving data from multiple nodes in the larger network. This segmentation allows the use of smaller, less complex hardware while maintaining the ability to process the original large-scale transform through coordinated computation across multiple processing units.
Solution Approach 2:
The smaller n-point processing units are designed to be universal and multi-functional, capable of handling computations that represent portions of the larger N-point transform. Each processing unit in the smaller network performs operations that contribute to the overall N-point transform result, allowing limited hardware resources to achieve comprehensive transform functionality through efficient resource utilization and data mapping.
3Productivity
If data is loaded into processing units for computation, then transform operations can be performed, but memory access time and bandwidth limitations increase latency
Solution Approach 1:
The patent introduces near-memory processing units as intermediaries between traditional memory and compute units. These processing units reside in the memory interface or memory controller, acting as intermediaries that can perform butterfly network computations directly on data stored in memory blocks. This eliminates the need to transfer data between separate memory and compute units, significantly reducing latency while maintaining high computation throughput.
Data Source
AI summary
Provided are computer program product, system, and method for mapping data for nodes in a first transform network to input nodes of near-memory processing units implementing a smaller transform network. A plurality of processing units, which are interconnected, receive input data for n input nodes for a second transform network to process at interlinked stages of nodes in the processing units. A mapping maps N input nodes for the first transform network to the n input nodes of the second transform network. N is greater than n and a plurality of the N input nodes of the first transform network map to one of the n input nodes of the second transform network. A transform manager uses the mapping to map the N input nodes to n input nodes and loads received input data for the n input nodes into the processing units to perform computations in the processing units.


