Processor-in-Memory Cluster with Look-Up Tables for Data-Intensive Applications
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current von Neumann computing models, used in CPUs and GPUs, suffer from high latency and inefficiency due to sequential data communication, limiting performance in data-intensive applications like Machine Learning and Data Encryption, and lack support for ultra-efficient hardware for operations like Convolutional Neural Networks.
Innovation Solution
A processor-in-memory (PIM) cluster with programmable look-up tables (LUTs) integrated within a DRAM chip, allowing for massively parallel computations by minimizing data communication latency and power consumption, and supporting operations beyond traditional CMOS logic-based computing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If data operands are stored in a separate memory chip and fetched through a narrow bandwidth communication channel, then memory capacity and processing capability can be independently optimized, but this results in high latency and inefficient data communication
Solution Approach 1:
The patent merges the processing elements and memory cells into a single integrated structure where processing cores are formed within the memory chip itself. This integration eliminates the separate memory chip and narrow bandwidth communication channel, allowing data processing to occur directly where data is stored, thereby dramatically improving data communication speed while maintaining manageable system complexity through unified architecture design.
2Productivity
If parallel processing elements are integrated inside the memory chip, then data communication latency and power dissipation are minimized, but the functionality is limited to specific operations
Solution Approach 1:
The patent implements universal processing cores within the memory chip that can perform multiple operational modes including logical operations, arithmetic operations, and data movement operations. Each processing core contains configurable logic units and data pathways that can be dynamically programmed to execute different computational tasks, thereby achieving high parallelization while maintaining operational versatility across different application domains.
Solution Approach 2:
The processing elements within the memory chip are designed with dynamic reconfigurability, allowing the same hardware structure to adapt its functionality based on computational requirements. The processing cores can be programmed to perform different operations and can dynamically allocate resources for parallel execution of multiple tasks, enabling both high productivity through parallelization and adaptability through operational flexibility.
3Speed
If large Look-Up Tables are used for performing large-scale multiplications, then multiplication operations can be accelerated, but the memory intensity increases and functionality is limited to multiplication only
Solution Approach 1:
Instead of using a single large Look-Up Table that consumes excessive memory resources, the patent segments the multiplication functionality into distributed processing elements within the memory chip. Each processing core contains smaller, optimized lookup structures and computation units that work in parallel to achieve large-scale multiplication acceleration. This segmentation reduces the memory footprint of individual lookup tables while maintaining overall multiplication performance through parallel execution across multiple cores.
Data Source
AI summary
A processing element includes a PIM cluster configured to read data from and write data to an adjacent DRAM subarray, wherein the PIM cluster has a plurality of processing cores, each processing core of the plurality of processing cores containing a look-up table, and a router connected to each processing core, wherein the router is configured to communicate data among each processing core; and a controller unit configured to communicate with the router, wherein the controller unit contains an executable program of operational decomposition algorithms. The look-up tables can be programmable. A DRAM chip including a plurality of DRAM banks, each DRAM bank having a plurality of interleaved DRAM subarrays and a plurality of the PIM clusters configured to read data from and write data to an adjacent DRAM subarray is disclosed.


