Neural Network Weight Mapping to Processing Core Array
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current integrated circuit architectures, particularly GPUs, are not optimized for handling complex machine learning algorithms used in machine perception technologies, leading to inefficiencies in processing sensor data for autonomous robotics and vehicles, which requires high-performance and real-time computing capabilities.
Innovation Solution
A dense algorithm and perception processing integrated circuit architecture with an array of processing cores, a dispatcher, and optimized memory management that enables fast Fourier transform (FFT) matrix multiply operations, bit-reversed input arrays, and efficient data movement instructions to perform parallel computations, reducing latency and enhancing processing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional GPUs are used for sensor data processing, then general-purpose computation capability is provided, but processing efficiency for complex machine learning algorithms is insufficient
Solution Approach 1:
The system segments the monolithic GPU architecture into an array of independent processing cores, each capable of executing machine learning algorithms autonomously. This segmentation allows each core to be optimized for specific ML workloads while maintaining overall system versatility for different sensor processing tasks.
Solution Approach 2:
Each processing core in the array is equipped with local memory and specialized instruction sets tailored for machine learning operations. This local quality enhancement enables faster data access and execution of ML algorithms without relying on centralized memory bottlenecks, thereby improving processing efficiency while maintaining adaptability across different ML models.
2Adaptability or versatility
If additional circuitry is assembled to handle advanced perception processing needs, then processing capabilities for sensor fusion and path planning are enhanced, but system complexity and inefficiency increase
Solution Approach 1:
The array of processing cores provides a universal platform that can handle diverse perception processing tasks including sensor fusion, path planning, and object detection through software configuration rather than hardware specialization. This multi-functionality approach enables the system to adapt to different processing needs without increasing physical complexity.
Solution Approach 2:
The system employs dynamic workload distribution and task scheduling mechanisms that allow processing cores to be allocated differently based on real-time computational demands. This dynamic allocation enables the same hardware architecture to efficiently handle varying complexity levels of perception tasks without requiring additional circuitry for each specific function.
3Adaptability or versatility
If additional circuitry is assembled to handle advanced perception processing needs, then processing capabilities for sensor fusion and path planning are enhanced, but computational inefficiency occurs
Solution Approach 1:
The system merges multiple processing functions into a unified array of cores that can execute machine learning algorithms for sensor fusion, path planning, and other perception tasks simultaneously. This consolidation eliminates the computational overhead and data transfer inefficiencies associated with separate dedicated circuits while maintaining full functional capability.
Solution Approach 2:
The processing array enables continuous computation pipelines where data can flow continuously through multiple processing stages without interruption or reconfiguration. This continuity is achieved through the homogeneous architecture that allows any core to take over computational tasks, eliminating idle time and maintaining high productivity across diverse perception processing workloads.
Data Source
AI summary
Systems and methods of configuring an array of processors of an integrated circuit includes identifying a fast Fourier transform (FFT) matrix multiply of input data, wherein the FFT matrix multiply of the input data includes a bit-reversed input array, configuring the array of processing cores based on the bit-reversed input array, wherein the configuring the array of processing cores includes storing the input bits of the bit-reversed input array within memory circuits of distinct processing cores of an array of processing cores of the integrated circuit based on an input bit mapping that identifies a pre-determined storage location within the array of processing cores of each input bit of the bit-reversed input array, and performing matrix multiply computations between weight stages of the FFT matrix multiply and the input bits of the bit-reversed input array stored within the memory circuits of the distinct processing cores.


