Neural Network Weight Mapping to Processing Core Array

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current integrated circuit architectures, particularly GPUs, are not optimized for handling complex machine learning algorithms used in machine perception technologies, leading to inefficiencies in processing sensor data for autonomous robotics and vehicles, which requires high-performance and real-time computing capabilities.

Innovation Solution

A dense algorithm and perception processing integrated circuit architecture with an array of processing cores, a dispatcher, and optimized memory management that enables fast Fourier transform (FFT) matrix multiply operations, bit-reversed input arrays, and efficient data movement instructions to perform parallel computations, reducing latency and enhancing processing efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional GPUs are used for sensor data processing, then general-purpose computation capability is provided, but processing efficiency for complex machine learning algorithms is insufficient

Engineering Contradiction:
Improveprocessing capability for machine learning algorithmsVSAvoidprocessing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system segments the monolithic GPU architecture into an array of independent processing cores, each capable of executing machine learning algorithms autonomously. This segmentation allows each core to be optimized for specific ML workloads while maintaining overall system versatility for different sensor processing tasks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each processing core in the array is equipped with local memory and specialized instruction sets tailored for machine learning operations. This local quality enhancement enables faster data access and execution of ML algorithms without relying on centralized memory bottlenecks, thereby improving processing efficiency while maintaining adaptability across different ML models.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If additional circuitry is assembled to handle advanced perception processing needs, then processing capabilities for sensor fusion and path planning are enhanced, but system complexity and inefficiency increase

Engineering Contradiction:
Improveprocessing capability for sensor fusion and path planningVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The array of processing cores provides a universal platform that can handle diverse perception processing tasks including sensor fusion, path planning, and object detection through software configuration rather than hardware specialization. This multi-functionality approach enables the system to adapt to different processing needs without increasing physical complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system employs dynamic workload distribution and task scheduling mechanisms that allow processing cores to be allocated differently based on real-time computational demands. This dynamic allocation enables the same hardware architecture to efficiently handle varying complexity levels of perception tasks without requiring additional circuitry for each specific function.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If additional circuitry is assembled to handle advanced perception processing needs, then processing capabilities for sensor fusion and path planning are enhanced, but computational inefficiency occurs

Engineering Contradiction:
Improveprocessing capability for sensor fusion and path planningVSAvoidcomputational efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system merges multiple processing functions into a unified array of cores that can execute machine learning algorithms for sensor fusion, path planning, and other perception tasks simultaneously. This consolidation eliminates the computational overhead and data transfer inefficiencies associated with separate dedicated circuits while maintaining full functional capability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The processing array enables continuous computation pipelines where data can flow continuously through multiple processing stages without interruption or reconfiguration. This continuity is achieved through the homogeneous architecture that allows any core to take over computational tasks, eliminating idle time and maintaining high productivity across diverse perception processing workloads.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS11392667B2Systems and methods for an intelligent mapping of neural network weights and input data to an array of processing cores of an integrated circuit
Publication Date: 2022.07.19 QUADRIC IO INC
  • US11392667B2 patent drawing
  • US11392667B2 patent drawing
  • US11392667B2 patent drawing

AI summary

Systems and methods of configuring an array of processors of an integrated circuit includes identifying a fast Fourier transform (FFT) matrix multiply of input data, wherein the FFT matrix multiply of the input data includes a bit-reversed input array, configuring the array of processing cores based on the bit-reversed input array, wherein the configuring the array of processing cores includes storing the input bits of the bit-reversed input array within memory circuits of distinct processing cores of an array of processing cores of the integrated circuit based on an input bit mapping that identifies a pre-determined storage location within the array of processing cores of each input bit of the bit-reversed input array, and performing matrix multiply computations between weight stages of the FFT matrix multiply and the input bits of the bit-reversed input array stored within the memory circuits of the distinct processing cores.