Dense Algorithm Integrated Circuit Multi-Directional DMA Transfers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current integrated circuitry for autonomous robotics and vehicles lacks robust processing capabilities to handle complex machine learning algorithms and sensor data processing efficiently, leading to inefficiencies in real-time computing and sensor signal processing.

Innovation Solution

A dense algorithm and perception processing integrated circuit architecture with array cores, border cores, and a dispatcher that enables accelerated memory transfers and efficient data processing through direct memory access (DMA) techniques, including multi-directional data accessing instructions and transpositional DMA instructions, to optimize data movement and processing within the circuit.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional GPU architectures are used for sensor processing, then basic computation capabilities are provided, but complex machine learning algorithms and sensor fusion tasks cannot be handled efficiently

Engineering Contradiction:
Improvecapability to handle complex machine learning algorithmsVSAvoidreal-time processing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system segments processing tasks into distinct functional units including array cores for parallel computation, border cores for data management, and specialized DMA controllers for memory operations. This segmentation allows each unit to be optimized for its specific function while working together to handle complex machine learning algorithms efficiently

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an n-dimensional memory architecture with multi-directional access capabilities, allowing data to be accessed along multiple axes simultaneously. This dimensional expansion enables efficient handling of multi-dimensional tensor operations required for machine learning while maintaining real-time processing speeds

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If additional disparate circuitry is assembled to a traditional GPU to handle sensor fusion and path planning, then processing capabilities are extended, but system complexity and inefficiency increase

Engineering Contradiction:
Improvesensor fusion and path planning capabilitiesVSAvoidfragmented and piecemeal circuit assembly
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges multiple processing functions including sensor data processing, machine learning inference, sensor fusion, and path planning into a single integrated circuit architecture. This unified design eliminates the need for assembling disparate circuitry while reducing overall system complexity and improving inter-component communication efficiency

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The array core architecture provides universal processing capabilities that can handle various machine learning algorithms and processing tasks through configuration rather than requiring separate dedicated hardware for each function. This multi-functionality extends adaptability without increasing physical circuit complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Device complexity

If general purpose CPUs are used for processing, then system simplicity is maintained, but computation throughput is insufficient for real-time machine learning

Engineering Contradiction:
Improveprocessing architecture simplicityVSAvoidcomputation throughput
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The system employs dynamic parallel processing where array cores can be dynamically configured and activated based on computational requirements. This dynamic architecture maintains simplicity when full parallelism is not needed while providing high throughput when complex machine learning operations require it, adapting to workload demands in real-time

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11907146B2Systems and methods for intelligently implementing concurrent transfers of data within a machine perception and dense algorithm integrated circuit
Publication Date: 2024.02.20 QUADRIC IO INC
  • US11907146B2 patent drawing
  • US11907146B2 patent drawing
  • US11907146B2 patent drawing

AI summary

System and method for implementing accelerated memory transfers in an integrated circuit includes identifying memory access parameters for configuring memory access instructions for accessing a target corpus of data from within a defined region of an n-dimensional memory; converting the memory access parameters to direct memory access (DMA) controller-executable instructions, wherein the converting includes: (i) defining dimensions of a data access tile based on a first parameter of the memory access parameters; (ii) generating multi-directional data accessing instructions that, when executed, automatically moves the data access tile along multiple distinct axes within the defined region of the n-dimensional memory based at least on a second parameter of the memory access parameters; transferring a corpus of data from the n-dimensional memory to a target memory based on executing the DMA controller-executable instructions.