On-Chip Memory Controller With Dual-Side Arbitration for Shared-Memory Access

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Edge devices face memory bandwidth bottlenecks when processing large-scale machine learning models, particularly for real-time data processing, which conventional methods like 3D packaging and on-chip caches fail to address effectively, leading to high latency and increased costs.

Innovation Solution

A high-performance on-chip memory controller architecture with backside and frontside arbitration controllers manages memory access for different hardware components, using separate protocols and priority-based arbitration to enhance data transfer efficiency, including data compression and decompression, and supports low-latency, high-throughput operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional memory controllers are used in edge devices, then device complexity is reduced, but memory bandwidth and throughput are insufficient for large-scale machine learning models

Engineering Contradiction:
Improvememory bandwidthVSAvoidmemory controller architecture
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The memory controller is divided into multiple arbitration controllers (first arbitration controller, second arbitration controller, third arbitration controller) that independently manage different memory bank groups. This segmentation enables parallel processing of memory access requests from multiple hardware components, significantly increasing memory bandwidth and throughput for machine learning workloads.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a multi-dimensional arbitration structure where different arbitration controllers operate at different levels (component-level, bank group-level, and memory bank-level). This dimensional expansion allows simultaneous handling of multiple access requests through different protocols (AXI, AHB, APB), resolving the bandwidth bottleneck without excessive complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If multiple hardware components access shared memory simultaneously, then computation parallelism is improved, but memory access conflicts and latency increase

Engineering Contradiction:
Improvecomputation parallelismVSAvoidmemory access latency
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

Multiple arbitration controllers act as intermediaries between hardware components and memory bank groups. Each arbitration controller receives access requests, determines priority based on protocols and components, and grants access rights to minimize conflicts. This intermediary layer coordinates simultaneous accesses, maintaining high computation parallelism while reducing memory access latency through intelligent arbitration.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system dynamically changes access parameters by assigning different protocols (AXI, AHB, APB) and priority levels to different hardware components based on their requirements. This parameter differentiation allows critical machine learning computations to receive higher priority access, reducing latency for time-sensitive operations while maintaining overall system parallelism.

Inventive Principle:
Principle #35Parameter changes

3Speed

If on-chip caches are used to reduce latency, then memory access speed is improved, but device cost and power consumption increase

Engineering Contradiction:
Improvememory access speedVSAvoidpower consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

Instead of implementing full on-chip caches for all memory access, the system uses selective arbitration that prioritizes and accelerates access for machine learning hardware components through dedicated protocols and direct memory bank group connections. This partial action approach provides cache-like performance benefits for critical workloads without the excessive power consumption and cost of comprehensive on-chip caching.

Inventive Principle:
Principle #16Partial or excessive action

4Productivity

If 3D packaging is used to increase memory bandwidth, then throughput is improved, but manufacturing complexity and cost increase

Engineering Contradiction:
Improvememory throughputVSAvoidmanufacturing complexity
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

The arbitration controllers are designed with multi-functionality, handling multiple protocols (AXI, AHB, APB) and managing different memory bank groups through unified control logic. This universal design achieves high throughput by efficiently coordinating access across diverse components without requiring complex 3D packaging, as the arbitration fabric can dynamically route and prioritize requests through existing 2D interconnect structures.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12417051B2High-performance on-chip memory controller
Publication Date: 2025.09.16 BLACK SESAME TECH INC
  • US12417051B2 patent drawing
  • US12417051B2 patent drawing
  • US12417051B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for controlling, by an on-chip memory controller, a plurality of hardware components that are configured to perform computations to access a shared memory. One of the on-chip memory controller includes at least one backside arbitration controller communicatively coupled with a memory bank group and a first hardware component, wherein the at least one backside arbitration controller is configured to perform bus arbitrations to determine whether the first hardware component can access the memory bank group using a first memory access protocol; and a frontside arbitration controller communicatively coupled with the memory bank group and a second hardware component, wherein the frontside arbitration controller is configured to perform bus arbitrations to determine whether the second hardware component can access the memory bank group using a second memory access protocol different from the first memory access protocol.