SoC Deep Learning Accelerator with Segmented RAM for Energy Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing integrated circuit devices face challenges in efficiently processing artificial neural networks (ANNs) due to high energy consumption and prolonged computation times.

Innovation Solution

The integration of a deep learning accelerator (DLA) with random access memory (RAM) in an integrated circuit device, featuring separate memory access connections and a camera interface for direct image data input, enables efficient parallel processing and reduced energy consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If a deep learning accelerator is integrated with random access memory in an integrated circuit device, then energy consumption is reduced and computation time is decreased, but device complexity increases

Engineering Contradiction:
Improveenergy consumptionVSAvoiddevice complexity
Core Design Contradiction:
Use of energy by moving objectVSDevice complexity

Solution Approach 1:

The patent combines the deep learning accelerator and random access memory into a single integrated circuit device, merging previously separate components (processor, memory, and deep learning processing units) into one unified system. This integration reduces the number of external connections and data transfers between separate components, thereby reducing energy consumption while accepting increased internal device complexity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The integrated circuit device is segmented into distinct functional units including a processor, deep learning processing units with specialized architectures, and random access memory. Each segment is optimized for specific tasks (general-purpose processing vs. deep learning computations), allowing efficient parallel operation and reduced energy consumption for deep learning workloads while managing device complexity through modular design

Inventive Principle:
Principle #1Segmentation

2Loss of time

If separate memory access connections are implemented, then computation time is reduced through parallel processing, but device complexity increases

Engineering Contradiction:
Improvecomputation timeVSAvoiddevice complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The memory access system is segmented into multiple independent connection paths, allowing different data streams to access memory simultaneously through separate channels. This parallel access capability reduces computation time for deep learning operations that require frequent data retrieval, while the segmented architecture manages complexity by dedicating specific paths to specific functions

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces memory controllers as intermediary components that manage and coordinate access to the random access memory through multiple connection paths. These controllers act as mediators between the deep learning processing units and memory, enabling parallel access while managing the complexity of coordinating multiple simultaneous memory operations

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If a camera interface is integrated for direct image data input, then productivity is improved through autonomous processing, but device complexity increases

Engineering Contradiction:
Improveprocessing throughputVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The camera interface is merged directly into the integrated circuit device, creating a unified system where image data can be captured and processed without being transferred to external devices. This integration enables autonomous processing of image data through the deep learning accelerator, improving productivity by eliminating external communication overhead while accepting increased device complexity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The integrated circuit device is designed to be self-sufficient by incorporating the camera interface, deep learning accelerator, and memory within a single device. This self-service capability allows the device to autonomously capture, process, and analyze image data without requiring external processing resources, thereby improving productivity through autonomous operation

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250117659A1System on a chip with deep learning accelerator and random access memory
Publication Date: 2025.04.10 MICRON TECHNOLOGY INC
  • US20250117659A1 patent drawing
  • US20250117659A1 patent drawing
  • US20250117659A1 patent drawing

AI summary

Systems, devices, and methods related to a deep learning accelerator and memory are described. An integrated circuit may be configured with: a central processing unit, a deep learning accelerator configured to execute instructions with matrix operands; random access memory configured to store first instructions of an artificial neural network executable by the deep learning accelerator and second instructions of an application executable by the central processing unit; one or connections among the random access memory, the deep learning accelerator and the central processing unit; and an input/output interface to an external peripheral bus. While the deep learning accelerator is executing the first instructions to convert sensor data according to the artificial neural network to inference results, the central processing unit may execute the application that uses inference results from the artificial neural network.