Scalable Neural Processor Circuit with Selective Engine Activation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in efficiently processing neural network tasks due to high bandwidth and power consumption when relying solely on central processing units (CPUs).

Innovation Solution

A scalable neural processor circuit with multiple engine circuits that can be selectively activated or deactivated for parallel processing, comprising a data buffer and a kernel extract circuit, to support the instantiation and execution of neural networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a central processing unit (CPU) and main memory are used to instantiate and execute machine learning systems, then ease of configuration is improved, but power consumption and CPU bandwidth increase significantly

Engineering Contradiction:
Improveease of configurationVSAvoidpower consumption
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The neural network processing functionality is segmented into multiple independent neural engine circuits (first neural engine circuit, second neural engine circuit, etc.) that can operate in parallel. Each engine circuit processes specific portions of the neural network computations, allowing the system to divide the computational workload and reduce the burden on any single component while maintaining ease of configuration through the task manager circuit.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A data buffer circuit is introduced as an intermediary between the neural engine circuits and the CPU/main memory. This data buffer stores input data and output data, reducing the bandwidth requirements between the CPU and neural engines. The data buffer acts as a mediator that decouples the computational engines from the memory subsystem, allowing the CPU to maintain ease of configuration while reducing overall system power consumption.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If a central processing unit (CPU) and main memory are used to instantiate and execute machine learning systems, then ease of configuration is improved, but CPU bandwidth is consumed significantly

Engineering Contradiction:
Improveease of configurationVSAvoidCPU bandwidth
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

The processing workload is segmented across multiple neural engine circuits that operate independently in parallel. Each engine circuit handles specific computational tasks, distributing the data processing load and reducing the aggregate bandwidth requirement on the CPU. The task manager circuit coordinates these segmented operations, maintaining ease of configuration while reducing CPU bandwidth consumption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The data buffer circuit serves as an intermediary that caches input data and output data locally. This reduces the amount of data that needs to be transferred through the CPU and main memory interface, thereby reducing CPU bandwidth requirements. The data buffer mediates between the neural engines and the memory subsystem, allowing configuration to remain easy while consuming less CPU bandwidth.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Speed

If multiple neural engine circuits are used for parallel processing, then processing speed is improved, but device complexity increases

Engineering Contradiction:
Improveprocessing speedVSAvoiddevice complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

Multiple neural engine circuits are designed with identical or highly similar architectures, where each circuit can perform the same set of operations (convolution, multiplication, accumulation). This universality allows the system to achieve parallel processing speedup while managing complexity through replication of standardized modules. The task manager circuit provides unified control for all engines, simplifying the overall system architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The neural network processing task is segmented into multiple independent units that can be executed by different neural engine circuits in parallel. Each engine circuit processes a specific portion of the computation, and the results are aggregated by the data buffer and task manager. This segmentation enables speedup through parallelism while keeping individual circuit complexity manageable through modular design.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250165747A1Scalable neural network processing engine
Publication Date: 2025.05.22 APPLE INC
  • US20250165747A1 patent drawing
  • US20250165747A1 patent drawing
  • US20250165747A1 patent drawing

AI summary

Embodiments relate to a neural processor circuit with scalable architecture for instantiating one or more neural networks. The neural processor circuit includes a data buffer coupled to a memory external to the neural processor circuit, and a plurality of neural engine circuits. To execute tasks that instantiate the neural networks, each neural engine circuit generates output data using input data and kernel coefficients. A neural processor circuit may include multiple neural engine circuits that are selectively activated or deactivated according to configuration data of the tasks. Furthermore, an electronic device may include multiple neural processor circuits that are selectively activated or deactivated to execute the tasks.