Denominator Circuit for Neural Network Pooling Operations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Artificial neural networks (ANNs) face challenges in efficiently performing pooling operations, particularly in determining the number of valid elements covered by a kernel, which is crucial for accurate calculations but consumes significant CPU bandwidth and power.

Innovation Solution

A denominator circuit is introduced that determines the number of valid elements by calculating two series of numbers representing kernel coverage in horizontal and vertical directions, which are then matrix multiplied to obtain the denominator values, allowing for accurate pooling operations while reducing computational load on the CPU.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If pooling operations are performed using CPU and main memory, then flexibility in configuring machine learning systems is improved, but CPU bandwidth consumption and power consumption increase significantly

Engineering Contradiction:
Improveflexibility in configuring machine learning systemsVSAvoidCPU bandwidth consumption and power consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent divides the pooling operation into two separate circuits: a first circuit that determines the number of valid elements in the row direction, and a second circuit that determines the number of valid elements in the column direction. These circuits independently calculate row-wise and column-wise valid element counts, then combine results to determine total valid elements covered by the kernel, reducing CPU involvement in each individual operation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dedicated pooling operation circuits as intermediary hardware components between the CPU and main memory. These circuits handle the computationally intensive pooling operations locally, using specialized hardware to calculate valid element counts and perform pooling operations, thereby reducing the burden on the CPU and main memory bandwidth

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If pooling operations are performed using CPU, then ease of operation is maintained, but power consumption increases

Engineering Contradiction:
Improveease of operationVSAvoidpower consumption
Core Design Contradiction:
Ease of operationVSUse of energy by stationary object

Solution Approach 1:

The patent introduces dedicated pooling operation circuits as intermediary hardware components between the CPU and main memory. These circuits handle the computationally intensive pooling operations locally, using specialized hardware to calculate valid element counts and perform pooling operations, thereby reducing the burden on the CPU and main memory bandwidth

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The pooling operation circuits are self-contained units that perform pooling operations autonomously without requiring continuous CPU intervention. The circuits independently determine valid element counts in row and column directions, combine these results, and execute pooling operations using their own dedicated resources, making the system self-sufficient for pooling tasks

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20210319077A1Circuit for performing pooling operation in neural processor
Publication Date: 2021.10.14 APPLE INC
  • US20210319077A1 patent drawing
  • US20210319077A1 patent drawing
  • US20210319077A1 patent drawing

AI summary

Embodiments relate to a denominator circuit that determines the number of valid elements of a data surface covered by a kernel depending on various locations of the kernel relative to the data surface. The denominator circuit includes a first circuit and a second circuit that have the same structure. The first circuit receives numbers representing different horizontal locations of a reference point in the kernel and generates a first matrix with first output elements corresponding to the different horizontal locations. The second circuit receives numbers representing different vertical locations of a reference point in the kernel and generates a second matrix with second output elements corresponding to the different vertical locations. A matrix multiplication of the first matrix and the second matrix is performed to obtain an array of valid elements covered by the kernel.