Pooling Chip With Dynamic Memory Output Modes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing hardware accelerators for pooling operations in neural networks either prioritize speed at the cost of versatility or vice versa, failing to balance both effectively.

Innovation Solution

A chip design incorporating a demultiplexer, dual memories for parallel and serial output, and a computation circuit that performs pooling operations efficiently, allowing for high-speed and versatile pooling operations by managing matrix outputs and computations effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If accelerators are designed to accelerate all pooling operations, then the overall operation speed is improved, but the versatility is reduced

Engineering Contradiction:
Improveoperation speedVSAvoidversatility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic configuration of memory output modes (parallel or serial) based on the specific pooling operation requirements. The system can adaptively switch between different output modes to match different pooling scenarios, thereby maintaining both high speed and versatility. This is achieved through control logic that selects the appropriate memory output mode according to the pooling operation parameters.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the operational parameters of the memory system by allowing it to switch between parallel output mode and serial output mode. This parameter change enables the same hardware infrastructure to support different pooling operation types and sizes, resolving the contradiction between speed optimization and versatility.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If accelerators are designed to increase operation speed of a particular pooling operation, then the operation speed is improved, but the versatility is reduced

Engineering Contradiction:
Improveoperation speedVSAvoidversatility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent designs a universal pooling accelerator that can handle multiple types of pooling operations (max pooling, average pooling, different pool sizes, different strides) through a single hardware architecture. The memory system's ability to switch between parallel and serial output modes provides multi-functionality, allowing the accelerator to adapt to various pooling operation requirements without sacrificing speed optimization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If dual memories with parallel and serial output are used, then both speed and versatility are maintained, but the device complexity increases

Engineering Contradiction:
Improveoperation speedVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges the functionality of multiple specialized memory units into a unified memory system that can operate in different modes. Instead of having separate parallel-output memory and serial-output memory units, the system combines them into a single memory infrastructure with reconfigurable output capabilities, thereby reducing overall device complexity while maintaining both speed and versatility.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20230267173A1Chip, Method, Accelerator, and System for Pooling Operation
Publication Date: 2023.08.24 SHENZHEN CORERAIN TECH CO LTD
  • US20230267173A1 patent drawing
  • US20230267173A1 patent drawing
  • US20230267173A1 patent drawing

AI summary

Disclosed are a chip, a method, an accelerator, and a system for pooling operation. The chip includes: a demultiplexer including a first input terminal, and first and second output terminals, and outputting a first matrix from the first input terminal via the first or second output terminals in response to a first control signal; a first memory connected to the first output terminal and outputting elements of the first matrix stored by the first memory in response to a second control signal; a second memory connected to the second output terminal and serially outputting elements of a second matrix in the first matrix stored by the second memory in response to a third control signal; and a computation circuit performing a pooling operation on the second matrix from the first memory or the second memory to obtain an operation result in response to a fourth control signal.