Fused Bitwise and FPL Data Path for Multi-Type Neural Network Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic devices require multiple hardware accelerators to process different data types, leading to increased size, manufacturing cost, and power consumption due to the need for various data types in neural networks like CNN, BNN, and TNN.

Innovation Solution

An electronic device is designed to compute inner products on binary, ternary, non-binary, and non-ternary data using a fused bitwise data path and a Full Precision Layer (FPL) data path, combining these paths to form a Processing Element (PE) that supports both bitwise and full precision operations without additional storage overhead.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple hardware accelerators are implemented to process different data types (binary, ternary, non-binary, non-ternary), then the processing capability and versatility are improved, but the device size, manufacturing cost, and power consumption increase

Engineering Contradiction:
Improveprocessing capability for different data typesVSAvoiddevice size
Core Design Contradiction:
Adaptability or versatilityVSArea of stationary object

Solution Approach 1:

The patent implements a universal processing element that can handle multiple data types (binary, ternary, non-binary, non-ternary) through a single hardware accelerator. This is achieved by designing a multi-functional architecture that dynamically adapts to different data types, eliminating the need for separate dedicated accelerators for each data type while maintaining comprehensive processing capability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges multiple data type processing capabilities into a single hardware accelerator by combining binary processing units, ternary processing units, and full precision layer units into one integrated structure. This consolidation reduces the overall device size while preserving the ability to process all four data types through shared hardware resources

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If multiple hardware accelerators are implemented to process different data types, then the processing capability is improved, but the manufacturing cost increases

Engineering Contradiction:
Improveprocessing capability for different data typesVSAvoidmanufacturing cost
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The universal processing element reduces manufacturing cost by consolidating multiple specialized accelerators into a single multi-functional unit. This approach decreases the total component count, simplifies the manufacturing process, and reduces assembly complexity while maintaining the ability to process all data types through configurable hardware logic

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

By merging binary, ternary, and full precision processing capabilities into one integrated hardware accelerator, the patent reduces the number of separate components that need to be manufactured and assembled. This consolidation directly lowers manufacturing costs through reduced material usage, simpler production workflows, and decreased assembly operations

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If multiple hardware accelerators are implemented to process different data types, then the processing capability is improved, but the power consumption increases

Engineering Contradiction:
Improveprocessing capability for different data typesVSAvoidpower consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by stationary object

Solution Approach 1:

The universal processing element reduces power consumption by enabling a single hardware accelerator to dynamically switch between different data type processing modes. This eliminates the need to operate multiple separate accelerators simultaneously, reducing overall power draw while maintaining the capability to process all four data types as needed through configurable operational states

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Area of stationary object

If a single hardware accelerator processes all data types, then the device size and power consumption are reduced, but the processing speed and efficiency may decrease

Engineering Contradiction:
Improvedevice sizeVSAvoidprocessing speed
Core Design Contradiction:
Area of stationary objectVSSpeed

Solution Approach 1:

The patent employs dynamic configuration within the universal processing element, allowing the hardware accelerator to adapt its internal structure and processing mode based on the input data type. This dynamic reconfiguration enables the single accelerator to optimize its processing speed for each specific data type (binary, ternary, non-binary, non-ternary) while maintaining a compact device size, rather than being locked into a fixed architecture

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12039430B2Electronic device and method for inference binary and ternary neural networks
Publication Date: 2024.07.16 SAMSUNG ELECTRONICS CO LTD
  • US12039430B2 patent drawing
  • US12039430B2 patent drawing
  • US12039430B2 patent drawing

AI summary

A method for computing an inner product on a binary data, a ternary data, a non-binary data, and a non-ternary data using an electronic device. The method includes calculating the inner product on a ternary data, designing a fused bitwise data path to support the inner product calculation on the binary data and the ternary data, designing a FPL data path to calculate an inner product between one of the non-binary data and the non-ternary data and one of the binary data and the ternary data, and distributing the inner product calculation for the binary data and the ternary data and the inner product between one of the non-binary data and the non-ternary data and one of the binary data and the ternary data in the fused bitwise data path and the FPL data path.