RBF Neural Network Hardware with Virtual CAM and Pre-Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current hardware implementations of neural network algorithms, such as Radial Basis Function (RBF) and k-Nearest Neighbor (kNN), face challenges in supporting probabilistic computations, multiple data types, and high-speed operations in multi-user and multi-purpose environments, lacking efficient pre- and post-processing capabilities and scalable architectures.

Innovation Solution

The development of enhanced RBF/RCE/kNN based architectures with integrated pre- and post-processing hardware, support for probabilistic computations, and scalable neural network designs that include features like K-Means clustering, recommendation engines, and virtual Content-Addressable Memory (CAM) operations, along with improved data handling and aggregation mechanisms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If hardware implementations use traditional multi-layer perceptron approaches, then analog and spiking neuron models can be achieved, but support for probabilistic computations and multiple data types is limited

Engineering Contradiction:
Improvesupport for probabilistic computations and multiple data typesVSAvoidhardware architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal hardware architecture that can perform multiple neural network algorithms (RBF, RCE, kNN, K-Means) and support various data types (integer, floating-point, probabilistic) through a single unified structure. The neuron array and distance calculation units are designed to handle different computational modes without requiring separate dedicated hardware for each algorithm or data type, thereby achieving multi-functionality while controlling complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The hardware architecture employs dynamic reconfiguration capabilities where the same physical hardware can switch between different computational modes (e.g., RBF mode, RCE mode, kNN mode) and data type handling (integer arithmetic, floating-point operations, probabilistic computations) based on input requirements. This dynamic adaptability allows the system to optimize its operation for different tasks without permanent hardware changes.

Inventive Principle:
Principle #15Dynamics

2Productivity

If neural network hardware performs high-speed operations, then processing rate increases, but pre- and post-processing capabilities become insufficient

Engineering Contradiction:
Improveprocessing rateVSAvoidpre- and post-processing capabilities
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent incorporates dedicated pre-processing hardware units that perform data normalization, feature extraction, and input vector preparation before the main neural network computation. These pre-processing operations are executed in parallel with the high-speed neuron array operations, ensuring that data is ready for processing without becoming a bottleneck to the overall processing rate.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The architecture introduces intermediate buffer memory and control logic units that mediate between the high-speed computation core and the slower pre/post-processing operations. These intermediary components allow the fast neuron array to operate continuously while pre-processing and post-processing tasks are performed on subsequent data batches or results, maintaining high throughput without sacrificing operational ease.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If neural network systems are designed for multi-user and multi-purpose environments, then versatility increases, but scalability and data handling efficiency decrease

Engineering Contradiction:
Improvemulti-user and multi-purpose supportVSAvoidscalability and data handling efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent divides the neural network hardware into multiple independent neuron arrays and distance calculation units that can operate in parallel. Each segment can be independently configured for different users or purposes, allowing the system to scale by activating only the necessary segments. This modular segmentation maintains high data handling efficiency by avoiding the need to process all data through a single monolithic structure.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The architecture implements a hierarchical structure where smaller neuron arrays can be nested within larger array configurations. This nesting allows the system to efficiently handle different data scales - small datasets can use only the necessary nested sub-arrays, while larger datasets can activate the full hierarchical structure, thereby maintaining scalability and efficiency across multi-user and multi-purpose scenarios.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS9747547B2Hardware enhancements to radial basis function with restricted coulomb energy learning and/or k-Nearest Neighbor based neural network classifiers
Publication Date: 2017.08.29 IN2H2
  • US9747547B2 patent drawing
  • US9747547B2 patent drawing
  • US9747547B2 patent drawing

AI summary

A nonlinear neuron classifier comprising a neuron array including a plurality of neuron chips each including a plurality of neurons of variable length and variable depth, the chips processing input vectors of variable length and variable depth that are input into the classifier for comparison against vectors stored in the classifier, wherein an NSP flag is set for a plurality of the neurons to indicate that only that plurality of neurons is to participate in the vector calculations. A virtual content addressable memory flag is set for certain of the neuron chips to enable functions including fast readout of data from the chips. Results of vector calculations are aggregated for fast readout for a host computer interfacing with the classifier.