Neuromorphic Voice Activation System Using Hierarchical Spiking Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems in portable devices face high power consumption, time delay, and computational costs due to continuous operation and difficulty in distinguishing voice commands from background noise, especially in noisy environments.

Innovation Solution

A neuromorphic voice activation system utilizing a hierarchical arrangement of first and second spiking neural networks, where the first network detects voice activity and activates the second network only when necessary, minimizing power consumption and rejecting background noise through selective bandwidth rejection in an artificial cochlea.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional speech recognition systems operate continuously to recognize voice commands, then voice command recognition capability is improved, but power consumption increases significantly

Engineering Contradiction:
Improvevoice command recognition capabilityVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The speech recognition system is divided into multiple independent spiking neural networks (SPNNs) organized in a hierarchical structure. Each SPNN processes specific aspects of audio signals (voice activity detection, speaker identification, command recognition), allowing the system to segment computational tasks and activate only the necessary networks based on detected voice patterns, thereby reducing overall power consumption while maintaining recognition capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts its operational state by transitioning between different processing modes. The hierarchical SPNN architecture allows the system to remain in a low-power sleep state during silence and only activate relevant neural networks when voice activity is detected, creating a dynamic power consumption profile that adapts to actual usage conditions rather than operating continuously at full capacity.

Inventive Principle:
Principle #15Dynamics

2Speed

If speech recognition processes audio signals in real-time, then response speed is improved, but computational cost and power consumption increase

Engineering Contradiction:
Improveresponse speedVSAvoidcomputational cost
Core Design Contradiction:
SpeedVSPower

Solution Approach 1:

The system replaces conventional software-based speech recognition algorithms with hardware-implemented spiking neural networks that use event-driven architecture. This substitution enables real-time processing by utilizing the natural timing of neural spikes to trigger computations only when needed, eliminating the need for continuous computational cycles and significantly reducing power consumption while maintaining fast response speeds.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The spiking neural networks operate using periodic action principles where neurons fire spikes only when triggered by input signals above threshold levels. This periodic activation pattern allows the system to process audio signals in real-time through event-driven computation, activating processing only when voice activity occurs rather than maintaining continuous computational load, thus reducing overall power consumption.

Inventive Principle:
Principle #19Periodic action

3Reliability

If the system processes audio from multiple microphones to improve noise filtering, then voice command accuracy in noisy environments is improved, but processing time and power consumption increase

Engineering Contradiction:
Improvevoice command accuracy in noisy environmentsVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The audio processing task is segmented across multiple SPNNs that each handle specific functions: some networks process signals from individual microphones while others integrate information from multiple sources. This segmentation allows parallel processing of audio streams from different microphones simultaneously, improving noise filtering accuracy without requiring sequential processing that would increase time delay.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary voice activity detection using SPNNs that analyze audio signals from multiple microphones before activating more complex processing networks. This preliminary action filters out background noise and identifies voice presence early in the processing chain, allowing subsequent processing stages to focus only on relevant audio segments, thereby reducing overall processing time while maintaining accuracy in noisy environments.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10157629B2Low power neuromorphic voice activation system and method
Publication Date: 2018.12.18 BRAINCHIP INC
  • US10157629B2 patent drawing
  • US10157629B2 patent drawing
  • US10157629B2 patent drawing

AI summary

The present invention provides a system and method for controlling a device by recognizing voice commands through a spiking neural network. The system comprises a spiking neural adaptive processor receiving an input stream that is being forwarded from a microphone, a decimation filter and then an artificial cochlea. The spiking neural adaptive processor further comprises a first spiking neural network and a second spiking neural network. The first spiking neural network checks for voice activities in output spikes received from artificial cochlea. If any voice activity is detected, it activates the second spiking neural network and passes the output spike of the artificial cochlea to the second spiking neural network that is further configured to recognize spike patterns indicative of specific voice commands. If the first spiking neural network does not detect any voice activity, it halts the second spiking neural network.