Neuromorphic Voice Activation System Using Hierarchical Spiking Neural Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems in portable devices face high power consumption, time delay, and computational costs due to continuous operation and difficulty in distinguishing voice commands from background noise, especially in noisy environments.
Innovation Solution
A neuromorphic voice activation system utilizing a hierarchical arrangement of first and second spiking neural networks, where the first network detects voice activity and activates the second network only when necessary, minimizing power consumption and rejecting background noise through selective bandwidth rejection in an artificial cochlea.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional speech recognition systems operate continuously to recognize voice commands, then voice command recognition capability is improved, but power consumption increases significantly
Solution Approach 1:
The speech recognition system is divided into multiple independent spiking neural networks (SPNNs) organized in a hierarchical structure. Each SPNN processes specific aspects of audio signals (voice activity detection, speaker identification, command recognition), allowing the system to segment computational tasks and activate only the necessary networks based on detected voice patterns, thereby reducing overall power consumption while maintaining recognition capability.
Solution Approach 2:
The system dynamically adjusts its operational state by transitioning between different processing modes. The hierarchical SPNN architecture allows the system to remain in a low-power sleep state during silence and only activate relevant neural networks when voice activity is detected, creating a dynamic power consumption profile that adapts to actual usage conditions rather than operating continuously at full capacity.
2Speed
If speech recognition processes audio signals in real-time, then response speed is improved, but computational cost and power consumption increase
Solution Approach 1:
The system replaces conventional software-based speech recognition algorithms with hardware-implemented spiking neural networks that use event-driven architecture. This substitution enables real-time processing by utilizing the natural timing of neural spikes to trigger computations only when needed, eliminating the need for continuous computational cycles and significantly reducing power consumption while maintaining fast response speeds.
Solution Approach 2:
The spiking neural networks operate using periodic action principles where neurons fire spikes only when triggered by input signals above threshold levels. This periodic activation pattern allows the system to process audio signals in real-time through event-driven computation, activating processing only when voice activity occurs rather than maintaining continuous computational load, thus reducing overall power consumption.
3Reliability
If the system processes audio from multiple microphones to improve noise filtering, then voice command accuracy in noisy environments is improved, but processing time and power consumption increase
Solution Approach 1:
The audio processing task is segmented across multiple SPNNs that each handle specific functions: some networks process signals from individual microphones while others integrate information from multiple sources. This segmentation allows parallel processing of audio streams from different microphones simultaneously, improving noise filtering accuracy without requiring sequential processing that would increase time delay.
Solution Approach 2:
The system performs preliminary voice activity detection using SPNNs that analyze audio signals from multiple microphones before activating more complex processing networks. This preliminary action filters out background noise and identifies voice presence early in the processing chain, allowing subsequent processing stages to focus only on relevant audio segments, thereby reducing overall processing time while maintaining accuracy in noisy environments.
Data Source
AI summary
The present invention provides a system and method for controlling a device by recognizing voice commands through a spiking neural network. The system comprises a spiking neural adaptive processor receiving an input stream that is being forwarded from a microphone, a decimation filter and then an artificial cochlea. The spiking neural adaptive processor further comprises a first spiking neural network and a second spiking neural network. The first spiking neural network checks for voice activities in output spikes received from artificial cochlea. If any voice activity is detected, it activates the second spiking neural network and passes the output spike of the artificial cochlea to the second spiking neural network that is further configured to recognize spike patterns indicative of specific voice commands. If the first spiking neural network does not detect any voice activity, it halts the second spiking neural network.


