Neural Network Keyphrase Detection Without DSP
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current keyphrase detection systems are resource-intensive, requiring complex computations and high power consumption due to their reliance on digital signal processors (DSPs), and they often need retraining for new keyphrases, while separate speech/non-speech detection modules increase computational complexity and power usage.
Innovation Solution
A neural network keyphrase detection system that uses a recursive neural network decoder to perform sub-phonetic-based detection with rejection modeling, eliminating the need for DSPs and integrating speech/non-speech detection, allowing for efficient processing on neural network accelerators.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If Gaussian mixture models (GMMs) are used to model the acoustics of keyphrase variations, then detection accuracy is improved, but device complexity and power consumption increase
Solution Approach 1:
The patent extracts only the essential acoustic features needed for keyphrase detection by using phone-based scoring and likelihood ratios, eliminating the need for complex GMMs while maintaining detection accuracy. This reduces model complexity and computational requirements.
Solution Approach 2:
The patent replaces expensive, complex GMM models with simpler, more computationally efficient models that can be quickly trained and deployed. The simpler models use basic acoustic scoring and likelihood ratio computations that are much less resource-intensive than GMMs.
2Measurement precision
If dense keyphrase decoding computations are performed on DSPs, then detection capability is improved, but power consumption increases
Solution Approach 1:
The patent extracts only the essential computational operations needed for keyphrase detection by using phone-based acoustic scoring and likelihood ratio computations, eliminating dense matrix operations typical of GMM-based systems. This reduces computational density and power consumption.
Solution Approach 2:
The patent replaces the mechanical DSP computation system with a neural network-based system that performs acoustic scoring and likelihood ratio computations. This substitution enables more efficient processing with lower power consumption while maintaining detection capability.
3Measurement precision
If separate speech/non-speech detection module is used, then speech detection accuracy is improved, but device complexity and power consumption increase
Solution Approach 1:
The patent merges speech/non-speech detection functionality directly into the keyphrase detection system by using the same acoustic scoring and likelihood ratio computations. This integration eliminates the need for a separate module while maintaining speech detection accuracy through the unified probabilistic framework.
Solution Approach 2:
The patent creates a universal keyphrase detection system that simultaneously performs both keyphrase recognition and speech/non-speech detection using the same acoustic model and likelihood ratio computations. This multi-functional approach reduces system complexity while maintaining both capabilities.
Data Source
AI summary
A method and system are directed to autonomous neural network keyphrase detection and includes generating and using a multiple element state score vector by using neural network operations and without substantial use of a digital signal processor (DSP) to perform the keyphrase detection.


