Sparse Phonetic Impulse Convolution for Low-Overhead Keyword Spotting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current spoken term detection systems require high computational overhead and external resources for processing large vocabulary databases, making them unsuitable for devices with limited processing resources, and struggle with out-of-vocabulary keywords and user variability.
Innovation Solution
A system that decomposes speech signals into a sparse set of phonetic impulses and convolves them with an ensemble of filters based on keyword models stored locally, allowing for rapid and efficient keyword identification without external resources, using Poisson rate parameters for keyword modeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If large vocabulary databases and lattice searching methods are used for spoken term detection, then recognition performance is improved, but computational overhead and processing time increase significantly
Solution Approach 1:
The patent segments the continuous speech signal into discrete phonetic events using a phonetic event detector. This segmentation transforms the continuous speech recognition problem into a discrete event detection problem, enabling efficient keyword spotting without requiring comprehensive lattice generation for the entire speech signal. The segmentation principle resolves the contradiction by processing only relevant portions of the speech signal.
Solution Approach 2:
The patent extracts only the essential phonetic events from the speech signal that are relevant to keyword detection, rather than processing the entire speech signal through comprehensive lattice searching. By extracting and focusing on discrete phonetic events, the system achieves fast keyword detection without the computational overhead of generating and searching through complete speech lattices.
2Measurement precision
If comprehensive lattice searching is performed for keyword detection, then detection accuracy is improved, but computational resources and processing time are excessively consumed
Solution Approach 1:
The system uses locally stored keyword models and phonetic event detection algorithms that operate independently without requiring external server resources. The device performs self-service keyword detection using its own computational resources, eliminating the need for cloud-based processing while maintaining detection accuracy through efficient local algorithm execution.
Solution Approach 2:
The patent changes the processing parameters from comprehensive lattice searching to discrete phonetic event detection. By transforming the problem from continuous signal processing to discrete event detection, the system achieves comparable detection accuracy with significantly reduced computational complexity and processing requirements.
3Power
If remote servers and network connectivity are used for speech processing, then processing power and vocabulary database access are improved, but system independence and processing speed are reduced
Solution Approach 1:
The system implements self-service by storing keyword models and phonetic event detection algorithms locally on the device. This enables the device to perform speech processing independently without requiring network connectivity or remote server access, maintaining both system independence and processing speed while utilizing local computational resources efficiently.
Data Source
AI summary
A system and method are provided for performing speech processing. A system includes an audio detection system configured to receive a signal including speech and a memory having stored therein a database of keyword models forming an ensemble of filters associated with each keyword in the database. A processor is configured to receive the signal including speech from the audio detection system, decompose the signal including speech into a sparse set of phonetic impulses, and access the database of keywords and convolve the sparse set of phonetic impulses with the ensemble of filters. The processor is further configured to identify keywords within the signal including speech based a result of the convolution and control operation the electronic system based on the keywords identified.


