Audio Keyword Search Phoneme Filtering for Low Power Devices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio keyword search techniques are computationally intensive and cannot be efficiently implemented on devices with limited processing power, such as mobile computing devices, due to the high processing demand of automatic speech recognition systems.

Innovation Solution

An audio keyword search system that learns a classifier function on posteriorgrams using machine learning techniques, allowing for the efficient detection of keywords in audio recordings by building a binary model to identify the presence of a set of keywords, which can be integrated into existing ASR systems without modifying their internals, thereby reducing the computational load.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a full automatic speech recognition process is implemented for keyword search, then keyword detection accuracy is improved, but processing power consumption increases significantly

Engineering Contradiction:
Improvekeyword detection accuracyVSAvoidprocessing power consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent divides the keyword search process into two segments: a fast filtering stage using phoneme-level acoustic models to quickly identify candidate segments, and a subsequent detailed recognition stage. This segmentation allows the system to avoid running the full computationally intensive ASR pipeline on all audio segments, thereby reducing overall processing power consumption while maintaining keyword detection accuracy for the identified candidates.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by implementing a simplified acoustic model that operates at phoneme level rather than full word-level recognition. This partial recognition approach processes only the essential phonetic features needed for keyword detection, avoiding the excessive computational resources required for complete speech-to-text conversion, thus reducing processing power consumption while maintaining sufficient accuracy for keyword search.

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If a computationally intensive multi-stage ASR process is used, then speech recognition accuracy is improved, but device performance is substantially hindered

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiddevice performance
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the ASR process into a fast phoneme-level filtering stage and a selective detailed recognition stage. The phoneme-level acoustic model quickly processes audio segments to identify candidates containing target keywords, while the full multi-stage ASR is applied only to these candidates. This segmentation maintains speech recognition accuracy for keywords while preserving device performance by avoiding unnecessary processing of non-relevant segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary action through a voice activity detector and phoneme-level acoustic model that pre-process and filter audio segments before they reach the full ASR pipeline. This preliminary filtering identifies and isolates segments containing potential keywords, allowing the computationally intensive multi-stage ASR to focus only on relevant portions, thereby maintaining accuracy while improving overall device performance.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If the ASR process is applied to all audio segments, then keyword detection completeness is improved, but processing time increases

Engineering Contradiction:
Improvekeyword detection completenessVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the audio processing into rapid phoneme-level analysis and selective detailed processing. The phoneme-level acoustic model quickly scans through all audio segments to identify candidates, maintaining detection completeness for keywords while dramatically reducing processing time compared to applying full ASR to every segment. The segmented approach ensures no keyword-containing segments are missed while minimizing overall processing time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary voice activity detection and phoneme-level acoustic modeling to all audio segments before detailed ASR processing. This preliminary action quickly identifies segments containing potential keywords, allowing the system to maintain detection completeness by covering all segments while reducing processing time by avoiding full ASR on non-relevant portions.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12020697B2Systems and methods for fast filtering of audio keyword search
Publication Date: 2024.06.25 RAYTHEON APPLIED SIGNAL TECHNOLOGY INC
  • US12020697B2 patent drawing
  • US12020697B2 patent drawing
  • US12020697B2 patent drawing

AI summary

An audio keyword searcher arranged to identify a voice segment of a received audio signal; identify, by an automatic speech recognition engine, one or more phonemes included in the voice segment; output, from the automatic speech recognition engine, the one or more phonemes to a keyword filter to detect whether the voice segment includes any of the one or more first keywords of the first keyword list and, if detected, output the one or more phonemes included in the voice segment to a decoder but, if not detected, not output the one or more phonemes included in the voice segment to the decoder. If the one or more phonemes are output to the decoder: generate a word lattice associated with the voice segment; search the word lattice for one or more second keywords, and determine whether the voice segment includes the one or more second keywords.