Voice Control Acoustic Recognition Parallel Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems for consumer electronics face challenges in achieving a balance between false reject and false accept rates, especially in noisy environments, and require manual triggers or internet access, limiting their hands-free functionality.

Innovation Solution

The system processes acoustic input signals using multiple recognition processes that start at different time points, with scores calculated across various states and offsets to improve recognition accuracy, allowing for hands-free operation and accurate word spotting.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition systems use a single recognition process with a fixed model, then the system complexity is low, but the recognition accuracy in noisy environments deteriorates

Engineering Contradiction:
Improverecognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the speech recognition process into multiple parallel recognition processes, each handling different segments or aspects of the acoustic input signal. Multiple acoustic models process the same input independently, and their results are combined to improve overall recognition accuracy while managing system complexity through modular design

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent combines results from multiple independent recognition processes through score integration. Each process generates confidence scores for recognized words, and these scores are merged to produce a final recognition result with higher accuracy than any single process could achieve alone

Inventive Principle:
Principle #5Merging (Combining)

2Ease of operation

If speech recognition systems process all acoustic input continuously, then hands-free operation is achieved, but the false accept rate increases in noisy environments

Engineering Contradiction:
Improvehands-free operationVSAvoidfalse accept rate
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements a preliminary voice trigger mechanism that activates the full recognition system only when a specific trigger word is detected. This preliminary action filters out ambient noise by requiring a deliberate activation command, reducing false accepts while maintaining hands-free operation capability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts its operation mode between trigger-based activation and continuous listening. The recognition process can switch between being dormant, actively listening for triggers, and fully engaged in command recognition, optimizing reliability based on operational context

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If speech recognition systems use complex acoustic models to improve accuracy, then word spotting precision improves, but the processing time increases

Engineering Contradiction:
Improveword spotting precisionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial processing by focusing recognition efforts on specific time segments where trigger words are likely to occur. Instead of uniformly processing the entire acoustic signal with full complexity, the system applies intensive processing only to relevant segments, reducing overall processing time while maintaining precision

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9484028B2Systems and methods for hands-free voice control and voice search
Publication Date: 2016.11.01 SENSORY INC
  • US9484028B2 patent drawing
  • US9484028B2 patent drawing
  • US9484028B2 patent drawing

AI summary

In one embodiment the present invention includes a method comprising receiving an acoustic input signal and processing the acoustic input signal with a plurality of acoustic recognition processes configured to recognize the same target sound. Different acoustic recognition processes start processing different segments of the acoustic input signal at different time points in the acoustic input signal. In one embodiment, initial states in the recognition processes may be configured on each time step.