Speech Processing Trigger for Targeted Error Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech processing systems face inefficiencies in identifying and correcting incorrect speech recognition results, leading to unnecessary processing of correctly recognized inputs and resource misallocation during training, as they review all data rather than focusing on specific incorrect results.

Innovation Solution

A method is introduced where a user can trigger the system to identify incorrect speech processing results using a special audio indicator, allowing the system to prioritize and focus training efforts on these specific errors, thereby avoiding the review of correctly processed data and improving overall efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the system reviews all speech processing data to identify incorrect results, then comprehensive error detection is achieved, but processing time and computational resources are wasted on correctly processed data

Engineering Contradiction:
Improveerror detection completenessVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the speech processing data stream into individual speech segments, each associated with a confidence score. The system selectively processes only those segments with confidence scores below a threshold, rather than reviewing all segments. This segmentation enables targeted error detection while avoiding unnecessary processing of correct segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by differentiating processing intensity based on segment characteristics. High-confidence segments receive minimal or no review, while low-confidence segments undergo detailed analysis. This localized approach ensures thorough error detection where needed while conserving resources on segments likely to be correct.

Inventive Principle:
Principle #3Local quality

2Stability of the object's composition

If the system processes all speech data uniformly, then consistent processing is maintained, but computational resources are misallocated to data that does not need correction

Engineering Contradiction:
Improveprocessing consistencyVSAvoidcomputational resource utilization
Core Design Contradiction:
Stability of the object's compositionVSUse of energy by moving object

Solution Approach 1:

The patent implements dynamic processing where the system adjusts its review strategy based on segment confidence scores. Instead of uniform processing, the system dynamically selects which segments require detailed analysis, allocating computational resources proportionally to the likely need for correction. This dynamic approach maintains processing consistency for high-confidence segments while intensifying analysis where errors are more probable.

Inventive Principle:
Principle #15Dynamics

3Productivity

If the system focuses training on specific incorrect results only, then training efficiency improves, but the ability to detect subtle errors across all data may be reduced

Engineering Contradiction:
Improvetraining efficiencyVSAvoiderror detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by using confidence scores as a pre-filter before detailed error analysis. Segments with low confidence scores are identified in advance as potential errors, allowing the training system to focus on these pre-selected candidates. This preliminary identification maintains high training efficiency while preserving the ability to detect subtle errors within the targeted segments.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9916826B1Targeted detection of regions in speech processing data streams
Publication Date: 2018.03.13 AMAZON TECH INC
  • US9916826B1 patent drawing
  • US9916826B1 patent drawing
  • US9916826B1 patent drawing

AI summary

In speech processing systems, a special audio trigger indication is configured to efficiently isolate and mark incorrect speech processing results. The trigger indication may be configured to be easily recognizable by a speech processing device under various ASR and acoustic conditions. Once a speech processing device recognizes the trigger indication, incorrectly processed speech processing results are marked and may be isolated and prioritized for review by training and upgrading processes.