Speech Characterization for ASR Grammar Enablement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice-controlled machines fail to effectively utilize speaker and utterance characteristics, such as age, gender, accent, and mood, in interpreting user speech, leading to inadequate user-specific responses and security controls.

Innovation Solution

Implementing speech characterization to condition automatic speech recognition and natural language processing, using speaker identification and utterance classification to influence grammar rules and statistical language models, thereby enabling more accurate and context-aware interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech characterization is implemented to condition ASR and NLP, then accuracy and context-relevance improve, but device complexity increases

Engineering Contradiction:
Improvespeech interpretation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs speech characterization (extracting speaker attributes like age, gender, accent, mood) before conditioning the ASR and NLP processing. This preliminary classification enables the system to pre-configure grammar rules and language models based on speaker characteristics, improving accuracy while managing complexity through staged processing

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies different grammar rules and language model parameters tailored to specific speaker characteristics. For example, different grammar configurations are used for different age groups, accents, or moods, allowing the system to optimize processing for each speaker type rather than using a one-size-fits-all approach

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If speech characterization is implemented to condition ASR and NLP, then user-specific responses improve, but processing time increases

Engineering Contradiction:
Improveuser-specific response capabilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

Speaker characterization is performed in advance to establish speaker profiles that include attributes like age, gender, accent, and mood. These pre-computed characteristics are then used to quickly condition the ASR and NLP processing, avoiding the need to analyze speech characteristics in real-time during each interaction

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts processing parameters based on speaker characteristics. For example, vocabulary lists, grammar rules, and language model weights are modified according to the speaker's accent, age group, or detected mood, enabling adaptive processing without requiring complete re-analysis of the speech signal

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If comprehensive speech characterization is performed, then interpretation accuracy improves, but computational power consumption increases

Engineering Contradiction:
Improveutterance classification accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system performs speech characterization to extract only the most relevant features for conditioning ASR and NLP, such as speaker age group, gender, accent, and mood. Rather than analyzing all possible speech attributes, the system focuses on key characteristics that have the greatest impact on interpretation accuracy, reducing unnecessary computational overhead

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10347245B2Natural language grammar enablement by speech characterization
Publication Date: 2019.07.09 SOUNDHOUND AI IP LLC
  • US10347245B2 patent drawing
  • US10347245B2 patent drawing
  • US10347245B2 patent drawing

AI summary

Either or both of voice speaker identification or utterance classification such as by age, gender, accent, mood, and prosody characterize speech utterances in a system that performs automatic speech recognition (ASR) and natural language processing (NLP). The characterization conditions NLP, either through application to interpretation hypotheses or to specific grammar rules. The characterization also conditions language models of ASR. Conditioning may comprise enablement and may comprise reweighting of hypotheses.