Context-Aware Speech Recognition Engine Adjustment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automated speech recognition engines face challenges in noisy environments and often misinterpret user commands due to lack of contextual understanding, leading to incorrect word or command recognition.

Innovation Solution

The system obtains contextual information about the user's activity, environment, and device state to adjust the automated speech recognition engine, biasing it towards more likely commands and words based on the context, thereby improving recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If automated speech recognition engine processes user speech input without contextual information, then device complexity is reduced, but speech recognition accuracy deteriorates in noisy environments

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by obtaining contextual information about the user's activity, environment, and device state before processing speech input. This contextual information is used to adjust the ASR engine's parameters and bias towards more likely commands, improving recognition accuracy before the speech input is even received.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes parameters of the ASR engine based on contextual information. Specifically, it adjusts the language model parameters, vocabulary weighting, and command probability distributions according to the detected context, allowing the engine to adapt its recognition behavior to different situations without increasing fundamental system complexity.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If automated speech recognition engine uses contextual information to adjust recognition, then speech recognition accuracy is improved, but processing time increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies partial action by selectively adjusting only the most relevant ASR parameters based on the detected context, rather than performing exhaustive analysis of all possible contextual factors. This allows the system to gain significant accuracy improvements while minimizing the time overhead of contextual processing.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If automated speech recognition engine is adjusted based on contextual information, then command recognition reliability is improved, but adaptability to different contexts is reduced

Engineering Contradiction:
Improvecommand recognition reliabilityVSAvoidcontext adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system implements dynamic adaptation by continuously monitoring contextual information and adjusting ASR engine parameters in real-time. The contextual analyzer dynamically updates the language model and vocabulary weighting based on current user activity, environment, and device state, allowing the system to adapt to different contexts while maintaining high reliability within each context.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11386886B2Adjusting speech recognition using contextual information
Publication Date: 2022.07.12 LENOVO SWITZERLAND INTERNATIONAL GMBH
  • US11386886B2 patent drawing
  • US11386886B2 patent drawing
  • US11386886B2 patent drawing

AI summary

An embodiment provides a method, including: obtaining, using a processor, contextual information relating to an information handling device; adjusting, using a processor, an automated speech recognition engine using the contextual information; receiving, at an audio receiver of the information handling device, user speech input; and providing, using a processor, recognized speech based on the user speech input received and the contextual information adjustment to the automated speech recognition engine. Other aspects are described and claimed.