Dual Speech Recognizers for Simultaneous OS and App Command Interpretation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems struggle to interpret global commands and application commands simultaneously, often leading to conflicts due to overlapping commands, phonetic similarities, and incompatibility of speech technologies, which results in ambiguity and a cumbersome user experience.

Innovation Solution

A speech-enabled system that employs two speech recognizers operating simultaneously to interpret operating system commands and application commands, using unrestricted speech grammars and natural language processing, with reserved words or cadence to differentiate between system and application commands, allowing users to speak freely without the need for a push-to-talk button.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a single speech recognizer is used, then the system is simple, but it cannot interpret both operating system commands and application commands simultaneously

Engineering Contradiction:
Improvecommand interpretation capabilityVSAvoidsystem structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent divides the speech recognition system into separate components: a first speech recognizer for operating system commands and a second speech recognizer for application commands. This segmentation allows each recognizer to specialize in specific command types, enabling simultaneous interpretation of both system and application commands without interference.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a universal command routing mechanism that directs speech input to the appropriate recognizer based on context. The system can handle multiple command types through a single integrated architecture, where the routing logic determines whether to send commands to the OS recognizer, application recognizer, or both simultaneously.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If speech recognizers operate simultaneously, then command interpretation is accurate, but ambiguity arises from overlapping commands and phonetic similarities

Engineering Contradiction:
Improvecommand recognition accuracyVSAvoidcommand disambiguation
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent introduces a command routing mechanism as an intermediary between the speech input and the recognizers. This mediator analyzes the speech input and determines the appropriate target (OS command, application command, or both), preventing direct conflicts between recognizers and enabling accurate simultaneous operation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system incorporates feedback mechanisms where the routing logic continuously monitors speech input characteristics and adjusts the distribution of commands to appropriate recognizers. This feedback loop resolves ambiguities by adapting to the context of spoken commands in real-time.

Inventive Principle:
Principle #23Feedback

3Ease of operation

If unrestricted speech grammars are used, then natural language processing is improved, but command differentiation between system and application becomes difficult

Engineering Contradiction:
Improvespeech input flexibilityVSAvoidcommand differentiation
Core Design Contradiction:
Ease of operationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies different speech grammar constraints to different recognizers based on their specific functions. The first recognizer uses grammar optimized for operating system commands while the second uses grammar for application commands. This local differentiation maintains unrestricted natural language input while enabling clear command differentiation through context-specific grammar rules.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10186262B2System with multiple simultaneous speech recognizers
Publication Date: 2019.01.22 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10186262B2 patent drawing
  • US10186262B2 patent drawing
  • US10186262B2 patent drawing

AI summary

A speech recognition system interprets both spoken system commands as well as application commands. Users may speak commands to an open microphone of a computing device that may be interpreted by at least two speech recognizers operating simultaneously. The first speech recognizer interprets operating system commands and the second speech recognizer interprets application commands. The system commands may include at least opening and closing an application and the application commands may include at least a game command or navigation within a menu. A reserve word may be used to identify whether the command is for the operation system or application. A user's cadence may also indicate whether the speech is a global command or application command. A speech recognizer may include a natural language software component located in a remote computing device, such as in the so-called cloud.