Dual Speech Recognizers for Simultaneous OS and App Command Interpretation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems struggle to interpret global commands and application commands simultaneously, often leading to conflicts due to overlapping commands, phonetic similarities, and incompatibility of speech technologies, which results in ambiguity and a cumbersome user experience.
Innovation Solution
A speech-enabled system that employs two speech recognizers operating simultaneously to interpret operating system commands and application commands, using unrestricted speech grammars and natural language processing, with reserved words or cadence to differentiate between system and application commands, allowing users to speak freely without the need for a push-to-talk button.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single speech recognizer is used, then the system is simple, but it cannot interpret both operating system commands and application commands simultaneously
Solution Approach 1:
The patent divides the speech recognition system into separate components: a first speech recognizer for operating system commands and a second speech recognizer for application commands. This segmentation allows each recognizer to specialize in specific command types, enabling simultaneous interpretation of both system and application commands without interference.
Solution Approach 2:
The patent implements a universal command routing mechanism that directs speech input to the appropriate recognizer based on context. The system can handle multiple command types through a single integrated architecture, where the routing logic determines whether to send commands to the OS recognizer, application recognizer, or both simultaneously.
2Measurement precision
If speech recognizers operate simultaneously, then command interpretation is accurate, but ambiguity arises from overlapping commands and phonetic similarities
Solution Approach 1:
The patent introduces a command routing mechanism as an intermediary between the speech input and the recognizers. This mediator analyzes the speech input and determines the appropriate target (OS command, application command, or both), preventing direct conflicts between recognizers and enabling accurate simultaneous operation.
Solution Approach 2:
The system incorporates feedback mechanisms where the routing logic continuously monitors speech input characteristics and adjusts the distribution of commands to appropriate recognizers. This feedback loop resolves ambiguities by adapting to the context of spoken commands in real-time.
3Ease of operation
If unrestricted speech grammars are used, then natural language processing is improved, but command differentiation between system and application becomes difficult
Solution Approach 1:
The patent applies different speech grammar constraints to different recognizers based on their specific functions. The first recognizer uses grammar optimized for operating system commands while the second uses grammar for application commands. This local differentiation maintains unrestricted natural language input while enabling clear command differentiation through context-specific grammar rules.
Data Source
AI summary
A speech recognition system interprets both spoken system commands as well as application commands. Users may speak commands to an open microphone of a computing device that may be interpreted by at least two speech recognizers operating simultaneously. The first speech recognizer interprets operating system commands and the second speech recognizer interprets application commands. The system commands may include at least opening and closing an application and the application commands may include at least a game command or navigation within a menu. A reserve word may be used to identify whether the command is for the operation system or application. A user's cadence may also indicate whether the speech is a global command or application command. A speech recognizer may include a natural language software component located in a remote computing device, such as in the so-called cloud.


