Mobile Voice Platform Speech Recognition Interface
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech-based human-machine interfaces in vehicles and cellular phones are limited in their ability to provide hands-free access to a wide range of services, relying heavily on screen interaction and not fully utilizing the capabilities of mobile devices with speech recognition and processing.
Innovation Solution
A method that establishes a short-range wireless connection between a mobile device and audio devices, using automated speech recognition to process user input and determine service requests, enabling hands-free access to various Internet-based and computer-based services through a mobile voice platform that includes a speech platform kernel and application interface suite, allowing interaction with both installed apps and cloud services.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If screen interaction is used for service selection, then user interface simplicity is improved, but driver distraction and hands-free capability deteriorate
Solution Approach 1:
The patent replaces the mechanical/screen-based interaction system with an acoustic field-based speech recognition system. Users communicate service selection requests through speech commands captured by a microphone, which are then processed by speech recognition software to identify and execute the desired service, eliminating the need for physical screen interaction
Solution Approach 2:
The patent introduces speech recognition technology as an intermediary between the user's spoken commands and the service execution system. This intermediary layer converts acoustic signals into actionable service requests, enabling hands-free operation while maintaining intuitive interaction
2Device complexity
If limited command set is used for speech interface, then system complexity is reduced, but service versatility deteriorates
Solution Approach 1:
The patent implements a universal speech recognition framework that can identify and execute multiple different services through a single unified interface. The speech recognition system processes diverse service requests (music playback, navigation, messaging, etc.) using the same acoustic input mechanism, allowing one system to perform many functions without requiring separate command sets for each service
Solution Approach 2:
The patent employs dynamic service identification where the system adapts its response based on the spoken command. The speech recognition module dynamically parses user requests and routes them to appropriate services, allowing the interface to flexibly handle various service types without requiring static, service-specific command structures
3Ease of operation
If basic mobile device functions are enabled, then ease of operation is improved, but functionality and integration deteriorate
Solution Approach 1:
The patent creates a universal speech-based control layer that provides easy operation for basic functions while simultaneously enabling access to advanced and integrated services. The same speech interface handles both simple tasks like calling and complex integrated services like navigation or messaging, maintaining ease of use across all functionality levels
Data Source
AI summary
A method of providing hands-free services using a mobile device having wireless access to computer-based services includes establishing a short range wireless connection between a mobile device and one or more audio devices that include at least a microphone and speaker; receiving at the mobile device speech inputted via the microphone from a user and sent via the short range wireless connection; wirelessly transmitting the speech input from the mobile device to a speech recognition server that provides automated speech recognition (ASR); receiving at the mobile device a speech recognition result representing the content of the speech input; determining a desired service by processing the speech recognition result using a first, service-identifying grammar; determining a user service request by processing at least some of the speech recognition result using a second, service-specific grammar associated with the desired service; initiating the user service request and receiving a service response; generating an audio message from the service response; and presenting the audio message to the user via the speaker.


