Network-Based Speech Processing API for Mobile User Interfaces
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Companies face a high barrier to entry for developing voice-enabled services due to the need for expensive customized systems, complex components, and specialized expertise, making it difficult for them to integrate voice recognition and synthesis into their user interfaces without significant investment.
Innovation Solution
A network-based architecture that uses a public, common network node to process speech and return text, allowing companies to implement voice-enabled services through a simple API, reducing the need for expensive engines and servers by enabling speech processing within the network accessible via IP protocol.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If companies implement traditional voice-enabled services with customized speech processing engines and servers, then speech recognition and synthesis functionality is achieved, but the cost and complexity of the system increases significantly
Solution Approach 1:
The patent extracts the complex speech processing engine from the client device and relocates it to a remote server. The client device only needs to capture audio and send it to the server, which performs the heavy speech recognition and synthesis processing. This extraction eliminates the need for companies to deploy and maintain complex speech processing engines on their own systems, thereby reducing device complexity while preserving speech processing functionality.
Solution Approach 2:
The patent introduces an intermediary web server that acts as a mediator between the client device and the speech processing engine. The server receives audio data from the client, processes it through the speech engine, and returns results. This intermediary layer simplifies the client device architecture while maintaining full speech processing capabilities through the server-based engine.
2Measurement precision
If companies develop custom voice-enabled services with specialized speech processing components, then accurate speech recognition is achieved, but the barrier to entry and initial investment increases
Solution Approach 1:
The patent creates a universal speech processing service that can be accessed by multiple different applications and devices through a common interface. The server-based architecture allows any company to utilize the same high-accuracy speech engine for different purposes (voice commands, transcription, search, etc.) without needing to develop or customize the engine themselves, thereby lowering the barrier to entry while maintaining recognition accuracy.
Solution Approach 2:
The patent enables companies to self-service speech processing needs by providing accessible APIs and web interfaces. Organizations can integrate speech recognition into their existing applications without requiring specialized expertise in speech processing, as the server handles all complex processing automatically. This self-service model reduces both the barrier to entry and initial investment requirements.
3Adaptability or versatility
If companies deploy speech processing engines and training data locally, then customized speech recognition is achieved, but the cost of hardware and expertise increases
Solution Approach 1:
The patent extracts the resource-intensive speech processing engine and training data from local deployment and relocates them to a centralized server. This allows companies to access customized speech recognition capabilities without investing in expensive local hardware or hiring specialized teams, as all processing resources are consolidated on the server infrastructure.
Solution Approach 2:
The patent merges multiple speech processing functions, training data, and computational resources into a single centralized server system. This consolidation allows multiple companies to share the same infrastructure, reducing individual resource investment while maintaining the ability to provide customized speech recognition through the unified platform.
Data Source
AI summary
Disclosed are systems, methods and computer-readable media for enabling speech processing in a user interface of a device. The method includes receiving an indication of a field and a user interface of a device, the indication also signaling that speech will follow, receiving the speech from the user at the device, the speech being associated with the field, transmitting the speech as a request to public, common network node that receives and processes speech, processing the transmitted speech and returning text associated with the speech to the device and inserting the text into the field. Upon a second indication from the user, the system processes the text in the field as programmed by the user interface. The present disclosure provides a speech mash up application for a user interface of a mobile or desktop device that does not require expensive speech processing technologies.


