Network-Based Speech Processing API for Mobile User Interfaces

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Companies face a high barrier to entry for developing voice-enabled services due to the need for expensive customized systems, complex components, and specialized expertise, making it difficult for them to integrate voice recognition and synthesis into their user interfaces without significant investment.

Innovation Solution

A network-based architecture that uses a public, common network node to process speech and return text, allowing companies to implement voice-enabled services through a simple API, reducing the need for expensive engines and servers by enabling speech processing within the network accessible via IP protocol.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If companies implement traditional voice-enabled services with customized speech processing engines and servers, then speech recognition and synthesis functionality is achieved, but the cost and complexity of the system increases significantly

Engineering Contradiction:
Improvespeech processing functionalityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the complex speech processing engine from the client device and relocates it to a remote server. The client device only needs to capture audio and send it to the server, which performs the heavy speech recognition and synthesis processing. This extraction eliminates the need for companies to deploy and maintain complex speech processing engines on their own systems, thereby reducing device complexity while preserving speech processing functionality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary web server that acts as a mediator between the client device and the speech processing engine. The server receives audio data from the client, processes it through the speech engine, and returns results. This intermediary layer simplifies the client device architecture while maintaining full speech processing capabilities through the server-based engine.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If companies develop custom voice-enabled services with specialized speech processing components, then accurate speech recognition is achieved, but the barrier to entry and initial investment increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidease of service deployment
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent creates a universal speech processing service that can be accessed by multiple different applications and devices through a common interface. The server-based architecture allows any company to utilize the same high-accuracy speech engine for different purposes (voice commands, transcription, search, etc.) without needing to develop or customize the engine themselves, thereby lowering the barrier to entry while maintaining recognition accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent enables companies to self-service speech processing needs by providing accessible APIs and web interfaces. Organizations can integrate speech recognition into their existing applications without requiring specialized expertise in speech processing, as the server handles all complex processing automatically. This self-service model reduces both the barrier to entry and initial investment requirements.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If companies deploy speech processing engines and training data locally, then customized speech recognition is achieved, but the cost of hardware and expertise increases

Engineering Contradiction:
Improvecustomized speech recognitionVSAvoidresource investment
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent extracts the resource-intensive speech processing engine and training data from local deployment and relocates them to a centralized server. This allows companies to access customized speech recognition capabilities without investing in expensive local hardware or hiring specialized teams, as all processing resources are consolidated on the server infrastructure.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent merges multiple speech processing functions, training data, and computational resources into a single centralized server system. This consolidation allows multiple companies to share the same infrastructure, reducing individual resource investment while maintaining the ability to provide customized speech recognition through the unified platform.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9530415B2System and method of providing speech processing in user interface
Publication Date: 2016.12.27 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9530415B2 patent drawing
  • US9530415B2 patent drawing
  • US9530415B2 patent drawing

AI summary

Disclosed are systems, methods and computer-readable media for enabling speech processing in a user interface of a device. The method includes receiving an indication of a field and a user interface of a device, the indication also signaling that speech will follow, receiving the speech from the user at the device, the speech being associated with the field, transmitting the speech as a request to public, common network node that receives and processes speech, processing the transmitted speech and returning text associated with the speech to the device and inserting the text into the field. Upon a second indication from the user, the system processes the text in the field as programmed by the user interface. The present disclosure provides a speech mash up application for a user interface of a mobile or desktop device that does not require expensive speech processing technologies.