Handheld Speech Recognition via Remote Intermediary

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Handheld devices lack the processing power to achieve the same speed and accuracy in speech recognition as desktop or server computers, limiting their speech recognition capabilities.

Innovation Solution

A method where a handheld device receives an audio signal, converts it into digital data, and transmits it to a more powerful computing device for processing, allowing real-time execution of instructions for tasks such as text conversion and command implementation, leveraging wireless communication to utilize the processing power of desktop or server computers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition processing is performed locally on the handheld device, then processing speed and accuracy improve, but device complexity and power consumption increase beyond acceptable limits for handheld form factors

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing power requirement
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a remote computing device as an intermediary to perform the computationally intensive speech recognition processing. The handheld device acts as a client that captures audio input and transmits it to the remote server, which then processes the speech and returns results. This mediator approach allows the handheld device to achieve high speech recognition accuracy without requiring complex local processing hardware.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If more processing resources are allocated to speech recognition on the handheld device, then speech recognition speed and accuracy improve, but the device becomes less portable and more expensive

Engineering Contradiction:
Improvespeech recognition speedVSAvoiddevice portability
Core Design Contradiction:
ProductivityVSWeight of moving object

Solution Approach 1:

The patent uses a remote computing server as a mediator to perform speech recognition processing externally. The handheld device maintains its portability by only performing lightweight functions such as audio capture and result display, while the computationally intensive speech recognition tasks are offloaded to the remote server, achieving high speech recognition speed without increasing device weight.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If complex speech recognition algorithms are implemented on the handheld device, then recognition accuracy improves, but processing power requirements exceed what is feasible for handheld devices

Engineering Contradiction:
Improvevoice interpretation accuracyVSAvoidprocessing power
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent introduces a remote computing device as an intermediary that hosts complex speech recognition algorithms. The handheld device communicates with this remote server, transmitting audio data and receiving interpreted results. This allows the system to achieve high voice interpretation accuracy through sophisticated algorithms while the handheld device itself requires minimal processing power, as the computational burden is handled by the remote intermediary.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9865263B2Real-time voice recognition on a handheld device
Publication Date: 2018.01.09 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9865263B2 patent drawing
  • US9865263B2 patent drawing
  • US9865263B2 patent drawing

AI summary

A method and apparatus for implementation of real-time speech recognition using a handheld computing apparatus are provided. The handheld computing apparatus receives an audio signal, such as a user's voice. The handheld computing apparatus ultimately transmits the voice data to a remote or distal computing device with greater processing power and operating a speech recognition software application, the speech recognition software application processes the signal and outputs a set of instructions for implementation either by the computing device or the handheld apparatus. The instructions can include a variety of items including instructing the presentation of a textual representation of dictation, or a function or command to be executed by the handheld device (such as linking to a website, opening a file, cutting, pasting, saving, or other file menu type functionalities), or by the computing device itself.