Mobile Speech Recognition via Phonetic Symbol Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition methods on mobile devices face challenges due to inadequate processing power and memory, requiring server-based solutions that incur costs and introduce latency, especially when dealing with large datasets or sensitive information, and result in data transmission distortions.

Innovation Solution

Converting speech input to a compact sequence of phonetic symbols on the mobile device for transmission over the network, allowing for efficient and cost-effective matching against large datasets on a server, with optional local refinement for improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speech waveform is transmitted over network to server for recognition, then recognition capability is improved, but transmission cost and latency increase

Engineering Contradiction:
Improverecognition capabilityVSAvoidlatency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts only the essential phonetic information from the speech waveform by converting it to phonetic symbols locally on the mobile device, rather than transmitting the complete waveform. This extraction reduces the data volume to be transmitted while preserving the core recognition information, thereby reducing latency and transmission costs.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The speech recognition process is segmented into two parts: local processing on the mobile device (speech to phonetic symbol conversion) and remote processing on the server (phonetic symbol matching). This segmentation allows the computationally intensive local conversion to be done efficiently on the device while only minimal data is transmitted to the server, resolving the latency-cost contradiction.

Inventive Principle:
Principle #1Segmentation

2Reliability

If speech waveform is transmitted over network to server for recognition, then recognition capability is improved, but transmission cost increases

Engineering Contradiction:
Improverecognition capabilityVSAvoidtransmission cost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent extracts only the essential phonetic information from the speech waveform by converting it to phonetic symbols locally on the mobile device, rather than transmitting the complete waveform. This extraction reduces the data volume to be transmitted while preserving the core recognition information, thereby reducing latency and transmission costs.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The phonetic symbol conversion is performed as a preliminary action locally on the mobile device before transmission to the server. This preliminary processing reduces the amount of data that needs to be transmitted over the network, significantly lowering transmission costs and energy consumption while maintaining recognition capability.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If large database is stored on mobile device for local recognition, then recognition speed is improved, but device complexity and memory requirements increase

Engineering Contradiction:
Improverecognition speedVSAvoidmemory requirements
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The speech recognition process is segmented into two parts: local processing on the mobile device (speech to phonetic symbol conversion) and remote processing on the server (phonetic symbol matching). This segmentation allows the computationally intensive local conversion to be done efficiently on the device while only minimal data is transmitted to the server, resolving the latency-cost contradiction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces phonetic symbols as an intermediary representation between the speech waveform and the large database of phonetic reference forms stored on the server. This intermediary format enables efficient matching without requiring the mobile device to store the entire large database, thus maintaining recognition speed while reducing local memory requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If speech waveform is transmitted over network, then recognition capability is improved, but data transmission distortions occur

Engineering Contradiction:
Improverecognition capabilityVSAvoidtransmission distortion
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent extracts only the essential phonetic information from the speech waveform by converting it to phonetic symbols locally on the mobile device, rather than transmitting the complete waveform. This extraction reduces the data volume to be transmitted while preserving the core recognition information, thereby reducing latency and transmission costs.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS9959870B2Speech recognition involving a mobile device
Publication Date: 2018.05.01 APPLE INC
  • US9959870B2 patent drawing
  • US9959870B2 patent drawing
  • US9959870B2 patent drawing

AI summary

A system and method of speech recognition involving a mobile device. Speech input is received (202) on a mobile device (102) and converted (204) to a set of phonetic symbols. Data relating to the phonetic symbols is transferred (206) from the mobile device over a communications network (104) to a remote processing device (106) where it is used (208) to identify at least one matching data item from a set of data items (114). Data relating to the at least one matching data item is transferred (210) from the remote processing device to the mobile device and presented (214) thereon.