Mobile Speech Recognition via Phonetic Symbol Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition methods on mobile devices face challenges due to inadequate processing power and memory, requiring server-based solutions that incur costs and introduce latency, especially when dealing with large datasets or sensitive information, and result in data transmission distortions.
Innovation Solution
Converting speech input to a compact sequence of phonetic symbols on the mobile device for transmission over the network, allowing for efficient and cost-effective matching against large datasets on a server, with optional local refinement for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech waveform is transmitted over network to server for recognition, then recognition capability is improved, but transmission cost and latency increase
Solution Approach 1:
The patent extracts only the essential phonetic information from the speech waveform by converting it to phonetic symbols locally on the mobile device, rather than transmitting the complete waveform. This extraction reduces the data volume to be transmitted while preserving the core recognition information, thereby reducing latency and transmission costs.
Solution Approach 2:
The speech recognition process is segmented into two parts: local processing on the mobile device (speech to phonetic symbol conversion) and remote processing on the server (phonetic symbol matching). This segmentation allows the computationally intensive local conversion to be done efficiently on the device while only minimal data is transmitted to the server, resolving the latency-cost contradiction.
2Reliability
If speech waveform is transmitted over network to server for recognition, then recognition capability is improved, but transmission cost increases
Solution Approach 1:
The patent extracts only the essential phonetic information from the speech waveform by converting it to phonetic symbols locally on the mobile device, rather than transmitting the complete waveform. This extraction reduces the data volume to be transmitted while preserving the core recognition information, thereby reducing latency and transmission costs.
Solution Approach 2:
The phonetic symbol conversion is performed as a preliminary action locally on the mobile device before transmission to the server. This preliminary processing reduces the amount of data that needs to be transmitted over the network, significantly lowering transmission costs and energy consumption while maintaining recognition capability.
3Productivity
If large database is stored on mobile device for local recognition, then recognition speed is improved, but device complexity and memory requirements increase
Solution Approach 1:
The speech recognition process is segmented into two parts: local processing on the mobile device (speech to phonetic symbol conversion) and remote processing on the server (phonetic symbol matching). This segmentation allows the computationally intensive local conversion to be done efficiently on the device while only minimal data is transmitted to the server, resolving the latency-cost contradiction.
Solution Approach 2:
The patent introduces phonetic symbols as an intermediary representation between the speech waveform and the large database of phonetic reference forms stored on the server. This intermediary format enables efficient matching without requiring the mobile device to store the entire large database, thus maintaining recognition speed while reducing local memory requirements.
4Reliability
If speech waveform is transmitted over network, then recognition capability is improved, but data transmission distortions occur
Solution Approach 1:
The patent extracts only the essential phonetic information from the speech waveform by converting it to phonetic symbols locally on the mobile device, rather than transmitting the complete waveform. This extraction reduces the data volume to be transmitted while preserving the core recognition information, thereby reducing latency and transmission costs.
Data Source
AI summary
A system and method of speech recognition involving a mobile device. Speech input is received (202) on a mobile device (102) and converted (204) to a set of phonetic symbols. Data relating to the phonetic symbols is transferred (206) from the mobile device over a communications network (104) to a remote processing device (106) where it is used (208) to identify at least one matching data item from a set of data items (114). Data relating to the at least one matching data item is transferred (210) from the remote processing device to the mobile device and presented (214) thereon.


