Voice Over SMS Phoneme Synthesis for Non-Literate Users
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In emerging markets with low literacy rates, the high cost of cell phone calls and the need for literacy to use text messaging limit the effectiveness of SMS as a communication alternative, making it inaccessible to many who cannot read or write.
Innovation Solution
A method that allows voice messages to be sent and received via SMS using phonetic recognition and synthesis, where the mobile device generates a non-text representation of an utterance, such as a phoneme string, and sends it over a wireless messaging channel, enabling non-literate users to communicate verbally through SMS.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If text messaging (SMS) is used as an alternative to voice calls, then communication cost is reduced and network capacity is improved, but accessibility is worsened for non-literate users who cannot read or write
Solution Approach 1:
The patent introduces an intermediary system that converts spoken language into text messages. A speech-to-text conversion module acts as a mediator between the user's voice and the SMS system, allowing non-literate users to communicate verbally while the system automatically transcribes and sends their messages through the existing SMS infrastructure.
Solution Approach 2:
The patent replaces the mechanical typing action with an automated speech recognition system. Instead of requiring manual keyboard input, the system uses acoustic signal processing and pattern recognition to convert spoken words into text, substituting a complex mechanical interaction with an automated electronic process.
2Quantity of substance
If voice codecs compress speech signals to low bit rates, then data transmission is minimized and network capacity is maximized, but speech quality deteriorates
Solution Approach 1:
The patent extracts only the essential phonetic and linguistic features from speech signals rather than transmitting the complete audio waveform. By identifying and transmitting only the critical elements needed for accurate text conversion (such as phoneme sequences and language model probabilities), the system achieves high-fidelity speech-to-text conversion with minimal data transmission.
Solution Approach 2:
The patent transforms the speech signal from the time domain to the feature domain, changing parameters from raw audio samples to extracted phonetic features. This parameter transformation allows the system to represent complex speech information using compact feature vectors that require far fewer bits to transmit while preserving the essential information needed for accurate text generation.
Data Source
AI summary
A method of operating a mobile communication device, the method involving: over a wireless messaging channel receiving a text message that contains a non-text representation of an utterance; extracting the non-text representation from the text message; synthesizing an audio representation of the spoken utterance from the non-text representation; and playing the synthesized audio representation through an audio output device on the mobile communication device.


