Captioned Telephone Speech Conversion for Users With Speech Disorders
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Individuals who are deaf or hard-of-hearing and have a speech disorder face challenges in verbal communication due to inaccurate pronunciation, making it difficult for peers to understand their voice during phone calls.
Innovation Solution
A captioned telephone service system that transcribes the user's voice into text and converts it into clear, articulate speech for the peer, utilizing a speech-to-text and text-to-speech handler, with a database for improving transcription accuracy over time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a DHH user with a speech disorder speaks directly to a peer during a phone call, then the user can communicate verbally, but the peer has difficulty understanding the user's voice due to inaccurate pronunciation
Solution Approach 1:
The patent introduces a CTS server as an intermediary between the user and peer. The server captures the user's speech, transcribes it to text using speech-to-text handlers, and converts it to clear synthesized speech using text-to-speech handlers. This intermediary process eliminates the pronunciation accuracy problem while maintaining verbal communication capability.
Solution Approach 2:
The patent replaces the mechanical speech production system (user's vocal apparatus) with an electronic system. The user's spoken words are converted to text signals, then to synthesized speech signals through text-to-speech conversion. This substitution transforms the physical speech mechanism into an electronic signal processing system, achieving clear articulation regardless of the user's speech disorder.
2Measurement precision
If the CTS system transcribes the user's voice into text and converts it to clear speech, then the peer can understand the user better, but the system complexity increases
Solution Approach 1:
The CTS server performs multiple functions within a single system: it transcribes peer's speech to text for the user, transcribes user's speech to text for the peer, and converts user's text to clear synthesized speech. This multi-functionality reduces the need for separate systems while achieving comprehensive communication support.
Solution Approach 2:
The CTS server acts as a central intermediary that handles all speech processing tasks. By consolidating speech-to-text and text-to-speech conversion functions in one server, the patent manages system complexity through centralized architecture rather than distributed multiple devices.
3Measurement precision
If the system uses a database to store user's spoken audios and corresponding texts for improving transcription accuracy, then the transcription quality improves over time, but the data storage and processing requirements increase
Solution Approach 1:
The system implements feedback by storing the user's spoken audio and corresponding transcribed text in a database. The speech-to-text handlers use this stored data to learn and improve transcription accuracy over time. The feedback loop allows the system to adapt to the user's specific speech patterns and pronunciation characteristics.
Solution Approach 2:
The system improves its own transcription capability by using its own data. The stored user audio and text pairs serve as training data that enables the speech-to-text handlers to automatically improve without external intervention. The system self-optimizes by leveraging its accumulated data.
Data Source
AI summary
A captioned telephone service (CTS) system provides a transcription service and a speech-to-text and text-to-speech converting service for the deaf or hard-of-hearing user with a speech disorder during a phone call between the user and the peer. The CTS system transcribes the peer's voice into text to be displayed on the user's device. The CTS system transcribes the user's voice into text based on the database storing user's spoken audios and corresponding texts, and converts the text into a clear and articulate speech to be sent to the peer's device instead of the user's voice in order to help the peer better understand what the user said. The CTS system includes a speech-to-text handler and a text-to-speech handler. The speech-to-text handler transcribes user's voice into text using the database, and the text-to-speech handler converts the text into a clear and articulate speech.


