Server-Based Voice Parameter Extraction for Instant Messaging TTS
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current instant messaging technologies lack effective solutions for visually impaired users, as they often require large storage resources on client devices and fail to produce speech that accurately represents the voice characteristics of message authors, leading to unintelligible communication.
Innovation Solution
The method involves server-side storage and analysis of a user's voice data, allowing the client device to generate speech that closely resembles the author's voice by transmitting synthesis parameters or phoneme samples, minimizing resource consumption and enabling distinctive voice recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If concatenative synthesis uses pre-recorded human voice samples to achieve natural-sounding speech, then speech naturalness is improved, but storage space requirements increase significantly
Solution Approach 1:
The patent extracts only the essential voice characteristics (formant parameters, pitch contours, timing information) from complete voice samples, storing these compressed parameters on the server instead of full audio recordings. This extraction approach maintains speech naturalness while dramatically reducing storage requirements.
Solution Approach 2:
The system creates parameter-based copies of voice characteristics rather than copying actual audio samples. These parameter sets can be transmitted and replayed to synthesize speech that mimics the original speaker's voice without requiring storage of the original audio files.
2Quantity of substance
If formant synthesis uses acoustic models to generate speech, then storage space is reduced, but speech naturalness deteriorates due to robotic-sounding output
Solution Approach 1:
The patent dynamically adjusts formant parameters, pitch contours, and timing parameters based on the stored voice characteristic data. By changing these parameters to match the speaker's natural speech patterns, the system achieves natural-sounding synthesis with minimal storage requirements.
Solution Approach 2:
The system performs preliminary analysis of the speaker's voice to extract and store characteristic parameters before speech synthesis is needed. This advance preparation allows the TTS engine to generate natural-sounding speech on-demand without requiring large pre-recorded databases.
3Manufacturing precision
If client devices store complete voice sample databases for TTS, then speech quality is improved, but device resource consumption increases
Solution Approach 1:
The patent introduces a server as an intermediary that stores and manages voice characteristic parameters. The server acts as a mediator between the speaker's voice and the client device's TTS engine, providing only the essential parameter data needed for high-quality synthesis without requiring the client to store complete voice databases.
4Adaptability or versatility
If TTS systems use large databases of speech samples to cover extensive vocabulary, then language coverage is improved, but processing time and computational resources increase
Solution Approach 1:
The patent segments the speech synthesis process into parameter extraction, parameter storage, and parameter-based synthesis stages. By separating these functions and storing only essential parameters, the system achieves extensive language coverage without the computational burden of processing large audio sample databases.
Data Source
AI summary
A system and method to allow an author of an instant message to enable and control the production of audible speech to the recipient of the message. The voice of the author of the message is characterized into parameters compatible with a formative or articulative text-to-speech engine such that upon receipt, the receiving client device can generate audible speech signals from the message text according to the characterization of the author's voice. Alternatively, the author can store samples of his or her actual voice in a server so that, upon transmission of a message by the author to a recipient, the server extracts the samples needed only to synthesize the words in the text message, and delivers those to the receiving client device so that they are used by a client-side concatenative text-to-speech engine to generate audible speech signals having a close likeness to the actual voice of the author.


