Server-Based Voice Parameter Extraction for Instant Messaging TTS

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current instant messaging technologies lack effective solutions for visually impaired users, as they often require large storage resources on client devices and fail to produce speech that accurately represents the voice characteristics of message authors, leading to unintelligible communication.

Innovation Solution

The method involves server-side storage and analysis of a user's voice data, allowing the client device to generate speech that closely resembles the author's voice by transmitting synthesis parameters or phoneme samples, minimizing resource consumption and enabling distinctive voice recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If concatenative synthesis uses pre-recorded human voice samples to achieve natural-sounding speech, then speech naturalness is improved, but storage space requirements increase significantly

Engineering Contradiction:
Improvespeech naturalnessVSAvoidstorage space
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential voice characteristics (formant parameters, pitch contours, timing information) from complete voice samples, storing these compressed parameters on the server instead of full audio recordings. This extraction approach maintains speech naturalness while dramatically reducing storage requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system creates parameter-based copies of voice characteristics rather than copying actual audio samples. These parameter sets can be transmitted and replayed to synthesize speech that mimics the original speaker's voice without requiring storage of the original audio files.

Inventive Principle:
Principle #26Copying

2Quantity of substance

If formant synthesis uses acoustic models to generate speech, then storage space is reduced, but speech naturalness deteriorates due to robotic-sounding output

Engineering Contradiction:
Improvestorage spaceVSAvoidspeech naturalness
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent dynamically adjusts formant parameters, pitch contours, and timing parameters based on the stored voice characteristic data. By changing these parameters to match the speaker's natural speech patterns, the system achieves natural-sounding synthesis with minimal storage requirements.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system performs preliminary analysis of the speaker's voice to extract and store characteristic parameters before speech synthesis is needed. This advance preparation allows the TTS engine to generate natural-sounding speech on-demand without requiring large pre-recorded databases.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If client devices store complete voice sample databases for TTS, then speech quality is improved, but device resource consumption increases

Engineering Contradiction:
Improvespeech qualityVSAvoiddevice resource consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent introduces a server as an intermediary that stores and manages voice characteristic parameters. The server acts as a mediator between the speaker's voice and the client device's TTS engine, providing only the essential parameter data needed for high-quality synthesis without requiring the client to store complete voice databases.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If TTS systems use large databases of speech samples to cover extensive vocabulary, then language coverage is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvelanguage coverageVSAvoidprocessing speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent segments the speech synthesis process into parameter extraction, parameter storage, and parameter-based synthesis stages. By separating these functions and storing only essential parameters, the system achieves extensive language coverage without the computational burden of processing large audio sample databases.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8224647B2Text-to-speech user's voice cooperative server for instant messaging clients
Publication Date: 2012.07.17 CERENCE OPERATING CO
  • US8224647B2 patent drawing
  • US8224647B2 patent drawing
  • US8224647B2 patent drawing

AI summary

A system and method to allow an author of an instant message to enable and control the production of audible speech to the recipient of the message. The voice of the author of the message is characterized into parameters compatible with a formative or articulative text-to-speech engine such that upon receipt, the receiving client device can generate audible speech signals from the message text according to the characterization of the author's voice. Alternatively, the author can store samples of his or her actual voice in a server so that, upon transmission of a message by the author to a recipient, the server extracts the samples needed only to synthesize the words in the text message, and delivers those to the receiving client device so that they are used by a client-side concatenative text-to-speech engine to generate audible speech signals having a close likeness to the actual voice of the author.