Synthetic Voice Customization via Vocal Sample Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-speech applications lack customization to mimic the voice of the sender, resulting in speech output that does not accurately represent the human voice of the message sender.

Innovation Solution

A system and method that measure vocal characteristics such as frequency, timbre, intensity, and rhythm from a recorded vocal sample to create a synthetic voice that approximates the sender's voice, allowing for personalized speech output in text messages, using formant or concatenative synthesis techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If standardized synthetic voices are used in text-to-speech applications, then the system is simple and easy to implement, but the speech output cannot be customized to match the sender's voice

Engineering Contradiction:
Improvevoice customizationVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary voice analysis by recording and analyzing the sender's voice sample in advance, extracting vocal characteristics such as pitch, tone, and rhythm. This preliminary action enables the system to store these characteristics for later use, allowing the text-to-speech application to generate customized speech output without requiring complex real-time processing during message transmission.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a synthetic voice model that copies the essential characteristics of the sender's voice. By analyzing the voice sample and generating a simplified digital representation of the sender's vocal patterns, the system can reproduce the sender's voice in text-to-speech output without needing to replicate the entire complexity of the human vocal apparatus.

Inventive Principle:
Principle #26Copying

2Measurement precision

If voice analysis and synthesis processing is performed, then the speech output accurately represents the sender's voice, but the processing time and computational resources increase

Engineering Contradiction:
Improvevoice characteristic accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs voice analysis and extracts vocal characteristics in advance, before the text-to-speech conversion is needed. By completing the complex analysis work beforehand and storing the results, the system minimizes processing time during actual message transmission while maintaining high accuracy in voice representation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system focuses on analyzing and reproducing only the most critical voice characteristics (such as pitch, tone, and rhythm) rather than attempting to capture every nuance of the voice. This partial action approach achieves sufficient accuracy for recognizable voice representation while significantly reducing computational complexity and processing time.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10614792B2Method and system for using a vocal sample to customize text to speech applications
Publication Date: 2020.04.07 MASON PAUL WENDELL
  • US10614792B2 patent drawing

AI summary

Apparatus and methods consistent with the present invention measure one or more of the characteristics of a voice recording and use such measurements to create a synthetic voice that approximates the recorded voice and uses such created synthetic voice to verbalize the content of an electronically conveyed written message such as an SMS text message. The vocal characteristics measured may include frequency, timbre, intensity, rhythm, and rate of speech as well as others.