Neural Network Speech Synthesis for Personalized Voice Delivery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for delivering voiced content lack the ability to dynamically alter speech patterns based on user behavior, experiences, and emotions, failing to provide personalized and contextually relevant voices.

Innovation Solution

A computer-implemented method that analyzes user profile data to recommend contextually applicable voices, transcribes voiced content into text, conditions a neural network with a voice sample to synthesize a waveform for artificial speech, and delivers the modified content, using advanced speech synthesis techniques like WaveNet or SampleRNN to mimic human-like voices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If pre-recorded voice settings are used, then voice delivery is consistent and reliable, but adaptability to user preferences and behaviors is limited

Engineering Contradiction:
Improveadaptability to user preferencesVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system dynamically adjusts voice characteristics by analyzing user profile data and behavior patterns in real-time, transitioning from static pre-recorded voices to adaptive synthesized voices that evolve with user preferences

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The neural network modifies voice parameters such as pitch, tone, and speech patterns based on analyzed user data, enabling continuous adaptation without requiring re-recording of voice settings

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If neural network speech synthesis is used, then personalized voice delivery is achieved, but processing time and computational resources increase

Engineering Contradiction:
Improvepersonalization capabilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of user profile data and pre-conditions the neural network with voice samples before actual speech generation is needed, reducing real-time processing requirements

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The neural network creates synthesized copies of voice waveforms based on trained patterns, allowing rapid generation of personalized voices without requiring actual human recording for each instance

Inventive Principle:
Principle #26Copying

3Ease of operation

If contextually relevant voice recommendations are implemented, then user experience is enhanced, but data analysis requirements and system complexity increase

Engineering Contradiction:
Improveuser experienceVSAvoiddata analysis complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system automatically analyzes user profile data and generates voice recommendations without requiring manual user input or configuration, enabling self-service personalization that reduces operational complexity

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10902841B2Personalized custom synthetic speech
Publication Date: 2021.01.26 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10902841B2 patent drawing
  • US10902841B2 patent drawing
  • US10902841B2 patent drawing

AI summary

Systems, methods, and computer program products customizing and delivering contextually relevant, artificially synthesized, voiced content that is targeted toward the individual user behaviors, viewing habits, experiences and preferences of each individual user accessing the content of a content provider. A network accessible profile service collects and analyzes collected user profile data and recommends contextually applicable voices based on the user's profile data. As user input to access voiced content or triggers voiced content maintained by a content provider, the voiced content being delivered to the user is a modified version comprising artificially synthesized human speech mimicking the recommended voice and delivering the dialogue of the voiced content, in a manner that imitates the sounds and speech patterns of the recommended voice.