Adaptive Prosody in Simulated Voice via Neural Network Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voicebot systems struggle to adapt the prosody of simulated voices in real-time to match user preferences, leading to reduced user acceptance and effectiveness in natural language conversations.

Innovation Solution

A neural network-based system that receives audio inputs from users, infers audio parameters such as pitch, speed, and pauses, and generates an adapted simulated voice in real-time to maximize user acceptance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If a simulated voice is used in voicebot systems, then automated audio conversation capability is improved, but user acceptance and effectiveness are reduced due to inability to adapt prosody in real-time

Engineering Contradiction:
Improveautomated audio conversation capabilityVSAvoidprosody adaptation capability
Core Design Contradiction:
Extent of automationVSAdaptability or versatility

Solution Approach 1:

The system dynamically adjusts prosody parameters (pitch, speed, pauses) in real-time during conversations based on user audio inputs. The neural network continuously infers audio parameters from user speech patterns and modifies the simulated voice accordingly, transforming a static voice system into a dynamic adaptive one that responds to user preferences during interaction

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes physical parameters of the simulated voice such as pitch, speaking speed, and pause duration based on inferred audio parameters from user inputs. By modifying these acoustic parameters in real-time, the system adapts the prosody to match user preferences while maintaining automated conversation capability

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If real-time prosody adaptation is implemented using neural networks, then user acceptance is improved, but system complexity and computational requirements increase

Engineering Contradiction:
Improveprosody adaptation capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system replaces traditional rule-based prosody adjustment mechanisms with a neural network-based inference system. Instead of using explicit programming rules to control voice parameters, the system uses a trained neural network that automatically infers appropriate audio parameters from user audio inputs, simplifying the control logic while enabling real-time adaptation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The neural network acts as an intermediary between user audio inputs and voice generation parameters. It processes user speech patterns and translates them into appropriate prosody adjustments, serving as a computational mediator that bridges the gap between raw audio input and adapted voice output without requiring direct complex control logic

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250037702A1Adaptive prosody in simulated voice
Publication Date: 2025.01.30 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250037702A1 patent drawing
  • US20250037702A1 patent drawing
  • US20250037702A1 patent drawing

AI summary

Systems and methods for real-time generation of adapted simulated voice are described. A processor can receive an audio input from a user. The processor can run a neural network using at least one property of the audio input to infer at least one audio parameter. The processor can generate a simulated voice using the at least one audio parameter. The processor can respond to the audio input using the simulated voice.