Recipient-Specific Voice Tone Generation for Telephony Calls

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Selecting the correct voice tone during voice communication over devices is difficult due to limited caller information and inconsistent training methods, leading to unmet needs for recipient-specific tone adjustment.

Innovation Solution

A system that extracts voice tone data from user samples, converts speech to text, and generates speech output using a text-to-speech model with adjusted tone, allowing users to save and apply specific tones for different recipients.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If traditional caller identification is used, then basic caller information is provided, but voice tone selection remains difficult and ineffective

Engineering Contradiction:
Improvecaller informationVSAvoidvoice tone selection
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system performs preliminary voice tone training during off-call periods, extracting voice tone characteristics from recorded samples and storing them for later use. This preliminary action prepares the tone data before actual communication needs arise, eliminating the need for difficult real-time tone selection.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically extracts voice tone characteristics from recorded voice samples without requiring manual intervention. The tone extraction and selection process is self-service, automatically matching appropriate tones to recipients based on stored training data, freeing users from manual tone selection efforts.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If voice tone training is provided, then tone selection capability is improved, but training consistency and effectiveness are insufficient

Engineering Contradiction:
Improvetone selection capabilityVSAvoidtraining consistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system replaces manual voice tone training with automated machine learning algorithms. Instead of relying on inconsistent human training methods, the system uses computational algorithms to extract voice tone characteristics from recorded samples, ensuring consistent and reliable tone generation across different recipients.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system creates copies of voice tone characteristics from recorded samples by extracting and storing tone parameters. These copied tone profiles are then applied to different recipients, ensuring consistent tone reproduction without requiring repeated manual training sessions.

Inventive Principle:
Principle #26Copying

3Ease of operation

If manual tone selection is required, then some tone control is achieved, but communication effectiveness remains unmet

Engineering Contradiction:
Improvetone controlVSAvoidcommunication effectiveness
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system incorporates feedback mechanisms where voice communication outcomes are analyzed to refine tone selection. By monitoring which tones produce better communication results, the system continuously improves its tone matching accuracy, enhancing overall communication effectiveness while maintaining ease of operation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system dynamically adjusts voice tones based on recipient-specific characteristics and communication context. Rather than using static, manually selected tones, the system adapts tones in real-time based on stored training data and communication patterns, significantly improving communication effectiveness.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250285610A1Recipient-specific voice tone adjustment in telephony
Publication Date: 2025.09.11 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250285610A1 patent drawing
  • US20250285610A1 patent drawing
  • US20250285610A1 patent drawing

AI summary

An embodiment extracts, from a plurality of voice samples, voice tone data. The embodiment converts, using a speech to text model, a speech input to corresponding text. The embodiment generates a speech output corresponding to the text, the speech output comprising audio generated from the text using a text to speech model and a voice tone generated using the voice tone data.