Automated Assistant Voice Adaptation for Natural Telephone Calls

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated assistants in automated telephone calls utilize a single robotic voice throughout the call, which can be off-putting to human users, reducing the likelihood of task completion and wasting computational and network resources.

Innovation Solution

The system dynamically adapts the voice used by the automated assistant during a call based on criteria such as the entity type, location, and interaction analysis, switching to alternative voices and adjusting prosodic properties to better match the representative's accent and cadence, and adapts the rendering of unique personal identifiers based on frequency, length, and representative type.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a single robotic voice is used throughout the automated telephone call, then the system complexity is reduced and computational resources are saved, but the user acceptance and task completion likelihood are reduced

Engineering Contradiction:
Improvevoice adaptabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic voice adaptation where the automated assistant switches between different voices and adjusts prosodic properties during the telephone call based on real-time analysis of the representative's accent and cadence. This dynamic adjustment allows the system to adapt to different speaking styles and accents, improving user acceptance while maintaining manageable system complexity through automated decision-making algorithms.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes multiple parameters including voice selection, prosodic properties (intonation, tone, stress, rhythm), and rendering methods (character-by-character vs non-character-by-character) to adapt to different representatives and accents. These parameter changes enable the system to match the representative's speaking style without requiring complete system redesign, thus improving adaptability while controlling complexity.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If a single robotic voice is used throughout the automated telephone call, then the computational resources are saved, but the task completion likelihood is reduced

Engineering Contradiction:
Improvetask completion likelihoodVSAvoidcomputational resources
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system continuously monitors the telephone call and analyzes the representative's accent and cadence in real-time. Based on this feedback, the automated assistant dynamically adjusts its voice and prosodic properties to better match the representative's speaking style. This feedback mechanism improves task completion likelihood by enhancing communication effectiveness while managing computational resources through targeted processing only when adaptation is needed.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary analysis of the representative's accent and cadence early in the call and pre-adjusts the voice parameters before critical interactions occur. This preliminary action ensures that the adapted voice is in place for important task discussions, improving task completion likelihood without requiring continuous heavy computational processing throughout the entire call duration.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If synthesized speech includes unique personal identifiers on character-by-character basis, then the speech clarity is improved, but the synthesis time and computational resources are increased

Engineering Contradiction:
Improvespeech clarityVSAvoidsynthesis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system dynamically changes the rendering parameter for unique personal identifiers based on the representative's accent characteristics and the context of the conversation. When high clarity is needed, it uses character-by-character rendering; when speed is prioritized, it uses non-character-by-character rendering. This parameter adaptation balances speech clarity with synthesis time without requiring fixed computational resource allocation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements dynamic adjustment of the personal identifier rendering method during the telephone call based on real-time analysis. The system can switch between character-by-character and non-character-by-character rendering approaches, and can adjust pause insertion, to optimize the balance between clarity and speed for each specific interaction context, thereby reducing unnecessary computational time while maintaining required speech quality.

Inventive Principle:
Principle #15Dynamics

4Ease of operation

If pauses are injected into synthesized speech, then the speech naturalness is improved, but the synthesis time is increased

Engineering Contradiction:
Improvespeech naturalnessVSAvoidsynthesis time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system dynamically adjusts pause insertion parameters based on the representative's cadence and speaking pace analysis. When the representative speaks slowly or with frequent pauses, the system injects similar pauses to match the natural rhythm. When the representative speaks quickly, the system reduces or eliminates pauses to maintain speed. This adaptive approach improves speech naturalness while minimizing unnecessary time addition through intelligent parameter selection.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250218423A1Dynamic adaptation of speech synthesis by an automated assistant during automated telephone call(s)
Publication Date: 2025.07.03 GOOGLE LLC
  • US20250218423A1 patent drawing
  • US20250218423A1 patent drawing
  • US20250218423A1 patent drawing

AI summary

Implementations are directed to dynamic adaptation of speech synthesis by an automated assistant during automated telephone call(s). In some implementations, processor(s) can select an initial voice to be utilized by the automated assistant in generating synthesized speech audio data and during an automated telephone call. However, during the automated telephone call, the processor(s) can determine to select an alternative voice to be utilized by the automated assistant in generating synthesized speech audio data and in continuing the automated telephone call. In additional or alternative implementations, and during the automated telephone call, the processor(s) can determine whether to generate any synthesized speech audio data that includes a unique personal identifier on a character-by-character basis or the unique personal identifier on a non-character-by-character basis. In additional or alternative implementations, and during the automated telephone call, the processor(s) can determine whether to inject pause(s) into any synthesized speech audio data that is generated.