Automated Assistant Voice Adaptation for Natural Telephone Calls
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated assistants in automated telephone calls utilize a single robotic voice throughout the call, which can be off-putting to human users, reducing the likelihood of task completion and wasting computational and network resources.
Innovation Solution
The system dynamically adapts the voice used by the automated assistant during a call based on criteria such as the entity type, location, and interaction analysis, switching to alternative voices and adjusting prosodic properties to better match the representative's accent and cadence, and adapts the rendering of unique personal identifiers based on frequency, length, and representative type.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single robotic voice is used throughout the automated telephone call, then the system complexity is reduced and computational resources are saved, but the user acceptance and task completion likelihood are reduced
Solution Approach 1:
The patent implements dynamic voice adaptation where the automated assistant switches between different voices and adjusts prosodic properties during the telephone call based on real-time analysis of the representative's accent and cadence. This dynamic adjustment allows the system to adapt to different speaking styles and accents, improving user acceptance while maintaining manageable system complexity through automated decision-making algorithms.
Solution Approach 2:
The system changes multiple parameters including voice selection, prosodic properties (intonation, tone, stress, rhythm), and rendering methods (character-by-character vs non-character-by-character) to adapt to different representatives and accents. These parameter changes enable the system to match the representative's speaking style without requiring complete system redesign, thus improving adaptability while controlling complexity.
2Reliability
If a single robotic voice is used throughout the automated telephone call, then the computational resources are saved, but the task completion likelihood is reduced
Solution Approach 1:
The system continuously monitors the telephone call and analyzes the representative's accent and cadence in real-time. Based on this feedback, the automated assistant dynamically adjusts its voice and prosodic properties to better match the representative's speaking style. This feedback mechanism improves task completion likelihood by enhancing communication effectiveness while managing computational resources through targeted processing only when adaptation is needed.
Solution Approach 2:
The system performs preliminary analysis of the representative's accent and cadence early in the call and pre-adjusts the voice parameters before critical interactions occur. This preliminary action ensures that the adapted voice is in place for important task discussions, improving task completion likelihood without requiring continuous heavy computational processing throughout the entire call duration.
3Measurement precision
If synthesized speech includes unique personal identifiers on character-by-character basis, then the speech clarity is improved, but the synthesis time and computational resources are increased
Solution Approach 1:
The system dynamically changes the rendering parameter for unique personal identifiers based on the representative's accent characteristics and the context of the conversation. When high clarity is needed, it uses character-by-character rendering; when speed is prioritized, it uses non-character-by-character rendering. This parameter adaptation balances speech clarity with synthesis time without requiring fixed computational resource allocation.
Solution Approach 2:
The patent implements dynamic adjustment of the personal identifier rendering method during the telephone call based on real-time analysis. The system can switch between character-by-character and non-character-by-character rendering approaches, and can adjust pause insertion, to optimize the balance between clarity and speed for each specific interaction context, thereby reducing unnecessary computational time while maintaining required speech quality.
4Ease of operation
If pauses are injected into synthesized speech, then the speech naturalness is improved, but the synthesis time is increased
Solution Approach 1:
The system dynamically adjusts pause insertion parameters based on the representative's cadence and speaking pace analysis. When the representative speaks slowly or with frequent pauses, the system injects similar pauses to match the natural rhythm. When the representative speaks quickly, the system reduces or eliminates pauses to maintain speed. This adaptive approach improves speech naturalness while minimizing unnecessary time addition through intelligent parameter selection.
Data Source
AI summary
Implementations are directed to dynamic adaptation of speech synthesis by an automated assistant during automated telephone call(s). In some implementations, processor(s) can select an initial voice to be utilized by the automated assistant in generating synthesized speech audio data and during an automated telephone call. However, during the automated telephone call, the processor(s) can determine to select an alternative voice to be utilized by the automated assistant in generating synthesized speech audio data and in continuing the automated telephone call. In additional or alternative implementations, and during the automated telephone call, the processor(s) can determine whether to generate any synthesized speech audio data that includes a unique personal identifier on a character-by-character basis or the unique personal identifier on a non-character-by-character basis. In additional or alternative implementations, and during the automated telephone call, the processor(s) can determine whether to inject pause(s) into any synthesized speech audio data that is generated.


