Dynamic Voice Assistant Speech Modulation for Accent Clarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice assistants often provide responses to user commands that are not understood due to pronunciation and accent differences, leading to user frustration.

Innovation Solution

A dynamic speech modulation system that learns a user's pronunciation and accent models to modify responses accordingly, using machine learning and natural language processing to enhance understanding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the voice assistant uses a fixed pronunciation model for responses, then the system complexity is reduced, but the user understanding of responses deteriorates due to pronunciation and accent differences

Engineering Contradiction:
Improveuser understanding of responsesVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic speech modulation by switching between different pronunciation models (e.g., neutral accent model and user accent model) based on real-time detection of user understanding. The system transitions from a static, fixed pronunciation model to a dynamic one that adapts to user needs, thereby improving response understandability while managing complexity through conditional application rather than continuous operation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the pronunciation parameter of the voice assistant's responses based on detected user comprehension levels. When understanding is poor, the system switches to a different accent model or modifies phonetic parameters to enhance clarity. This parameter adjustment allows the system to optimize response delivery without requiring complete system redesign, thus improving reliability while controlling complexity.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If the voice assistant provides multiple responses to ensure understanding, then the user satisfaction improves, but the time to provide responses increases

Engineering Contradiction:
Improveuser satisfactionVSAvoidresponse time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system employs feedback mechanisms where the voice assistant detects user understanding in real-time and adjusts responses accordingly. This feedback loop allows the system to provide additional responses only when necessary (when understanding is detected as insufficient), rather than consistently providing multiple responses. Thus, user satisfaction is improved through adaptive response strategies while minimizing time loss by avoiding unnecessary repeated responses.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The response strategy dynamically adjusts based on detected user comprehension. The system transitions from a static response protocol to a dynamic one that provides multiple responses only when needed, thereby improving satisfaction through adaptability while reducing time consumption by avoiding unnecessary repeated responses in cases where the first response is already understood.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If the voice assistant learns user pronunciation models dynamically, then the adaptability to user accents improves, but the processing time increases

Engineering Contradiction:
Improveadaptability to user accentsVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary learning of user pronunciation models during initial interactions or idle periods, storing the accent characteristics in advance. This preliminary action allows the system to quickly apply the learned models during actual response generation without performing time-consuming real-time analysis, thus improving adaptability to user accents while minimizing processing time during critical response delivery moments.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically balances between using pre-learned pronunciation models and performing real-time adaptation based on user feedback. When a user accent model is available, the system quickly switches to using it rather than performing lengthy real-time learning processes. This dynamic strategy improves adaptability by leveraging previously gathered data while reducing processing time by avoiding redundant real-time analysis.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12444414B2Dynamic virtual assistant speech modulation
Publication Date: 2025.10.14 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12444414B2 patent drawing
  • US12444414B2 patent drawing
  • US12444414B2 patent drawing

AI summary

A method, computer system, and a computer program product for dynamic speech modulation is provided. The present invention may include transmitting a first response to a received command. The present invention may include determining the first response is not understood by a user. The present invention may include transmitting a second response to the received command.