Machine-Learning Accent Conversion for Clearer Virtual Conferences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Participants in virtual conferences with different native languages and accents face difficulty understanding each other due to varying speech patterns, leading to communication barriers.

Innovation Solution

A virtual conference provider employs trained machine-learning models to convert speech with one accent to another while retaining the original speaker's voice identity, using recorded speech samples from multiple speakers to generate training data for accent conversion models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speech is transmitted with original accents from speakers with different native languages, then speaker identity characteristics are preserved, but communication understanding deteriorates

Engineering Contradiction:
Improvespeaker identity preservationVSAvoidcommunication understanding
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent introduces an accent conversion model as an intermediary component between the audio stream input and output. This model transforms the accent characteristics of speech while preserving the underlying voice identity, acting as a mediator that reconciles the conflict between maintaining speaker recognition and improving comprehensibility for participants with different linguistic backgrounds

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the accent parameter of speech signals while maintaining other acoustic parameters that define speaker identity. By selectively modifying only the accent-related spectral and temporal parameters through machine learning models, the system achieves improved understanding without sacrificing speaker recognition

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If accent conversion is implemented using machine-learning models, then communication clarity improves, but system complexity increases

Engineering Contradiction:
Improvecommunication clarityVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent creates synthetic training data by copying and transforming existing speech samples. The system generates artificial speech datasets with various accent combinations through voice conversion techniques, eliminating the need for extensive manual recording sessions and reducing the complexity of data collection infrastructure

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The accent conversion models are trained in advance using pre-generated synthetic training data before deployment. This preliminary training phase allows the system to learn accent transformation patterns offline, so that during actual conference operations, the conversion process is computationally efficient and does not add significant real-time complexity

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If speech samples from multiple speakers with different accents are recorded for training, then model training data quality improves, but data collection time increases

Engineering Contradiction:
Improvetraining data qualityVSAvoiddata collection time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system generates synthetic training speech samples by copying existing recordings and applying voice conversion algorithms to create artificial accents. This copying approach allows the generation of diverse multi-accent training datasets from a limited set of original recordings, dramatically reducing the time required to collect training data from multiple speakers

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of manually recording speech from multiple native speakers with different accents with an automated computational system. The voice conversion algorithms automatically transform recordings into various accent variants, substituting time-consuming human recording sessions with rapid computational synthesis

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12375624B2Accent conversion for virtual conferences
Publication Date: 2025.07.29 ZOOM COMMUNICATIONS INC
  • US12375624B2 patent drawing
  • US12375624B2 patent drawing
  • US12375624B2 patent drawing

AI summary

One example method includes receiving, during a virtual conference hosted by a virtual conference provider, a first audio stream comprising speech having first speech patterns according to a first accent, the first audio stream received from a first client device associated with a first participant in the virtual conference; generating, by a first trained machine learning (“ML”) model, a second audio stream comprising the speech having second speech patterns according to a second accent; and outputting the second audio stream.