Machine-Learning Accent Conversion for Clearer Virtual Conferences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Participants in virtual conferences with different native languages and accents face difficulty understanding each other due to varying speech patterns, leading to communication barriers.
Innovation Solution
A virtual conference provider employs trained machine-learning models to convert speech with one accent to another while retaining the original speaker's voice identity, using recorded speech samples from multiple speakers to generate training data for accent conversion models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech is transmitted with original accents from speakers with different native languages, then speaker identity characteristics are preserved, but communication understanding deteriorates
Solution Approach 1:
The patent introduces an accent conversion model as an intermediary component between the audio stream input and output. This model transforms the accent characteristics of speech while preserving the underlying voice identity, acting as a mediator that reconciles the conflict between maintaining speaker recognition and improving comprehensibility for participants with different linguistic backgrounds
Solution Approach 2:
The system changes the accent parameter of speech signals while maintaining other acoustic parameters that define speaker identity. By selectively modifying only the accent-related spectral and temporal parameters through machine learning models, the system achieves improved understanding without sacrificing speaker recognition
2Ease of operation
If accent conversion is implemented using machine-learning models, then communication clarity improves, but system complexity increases
Solution Approach 1:
The patent creates synthetic training data by copying and transforming existing speech samples. The system generates artificial speech datasets with various accent combinations through voice conversion techniques, eliminating the need for extensive manual recording sessions and reducing the complexity of data collection infrastructure
Solution Approach 2:
The accent conversion models are trained in advance using pre-generated synthetic training data before deployment. This preliminary training phase allows the system to learn accent transformation patterns offline, so that during actual conference operations, the conversion process is computationally efficient and does not add significant real-time complexity
3Manufacturing precision
If speech samples from multiple speakers with different accents are recorded for training, then model training data quality improves, but data collection time increases
Solution Approach 1:
The system generates synthetic training speech samples by copying existing recordings and applying voice conversion algorithms to create artificial accents. This copying approach allows the generation of diverse multi-accent training datasets from a limited set of original recordings, dramatically reducing the time required to collect training data from multiple speakers
Solution Approach 2:
The patent replaces the mechanical process of manually recording speech from multiple native speakers with different accents with an automated computational system. The voice conversion algorithms automatically transform recordings into various accent variants, substituting time-consuming human recording sessions with rapid computational synthesis
Data Source
AI summary
One example method includes receiving, during a virtual conference hosted by a virtual conference provider, a first audio stream comprising speech having first speech patterns according to a first accent, the first audio stream received from a first client device associated with a first participant in the virtual conference; generating, by a first trained machine learning (“ML”) model, a second audio stream comprising the speech having second speech patterns according to a second accent; and outputting the second audio stream.


