Synthetic Speech Waveforms for Privacy-Safe Acoustic Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing acoustic models in speech processing rely on personal data from users, which raises privacy concerns and limits operational flexibility, especially when dealing with sensitive populations or data that needs to be aged out.

Innovation Solution

The method generates pseudo-speaker-specific text-to-speech voices that do not use personal data, using text-to-speech datasets and automatic speech recognition data to create waveforms for training acoustic models, allowing for flexible model training and data augmentation without storing private user data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If recorded speech from individual users is used to train acoustic models, then model training accuracy is improved, but user privacy and data security are compromised

Engineering Contradiction:
Improvemodel training accuracyVSAvoiduser privacy risk
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic speech data that copies the statistical properties and acoustic characteristics of real user speech without using actual recorded speech. A text-to-speech system generates synthetic waveforms that mimic the distribution and features of target speaker speech, allowing model training while eliminating direct use of personal data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces a text-to-speech synthesis system as an intermediary between real speech data and acoustic model training. This intermediary generates synthetic speech that serves as a proxy for real speech, enabling training without direct exposure to sensitive personal data while preserving the necessary statistical properties for accurate model learning.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If user speech data is stored for acoustic model training, then training data availability is improved, but operational flexibility and data management complexity increase

Engineering Contradiction:
Improvetraining data availabilityVSAvoidoperational flexibility
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent enables the system to generate its own training data on-demand through text-to-speech synthesis without relying on stored datasets. The system can self-generate unlimited synthetic speech data with various characteristics as needed, eliminating the need for data storage infrastructure and associated management complexities while maintaining data availability.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent allows dynamic control of synthetic speech generation parameters (such as speaker characteristics, audio conditions, and speech styles) to adapt to different training requirements. This enables flexible data generation for various scenarios without being constrained by pre-stored datasets, enhancing operational versatility while reducing data management burden.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11545136B2System and method using parameterized speech synthesis to train acoustic models
Publication Date: 2023.01.03 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11545136B2 patent drawing
  • US11545136B2 patent drawing
  • US11545136B2 patent drawing

AI summary

A method for removing private data from an acoustic model includes capturing speech from a large population of users, creating a text-to-speech voice from at least a portion of the large population of users, discarding speech data from a database of speech, creating text-to-speech waveforms from the text-to-speech voice and the new database of speech with the discarded speech data and generating an automatic speech recognition model using the text-to-speech waveforms.