Privacy-Sensitive Speech Model Creation via User Model Aggregation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face privacy concerns due to the collection of personal and confidential information during training data collection, making it difficult to acquire user permission for data storage, especially with increasing security and privacy concerns.

Innovation Solution

A system that processes audio data locally to create derived statistics or updated acoustic models, which are then sent to a server for aggregation, ensuring that personal information is not stored on third-party servers, and using algorithms like ADMM to maximize the objective function while maintaining privacy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio data and transcriptions are collected from users to train acoustic models, then the quality and performance of speech recognition systems is improved, but privacy concerns arise due to storage of personal and confidential information

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprivacy concerns
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the essential statistical properties and acoustic parameters from the audio data, separating the useful training information from the sensitive personal content. By taking out only what is needed (acoustic statistics, spectral features) and leaving behind the identifiable personal information, the system achieves both improved speech recognition and privacy protection.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the audio data from its original form into different parameter representations (spectral parameters, statistical moments, acoustic features) that retain the acoustic information needed for training while removing personally identifiable information. This parameter transformation allows the same data to serve dual purposes: maintaining recognition accuracy while protecting user privacy.

Inventive Principle:
Principle #35Parameter changes

2Object-affected harmful factors

If user permission is required for data storage due to privacy concerns, then privacy protection is improved, but it becomes difficult or impossible to acquire sufficient training data

Engineering Contradiction:
Improveprivacy protectionVSAvoidtraining data volume
Core Design Contradiction:
Object-affected harmful factorsVSQuantity of substance

Solution Approach 1:

The patent introduces an intermediary processing step that transforms raw audio data into anonymized statistical representations before storage or transmission. This intermediary form acts as a mediator between the need for extensive training data and the requirement for privacy protection, allowing data to be shared and aggregated without requiring explicit user permission for each data point.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates copies of the acoustic information in the form of statistical parameters and spectral features that are functionally equivalent for training purposes but contain no personally identifiable information. These copies can be freely shared and aggregated across multiple users, providing sufficient training data volume without compromising privacy.

Inventive Principle:
Principle #26Copying

3Measurement precision

If raw audio data is stored on third-party servers for model training, then acoustic model quality is improved, but security risks increase due to storage of sensitive information

Engineering Contradiction:
Improveacoustic model qualityVSAvoiddata security
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent extracts only the essential acoustic statistics and parameters needed for model training, removing all personally identifiable information before storage on third-party servers. This extraction process maintains acoustic model quality while eliminating security risks associated with storing sensitive audio data.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent treats the anonymized statistical data as disposable training material that can be freely stored, shared, and discarded on third-party servers without security concerns. Since the data contains no personal information, it can be handled like temporary computational artifacts rather than sensitive personal data, greatly reducing security requirements.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Data Source

PatentUS9424836B2Privacy-sensitive speech model creation via aggregation of multiple user models
Publication Date: 2016.08.23 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9424836B2 patent drawing
  • US9424836B2 patent drawing
  • US9424836B2 patent drawing

AI summary

Techniques disclosed herein include systems and methods for privacy-sensitive training data collection for updating acoustic models of speech recognition systems. In one embodiment, the system locally creates adaptation data from raw audio data. Such adaptation can include derived statistics and/or acoustic model update parameters. The derived statistics and/or updated acoustic model data can then be sent to a speech recognition server or third-party entity. Since the audio data and transcriptions are already processed, the statistics or acoustic model data is devoid of any information that could be human-readable or machine readable such as to enable reconstruction of audio data. Thus, such converted data sent to a server does not include personal or confidential information. Third-party servers can then continually update speech models without storing personal and confidential utterances of users.