Obfuscating Training Data for Audio Analysis Privacy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Audio analysis systems, such as speech recognition and language identification, rely heavily on large training datasets, which can be costly and time-consuming to source, and often contain confidential information that cannot be shared without consent, posing a challenge in maintaining data privacy.

Innovation Solution

The method involves obfuscating training data by randomizing and reorganizing annotated feature vectors generated from audio and text transcripts, ensuring that the audio analysis system cannot determine the content or subject matter of the original data, while still utilizing the data for model training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If large training datasets are used to improve accuracy of audio analysis systems, then the accuracy and representativeness of the system improves, but data privacy and confidentiality are compromised

Engineering Contradiction:
ImproveaccuracyVSAvoiddata privacy risk
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the essential acoustic features and state information from the training data while removing or obscuring the original audio content and text transcripts. This allows the system to retain the useful training information needed for accuracy while eliminating the confidential content that poses privacy risks.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates obfuscated copies of the training data that preserve the acoustic patterns and linguistic features needed for model training, but replace the original confidential audio and text with randomized or synthetic representations that cannot be traced back to the source material.

Inventive Principle:
Principle #26Copying

2Productivity

If existing confidential conversations and transcripts are used for training, then the cost and time of data collection is reduced, but consent and privacy requirements cannot be met

Engineering Contradiction:
Improvedata collection efficiencyVSAvoidprivacy compliance
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces an obfuscation process as an intermediary between the confidential training data and the audio analysis system. This intermediary transforms the data into an intermediate representation that maintains training value while removing privacy concerns, enabling the use of existing confidential data without direct consent.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameters of the training data by transforming audio signals into acoustic feature vectors and text into state sequences, altering the data representation from recognizable content to abstract numerical representations that preserve statistical patterns but eliminate identifiable information.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If original training data is shared with audio analysis systems, then model training quality improves, but data security and confidentiality controls are violated

Engineering Contradiction:
Improvemodel training qualityVSAvoiddata security risk
Core Design Contradiction:
Manufacturing precisionVSObject-generated harmful factors

Solution Approach 1:

The patent converts the harmful aspect of confidential data (its sensitivity and identifiability) into a benefit by using the obfuscation process to create training data that is equally valuable for model training but inherently secure. The very process that removes identifying information preserves the acoustic and linguistic patterns needed for high-quality training.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentEP3262634B1Obfuscating training data
Publication Date: 2019.04.03 LONGSAND LTD
  • EP3262634B1 patent drawingFigure 1
  • EP3262634B1 patent drawingFigure 2
  • EP3262634B1 patent drawingFigure 3

AI summary

Examples disclosed herein involve obfuscating training data. An example method includes computing a sequence of acoustic features from audio data of training data, the training data comprising the audio data and a corresponding text transcript; mapping the acoustic features to acoustic model states to generate annotated feature vectors, the annotated feature vectors comprising the acoustic features and corresponding context from the text transcript; and providing a randomized sequence of the annotated feature vectors as obfuscated training data to an audio analysis system.