Anonymized Speech Recognition via Frequency Domain Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems face challenges in complying with regulatory requirements, such as COPPA, which consider child speech data as personal information, necessitating rigorous consent methods when sharing with third parties, while also needing to perform interactive operations without capturing and storing raw audio files.

Innovation Solution

A method for anonymizing speech data by converting raw waveforms into frequency components and removing identifying features, allowing for speech recognition without retaining raw audio files, thus enabling compliance with less stringent consent requirements and protecting user privacy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If raw audio files are captured and stored for speech recognition, then speech recognition functionality is maintained, but user privacy protection deteriorates and regulatory compliance becomes difficult

Engineering Contradiction:
Improvespeech recognition functionalityVSAvoiduser privacy exposure
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts and removes identifying features from speech waveforms by converting to frequency domain and eliminating specific frequency components that carry speaker identification information while preserving speech content, thereby maintaining speech recognition functionality while protecting user privacy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The speech waveform is segmented into frequency components through Fourier transform, allowing selective removal of identifying frequency components while retaining speech recognition-relevant features, thus separating privacy-protection-critical elements from functionality-critical elements

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If speech data is shared with third parties for processing, then speech recognition accuracy is improved, but regulatory compliance requirements become more stringent

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidconsent acquisition complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary anonymization processing on speech waveforms before sharing with third parties, converting to frequency domain and removing identifying features in advance, so that third parties receive only anonymized data requiring milder consent under COPPA regulations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary anonymization processing step between speech capture and third-party sharing, using frequency domain transformation and selective component removal as a mediator to decouple speech recognition accuracy from privacy exposure risks

Inventive Principle:
Principle #24Intermediary (Mediator)

3Object-affected harmful factors

If identifying features are removed from speech waveforms, then user privacy protection is improved, but speech data utility for analysis deteriorates

Engineering Contradiction:
Improveuser privacy exposureVSAvoidspeech data utility
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent applies local quality modification by selectively removing only specific frequency components that carry speaker identification information while preserving other frequency components that contain speech content, achieving differential treatment of waveform components based on their functional importance

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9437207B2Feature extraction for anonymized speech recognition
Publication Date: 2016.09.06 CHATTERBOX CAPITAL LLC
  • US9437207B2 patent drawing
  • US9437207B2 patent drawing
  • US9437207B2 patent drawing

AI summary

Various of the disclosed embodiments relate to systems and methods for extracting audio information, e.g. a textual description of speech, from a speech recording while retaining the anonymity of the speaker. In certain embodiments, a third party may perform various aspects of the anonymization and speech processing. Certain embodiments facilitate anonymization in compliance with various legislative requirements even when third parties are involved.