Multi-Channel Speech Privacy Processing for ASR Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech processing systems face challenges in protecting Personal Identifiable Information (PII) while maintaining accurate automated speech recognition (ASR) performance, as current voice style transfer and voice conversion methods often degrade ASR accuracy by not effectively modeling background acoustics and may result in voices that are not well synthesized or indistinguishable from training data.

Innovation Solution

A multi-channel speech privacy process that uses a voice style transfer system to map the original speaker's voice to a licensed voice, incorporating a text-to-speech system and speaker selection to enhance privacy, while preserving and improving Word Error Rate (WER) by selectively combining original and de-identified speech data, and using a canonical target speaker for voice conversion to ensure better anonymization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If voice style transfer or voice conversion algorithms are used to convert the voice to that of a licensed actor, then privacy protection is improved, but automated speech recognition accuracy deteriorates

Engineering Contradiction:
Improveprivacy protectionVSAvoidautomated speech recognition accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent divides the speech processing into separate channels: one channel processes the original speech signal for ASR accuracy, while another channel applies voice style transfer for privacy protection. This segmentation allows each channel to optimize for its specific function without compromising the other.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary multi-channel processing system that receives the original speech signal and produces multiple output channels. This intermediary system enables the speech content to be preserved for ASR while simultaneously generating a privacy-protected version through voice conversion.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If voice conversion is applied to protect speaker identity, then speaker anonymity is improved, but speech processing quality deteriorates

Engineering Contradiction:
Improvespeaker anonymityVSAvoidspeech processing quality
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent segments the speech processing system into multiple independent channels, where one channel maintains the original speech quality for processing tasks while another channel applies voice conversion for anonymity. This allows speech processing quality to be preserved in the original channel while achieving anonymity in the converted channel.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a dimensional aspect to speech processing by creating parallel processing channels. Instead of modifying the original signal in place, the system processes the speech in multiple dimensions (original and converted channels), allowing quality preservation in one dimension while achieving anonymity in another.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If conventional voice conversion methods are used, then privacy protection is achieved, but Word Error Rate increases

Engineering Contradiction:
Improveprivacy protectionVSAvoidWord Error Rate
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent implements segmentation by creating separate processing channels for privacy protection and speech recognition. The original speech channel maintains high ASR accuracy while the converted speech channel provides privacy protection, thereby reducing Word Error Rate in the original channel while achieving privacy goals.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies voice conversion selectively rather than universally. By using partial action (applying conversion only where needed for privacy while preserving original signals for ASR), the system achieves privacy protection without excessively degrading speech recognition accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20240296826A1System and Method for Multi-Channel Speech Privacy Processing
Publication Date: 2024.09.05 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240296826A1 patent drawing
  • US20240296826A1 patent drawing
  • US20240296826A1 patent drawing

AI summary

A method, computer program product, and computing system for receiving a speech signal from a single microphone. A sensitive speech component is identified from the speech signal. In response to identifying the sensitive speech component, a filtered speech signal is generated by removing the sensitive speech component from the speech signal. A voice style transfer of the speech signal is generated. Speech processing is performed on the filtered speech signal and the voice style transfer of the speech signal.