Pseudo-Speech Masking for Acoustic Continuity in PII Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for securing sensitive content in speech signals, such as Personal Identifiable Information (PII) and Personal Health Information (PHI), introduce signal discontinuities and acoustic mismatches due to the surrogation of text without corresponding audio, affecting downstream speech processing systems.

Innovation Solution

Implementing a voice converter system that generates pseudo-speech representations for sensitive content, which are unintelligible and reduce acoustic quality, while maintaining consistent speech processing by using predefined acoustic signatures, spectral modifications, and watermarking to ensure privacy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If surrogation of PII elements in transcribed training data is performed without corresponding audio, then privacy protection is improved, but signal discontinuities appear in the speech signals processed by downstream speech processing system

Engineering Contradiction:
Improveprivacy protectionVSAvoidsignal continuity
Core Design Contradiction:
ReliabilityVSStability of the object's composition

Solution Approach 1:

The patent introduces pseudo-speech representations as an intermediary element that mediates between the need for privacy protection and signal continuity. These pseudo-speech segments act as placeholders that maintain the temporal and acoustic structure of the original signal while containing no real PII, thus preventing signal discontinuities that would occur with simple surrogation or redaction.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms sensitive speech segments into pseudo-speech representations by modifying acoustic parameters such as pitch, timbre, and spectral characteristics. This parameter transformation preserves the temporal and structural properties of the original signal (maintaining continuity) while rendering the content unintelligible and protecting privacy.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If text-level surrogation is applied without corresponding audio replacement, then PII removal is achieved, but acoustic mismatching occurs in downstream speech processing

Engineering Contradiction:
ImprovePII removal effectivenessVSAvoidacoustic matching accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

Pseudo-speech representations serve as an intermediary that bridges the gap between text-level surrogation and audio processing. Instead of leaving audio gaps or using unrelated audio segments, the system generates pseudo-speech that acoustically matches the temporal and spectral characteristics of the original segment, ensuring consistent acoustic processing downstream while still removing PII.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system applies parameter changes to transform sensitive speech into pseudo-speech by modifying acoustic features such as fundamental frequency, spectral envelope, and temporal dynamics. These parameter transformations ensure that the pseudo-speech segments acoustically match the surrounding non-sensitive segments, eliminating acoustic mismatching issues while maintaining PII removal effectiveness.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If speech signals are processed to remove sensitive content, then privacy protection is improved, but speech processing accuracy may deteriorate

Engineering Contradiction:
Improveprivacy protectionVSAvoidspeech processing accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent applies controlled parameter changes to transform sensitive speech into pseudo-speech representations. By carefully selecting which acoustic parameters to modify (such as pitch and spectral characteristics) while preserving temporal structure and duration, the system maintains speech processing accuracy for non-sensitive portions while protecting sensitive content through parameter transformation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12531051B2System and method for secure processing of speech signals using pseudo-speech representations
Publication Date: 2026.01.20 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12531051B2 patent drawing
  • US12531051B2 patent drawing
  • US12531051B2 patent drawing

AI summary

A method, computer program product, and computing system for processing a speech signal. A sensitive portion of the speech signal is identified. A pseudo-speech representation of the sensitive portion is generated using a voice converter system. Speech processing is performed on the speech signal and the pseudo-speech representation of the sensitive portion using a speech processing system.