Audio Scrambling for Speech Privacy with Environmental Sound Preservation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic devices face challenges in scrambling audio data to prevent unwanted speech from being perceived while allowing environmental sounds to be retained, as the frequency bands of speech and environmental sounds often overlap, making it difficult to distinguish and separate them effectively.

Innovation Solution

An electronic device with a processor that acquires video and audio data, scrambles the audio data in non-time order, and stores it alongside the video data, using a predetermined or varying time interval, such as through a Hash-based Message Authentication Codes (HMAC)-based One-time Password (HOTP) algorithm, to ensure that speech content remains ambiguous while environmental sounds remain perceptible.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If audio data is scrambled to prevent speech from being perceived, then speech privacy is improved, but environmental sound perceptibility deteriorates

Engineering Contradiction:
Improvespeech privacy protectionVSAvoidenvironmental sound distortion
Core Design Contradiction:
Object-affected harmful factorsVSObject-generated harmful factors

Solution Approach 1:

The patent segments the audio data into speech components and environmental sound components, applying different processing strategies to each. The speech components are scrambled using time-reversal and frequency transformation to protect privacy, while the environmental sound components are preserved or minimally processed to maintain perceptibility. This segmentation allows selective scrambling that addresses the contradiction between speech privacy and environmental sound preservation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by treating different parts of the audio spectrum differently. High-frequency components associated with speech are subjected to strong scrambling transformations, while low-frequency components associated with environmental sounds are preserved or lightly processed. This localized differentiation enables the system to protect speech privacy without significantly degrading environmental sound quality.

Inventive Principle:
Principle #3Local quality

2Object-affected harmful factors

If frequency band separation is used to distinguish speech and environmental sound, then speech privacy is improved, but processing complexity increases

Engineering Contradiction:
Improvespeech privacy protectionVSAvoidsignal processing complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The patent employs time-reversal vibration patterns as a core scrambling mechanism. By reversing the time-order of audio segments and applying vibrational transformation patterns, the system creates scrambled speech that is unintelligible while preserving the temporal structure needed for environmental sound recognition. This mechanical vibration approach provides effective scrambling without requiring complex frequency-domain analysis.

Inventive Principle:
Principle #18Mechanical vibration

Solution Approach 2:

The patent changes temporal parameters of the audio signal through time-reversal and variable time-interval scrambling. Instead of relying on complex frequency band separation, the system transforms the temporal structure of speech while preserving environmental sound characteristics. This parameter transformation approach simplifies the overall processing architecture while maintaining speech privacy protection.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11991267B2Electronic device for speech security with less environmental sound distortion and operating method thereof
Publication Date: 2024.05.21 THINKWARE
  • US11991267B2 patent drawing
  • US11991267B2 patent drawing
  • US11991267B2 patent drawing

AI summary

An electronic device and an operating method thereof according to various example embodiments may be configured to acquire video data and audio data, to scramble audio data of a time interval (ΔT) in non-time order, and to store the video data and the scrambled audio data together. According to an example embodiment, the time interval (ΔT) may be fixed to be the same with respect to continuous audio data. According to another example embodiment, the time interval (ΔT) may vary using a function of receiving an input of a secret key and order and outputting a unique value, such as a Hash-based Message Authentication Codes (HMAC)-based One-time Password (HOTP) algorithm.