Audio Firewall for Privacy-Safe Voice Command Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Smart audio systems continuously monitor audio data for wake words, leading to unintended sensitive information being sent to remote servers, which can be vulnerable to eavesdropping.

Innovation Solution

An audio firewall system that processes voice inputs locally, identifies wake words to determine whether to process requests locally or anonymize them before sending to remote servers, thereby preventing sensitive information from being transmitted.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the smart audio system continuously sends audio data to the remote server system, then the cloud service can process voice commands, but sensitive information may be transmitted over the network and subjected to eavesdropping

Engineering Contradiction:
Improvevoice command processing capabilityVSAvoidprivacy security risk
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent introduces a local processing intermediary (the audio firewall system) between the microphone and the remote server. This intermediary captures audio data locally, processes it through speech-to-text conversion and wake word detection, and only transmits necessary data to the cloud. This mediator protects privacy by preventing direct transmission of raw audio data while still enabling cloud-based voice command processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the smart audio system continuously monitors audio data for wake words, then wake word detection capability is improved, but the system transmits unnecessary audio data to the remote server

Engineering Contradiction:
Improvewake word detection accuracyVSAvoidnetwork transmission resource consumption
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent applies preliminary action by performing wake word detection locally before any network transmission occurs. The system converts audio data to text locally, detects wake words in the transcribed text, and only transmits data when a wake word is actually detected. This preliminary local processing prevents unnecessary network transmissions while maintaining accurate wake word detection capability.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If the system processes all audio data through the remote server, then comprehensive voice processing is achieved, but network bandwidth and server resources are continuously consumed

Engineering Contradiction:
Improvevoice processing capabilityVSAvoidnetwork and server resource usage
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent segments the voice processing function into two parts: local processing (audio-to-text conversion, wake word detection, and command interpretation) and cloud processing (only when wake words are detected). This segmentation allows the system to maintain comprehensive voice processing capability while dramatically reducing network resource consumption by handling most processing locally.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12283277B2Audio firewall
Publication Date: 2025.04.22 NICE NORTH AMERICA LLC
  • US12283277B2 patent drawing
  • US12283277B2 patent drawing
  • US12283277B2 patent drawing

AI summary

An audio firewall system has a microphone that generates audio data. A speech-to-text engine converts the audio data to text data. The text data is parsed for a service wake word and corresponding content data. The service wake word identifies one of a local security system and a remote assistant server. A text-to-speech engine converts the service wake word and the corresponding content data to converted audio data. The converted audio data is provided to the remote assistant server. The content data is provided to the local security system. The audio firewall system receives a response from the remote assistant server or the local security system and outputs an audio signal corresponding to the response.