Voice Interface Speaker Differentiation via Sub-Bass Energy Metrics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice interfaces in IoT devices lack effective differentiation between human and electronic speakers, leading to potential unauthorized actions and security vulnerabilities, as any sound-emitting device within range can cause these systems to perform unintended operations.

Innovation Solution

The method programmatically detects sub-bass over-excitation, a feature inherent to electronic speakers due to their design, to differentiate between human and electronic speakers, using signal processing and energy balance metrics to identify and prevent adversarial requests and replayed audio.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice interfaces accept any sound-emitting device as input, then ease of operation is improved, but security is worsened due to unauthorized actions by electronic speakers

Engineering Contradiction:
Improveease of useVSAvoidsecurity
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent changes the detection parameter from general audio signal analysis to specifically targeting sub-bass frequency characteristics (20-80 Hz). By monitoring the energy balance metric in this specific frequency range, the system can distinguish electronic speakers from human voices while maintaining ease of voice interface operation. This parameter-specific approach resolves the contradiction by adding security without complicating the user experience.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If the system implements speaker differentiation to improve security, then reliability is improved, but device complexity increases due to additional detection mechanisms

Engineering Contradiction:
ImprovesecurityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts only the critical sub-bass frequency components (20-80 Hz) from the full audio spectrum for analysis. By focusing detection efforts on this specific frequency range where electronic speakers exhibit characteristic over-excitation, the system achieves reliable speaker differentiation without implementing complex full-spectrum analysis. This extraction approach improves security while minimizing the addition of system complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If the system uses sub-bass detection to identify electronic speakers, then measurement precision is improved, but difficulty of detecting and measuring increases due to specialized signal processing requirements

Engineering Contradiction:
Improvedetection accuracyVSAvoiddetection complexity
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the audio frequency spectrum into distinct bands, specifically isolating the sub-bass range (20-80 Hz) for targeted analysis. By dividing the detection task into frequency-specific segments rather than analyzing the entire spectrum simultaneously, the system achieves high measurement precision in identifying electronic speakers. This segmentation approach improves detection accuracy while managing the complexity through focused, band-limited signal processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11176960B2Method and apparatus for differentiating between human and electronic speaker for voice interface security
Publication Date: 2021.11.16 UNIV OF FLORIDA RESEARCH FOUNDATION INC
  • US11176960B2 patent drawing
  • US11176960B2 patent drawing
  • US11176960B2 patent drawing

AI summary

A system for distinguishing between a human voice generated command and an electronic speaker generated command is provided. An exemplary system comprises a microphone array for receiving an audio signal collection, preprocessing circuitry configured for converting the audio signal collection into processed recorded audio signals, energy balance metric determination circuitry configured for calculating a final energy balance metric based on the processed recorded audio signals, and energy balance metric evaluation circuitry for outputting a command originator signal based at least in part on the final energy balance metric.