Loudspeaker Playback Detection via Frequency Band Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice biometrics systems are vulnerable to replay attacks where a malicious party records and plays back an enrolled user's speech to gain unauthorized access, and existing methods fail to effectively distinguish between live and recorded speech.

Innovation Solution

A method that analyzes audio signals by separating them into frequency bands, comparing signal content, and identifying frequency-based variations indicative of loudspeaker use, such as non-linearities greater at lower frequencies, to determine if the sound was generated by a loudspeaker, using statistical metrics and machine learning techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice biometrics systems accept speech input for authentication, then user convenience and access speed are improved, but the system becomes vulnerable to replay attacks using recorded speech

Engineering Contradiction:
Improveuser convenienceVSAvoidsecurity against replay attacks
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system performs preliminary analysis of the audio signal to detect characteristics of loudspeaker playback before making an authentication decision. By analyzing frequency band variations and non-linearities in advance, the system can reject replay attacks before they compromise security, while still maintaining convenient voice-based authentication for legitimate users.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If the system analyzes audio signals in detail to detect replay attacks, then security is improved, but processing time and computational complexity increase

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The audio signal is segmented into multiple frequency bands for parallel analysis. By dividing the spectrum into distinct bands and analyzing each independently for non-linearities and playback characteristics, the system achieves comprehensive detection coverage while enabling efficient parallel processing that reduces overall computation time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The analysis focuses on specific local characteristics of the audio signal that are most indicative of playback, such as non-linearities in particular frequency bands and specific statistical metrics. Rather than analyzing the entire signal uniformly, the system concentrates computational effort on the most discriminative features, improving detection accuracy while reducing overall processing requirements.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11051117B2Detection of loudspeaker playback
Publication Date: 2021.06.29 CIRRUS LOGIC INC
  • US11051117B2 patent drawing
  • US11051117B2 patent drawing
  • US11051117B2 patent drawing

AI summary

A method of determining whether a sound has been generated by a loudspeaker comprises receiving an audio signal representing at least a part of the sound. The audio signal is separated into different frequency bands. The signal content of different frequency bands are compared. Based on said comparison, frequency-based variations in signal content indicative of use of a loudspeaker are identified.