Audio Fingerprinting Using Spectral Peaks for Robocall Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems struggle to effectively and efficiently analyze communications media, such as robocalls, to generate fingerprints that can identify similar calls while maintaining privacy and efficiently storing and retrieving this information for detection and classification.

Innovation Solution

A method involving the analysis of power spectral density values to generate audio fingerprints by identifying dominant frequency peaks and positions in the audio signal, which are then used to create a fingerprint-set that can identify similar communications, with privacy preservation and efficient storage and retrieval.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If audio fingerprinting is used to identify robocalls by analyzing media content, then detection accuracy is improved, but system complexity increases and processing time is extended

Engineering Contradiction:
Improverobocall detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts only the essential acoustic features (power spectral density values, dominant frequency peaks, time positions) from the audio signal to create fingerprints, rather than analyzing the entire audio content. This extraction approach maintains detection accuracy while significantly reducing system complexity and processing requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The audio signal is divided into multiple time segments, and fingerprints are generated for each segment independently. This segmentation allows the system to process audio incrementally, reducing the computational burden on any single processing unit while maintaining overall detection accuracy through aggregation of segment-level results.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If detailed audio analysis is performed to generate accurate fingerprints, then detection precision is improved, but data storage requirements increase

Engineering Contradiction:
Improvefingerprint matching precisionVSAvoiddata storage volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the critical parameters (power spectral density values, dominant frequency peaks, and their time positions) from the full audio signal to create compact fingerprints. This extraction maintains the essential information needed for accurate robocall identification while reducing the data volume to a minimal set of numerical values that can be efficiently stored and compared.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If fuzzy fingerprint matching is implemented to handle degraded communications, then detection robustness is improved, but processing time increases

Engineering Contradiction:
Improvedetection robustness to degradationVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements fuzzy matching by allowing approximate matches on the extracted fingerprint parameters rather than requiring exact matches. This partial matching approach tolerates the variations introduced by network degradation while maintaining efficient processing by comparing only the essential extracted features rather than performing exhaustive analysis of degraded audio signals.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12586592B2Methods and apparatus for generating audio fingerprints for calls using power spectral density values
Publication Date: 2026.03.24 RIBBON COMMUNICATIONS OPERATING CO INC
  • US12586592B2 patent drawing
  • US12586592B2 patent drawing
  • US12586592B2 patent drawing

AI summary

The present invention relates to methods, systems, and apparatus for processing audio signals. An exemplary method embodiment includes the steps of: removing silence from an audio signal; determining, for a plurality of time segments of the audio signal, power spectral density values of the audio signal for each of a plurality of N different frequency bins, N being an integer greater than 1; identifying (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks; and generating a first audio fingerprint from at least some of the identified plurality of dominant frequency peaks and the identified positions in the audio signal corresponding to the identified peaks. In various embodiments, audio fingerprints are generated from an audio signal of call and then used to determine if the call is a robocall or SPAM call.