Audio Fingerprinting Using Spectral Peaks for Robocall Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems struggle to effectively and efficiently analyze communications media, such as robocalls, to generate fingerprints that can identify similar calls while maintaining privacy and efficiently storing and retrieving this information for detection and classification.
Innovation Solution
A method involving the analysis of power spectral density values to generate audio fingerprints by identifying dominant frequency peaks and positions in the audio signal, which are then used to create a fingerprint-set that can identify similar communications, with privacy preservation and efficient storage and retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If audio fingerprinting is used to identify robocalls by analyzing media content, then detection accuracy is improved, but system complexity increases and processing time is extended
Solution Approach 1:
The patent extracts only the essential acoustic features (power spectral density values, dominant frequency peaks, time positions) from the audio signal to create fingerprints, rather than analyzing the entire audio content. This extraction approach maintains detection accuracy while significantly reducing system complexity and processing requirements.
Solution Approach 2:
The audio signal is divided into multiple time segments, and fingerprints are generated for each segment independently. This segmentation allows the system to process audio incrementally, reducing the computational burden on any single processing unit while maintaining overall detection accuracy through aggregation of segment-level results.
2Measurement precision
If detailed audio analysis is performed to generate accurate fingerprints, then detection precision is improved, but data storage requirements increase
Solution Approach 1:
The patent extracts only the critical parameters (power spectral density values, dominant frequency peaks, and their time positions) from the full audio signal to create compact fingerprints. This extraction maintains the essential information needed for accurate robocall identification while reducing the data volume to a minimal set of numerical values that can be efficiently stored and compared.
3Reliability
If fuzzy fingerprint matching is implemented to handle degraded communications, then detection robustness is improved, but processing time increases
Solution Approach 1:
The patent implements fuzzy matching by allowing approximate matches on the extracted fingerprint parameters rather than requiring exact matches. This partial matching approach tolerates the variations introduced by network degradation while maintaining efficient processing by comparing only the essential extracted features rather than performing exhaustive analysis of degraded audio signals.
Data Source
AI summary
The present invention relates to methods, systems, and apparatus for processing audio signals. An exemplary method embodiment includes the steps of: removing silence from an audio signal; determining, for a plurality of time segments of the audio signal, power spectral density values of the audio signal for each of a plurality of N different frequency bins, N being an integer greater than 1; identifying (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks; and generating a first audio fingerprint from at least some of the identified plurality of dominant frequency peaks and the identified positions in the audio signal corresponding to the identified peaks. In various embodiments, audio fingerprints are generated from an audio signal of call and then used to determine if the call is a robocall or SPAM call.


