Audio Fingerprint Call State Detection for Human vs Machine Responses

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for determining the state of a telephone call, such as whether it connects to a human or machine, are inefficient, inaccurate, and resource-intensive, often failing to distinguish between different types of machine responses and human responses, leading to suboptimal decision-making and poor user experience.

Innovation Solution

A system that generates audio fingerprints of calls, classifies them using machine learning models, and matches them against a fingerprint database to identify actionable machine responses, non-actionable machine responses, or human responses, enabling quick and accurate detection and appropriate actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional call state detection methods are used, then the system can determine whether a call connected, but the detection is time-consuming, processing-intensive, and inaccurate in distinguishing between different response types

Engineering Contradiction:
Improveaccuracy of call state detectionVSAvoidtime for call state detection
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the call audio into multiple samples and generates fingerprints for each segment independently. This allows parallel processing of different audio segments, reducing overall detection time while maintaining accuracy through comprehensive analysis of the entire call duration

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates audio fingerprints as compact representations (copies) of the original audio samples. These fingerprints serve as efficient proxies that capture essential acoustic features without requiring analysis of the full audio data, significantly reducing processing time and computational resources

Inventive Principle:
Principle #26Copying

2Measurement precision

If traditional call state detection methods are used, then the system can determine call connection status, but the processing is resource-intensive and inaccurate in distinguishing actionable machine responses from human responses

Engineering Contradiction:
Improveaccuracy of response type classificationVSAvoidprocessing resources for call analysis
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent creates compact audio fingerprints that serve as efficient representations of the original audio. These fingerprints reduce the data volume requiring processing while preserving the essential acoustic features needed for accurate classification of response types

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces traditional complex audio analysis methods with a fingerprint-based matching system. Instead of analyzing the full audio signal through intensive processing, the system uses fingerprint comparison against a database of known response patterns, significantly reducing computational resources while improving accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If the system analyzes full audio samples to distinguish between intermediary audio and target call recipient audio, then accuracy may improve, but the processing time and computational load increase significantly

Engineering Contradiction:
Improveaccuracy of audio source identificationVSAvoidcall processing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent creates compact fingerprints as representations of full audio samples. This allows the system to maintain high accuracy in identifying audio sources while dramatically reducing the computational burden compared to analyzing complete audio files

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent divides calls into multiple audio samples and generates fingerprints for each segment. This segmentation enables parallel processing of multiple call segments simultaneously, increasing overall system throughput while maintaining comprehensive analysis coverage

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12598253B2Systems and methods for media analysis for call state detection
Publication Date: 2026.04.07 INFOBIP LTD
  • US12598253B2 patent drawing
  • US12598253B2 patent drawing
  • US12598253B2 patent drawing

AI summary

A method includes: initiating a call to a telephone number based on a request from a call originator; generating a fingerprint of an audio sample of the call; matching the fingerprint to a record in a fingerprint database; determining, based on the matched record, whether the recorded call is an actionable machine response, a non-actionable machine response, or a human response, as a detected response; providing the detected response to the call originator; receiving, based on the providing the detected response, an action from the call originator; and performing the action associated with the call.