Voice Energy Dialog Analysis for Human Answer Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for analyzing voice calls in call centers are computationally intensive and require complex speech-to-text conversions, making them inefficient and costly.

Innovation Solution

A computer-implemented method that analyzes voice energy levels of callers and agents using decibel measurements to determine key characteristics of phone calls, such as dialog and talkover, without the need for transcription, enabling efficient real-time analysis of call interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If artificial intelligence and machine learning techniques are used to conduct natural language processing using call transcripts, then analysis accuracy is improved, but computational load and complexity increase

Engineering Contradiction:
Improveanalysis accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential acoustic features (voice energy levels, silence detection, talk-over detection) from the complex audio signal, eliminating the need for full speech-to-text conversion and natural language processing. This extraction approach maintains sufficient analysis accuracy while dramatically reducing computational complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of converting speech to text and then analyzing the text content, the patent inverts the approach by directly analyzing acoustic properties of the voice signal. This inversion bypasses the computationally intensive speech-to-text conversion step while still enabling effective call analysis.

Inventive Principle:
Principle #13The other way round (Inversion)

2Ease of operation

If speech-to-text conversions are performed to prepare transcripts for analysis, then analysis capability is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveanalysis capabilityVSAvoidprocessing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent extracts key interaction characteristics directly from acoustic features such as voice energy levels, silence periods, and talk-over events. This extraction method provides sufficient analysis capability for call center applications without requiring full speech-to-text conversion, thereby reducing processing time.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses simple, computationally inexpensive acoustic measurements (decibel levels, silence detection) instead of expensive and time-consuming speech-to-text conversion. These simple measurements provide the necessary information for effective call analysis.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Measurement precision

If trained machine learning models are deployed to evaluate call transcripts, then prediction accuracy is improved, but training requirements and feature engineering complexity increase

Engineering Contradiction:
Improveprediction accuracyVSAvoidmodel deployment complexity
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent employs simple threshold-based detection rules that do not require external training data or model training phases. The system self-configures by applying predetermined thresholds to acoustic features, eliminating the need for separate training and feature engineering processes while maintaining practical prediction accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the parameters being measured from complex linguistic features to simple acoustic parameters (voice energy level, silence duration, talk-over frequency). This parameter change simplifies the analysis model while maintaining the ability to detect key call characteristics.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12381973B2Dialog analysis using voice energy level
Publication Date: 2025.08.05 INVOCA INC
  • US12381973B2 patent drawing
  • US12381973B2 patent drawing
  • US12381973B2 patent drawing

AI summary

A computer-implemented method for analyzing whether a phone call is answered by a human agent. The computer-implemented method receives phone call audio data of the phone call and separates the phone call audio data into caller stream data and agent stream data that each includes a plurality of frames and calculates decibel level for each frame. In response to measuring alternating groups of frames in the caller stream data and agent stream data that exceed a dialog decibel threshold, the computer-implemented method further identifies a dialog in the phone call audio data. In response to measuring decibel levels that exceed the dialog decibel threshold in corresponding frames in both the caller stream data and agent stream data, the computer-implemented method further identifies talkover in the phone call audio data. Furthermore, in response to identifying the dialog and if a level of talkover in the phone call audio data does not exceed a talkover threshold, the computer-implemented method determines the call is answered by the human agent.