Speech Source Classification Using ASR and Sliding Window Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for distinguishing between human and machine-generated speech require extensive training data and are not suitable for real-time or almost real-time applications, limiting their feasibility in situations where rapid classification is necessary, such as automating hold times or screening spam calls.

Innovation Solution

A system that processes audio segments into text using ASR, applies a sliding window analysis with configurable size, and evaluates key features like keyword presence, speech speed, and silence patterns to classify speech as human or machine-generated, enabling near real-time detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional approaches use extensive training data and speaker identification methods, then accuracy in distinguishing human from machine-generated speech is improved, but the time required for determination increases making real-time classification infeasible

Engineering Contradiction:
Improveaccuracy of speech source determinationVSAvoidtime required for speech classification
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts only the most critical acoustic features (pitch, formant frequencies, spectral characteristics) that are most indicative of human versus machine-generated speech, rather than analyzing the entire speech signal or using comprehensive training data. This selective extraction maintains classification accuracy while dramatically reducing processing time for real-time application.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The speech signal is segmented into short analysis windows (e.g., 20-50 milliseconds) that can be processed independently and rapidly. This segmentation allows the system to make classification decisions on small, manageable portions of speech in real-time, rather than requiring analysis of entire sentences or paragraphs.

Inventive Principle:
Principle #1Segmentation

2Reliability

If conventional methods analyze entire sections of text or audio, then comprehensive classification is achieved, but the approach cannot be applied in real-time situations

Engineering Contradiction:
Improvecompleteness of speech analysisVSAvoidreal-time classification capability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary analysis by immediately extracting key acoustic features as speech arrives, rather than waiting to collect complete sentences or sections. This preliminary extraction of critical features enables real-time classification decisions to be made as speech is being spoken, maintaining both reliability and productivity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial action by analyzing only the most discriminative acoustic features rather than performing complete speech analysis. By focusing on specific formant frequencies, pitch contours, and spectral characteristics that are most indicative of human versus machine speech, the system achieves reliable classification without the computational burden of comprehensive analysis.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12579986B2Systems and methods for distinguishing between human speech and machine generated speech
Publication Date: 2026.03.17 OUTBOUND AI INC
  • US12579986B2 patent drawing
  • US12579986B2 patent drawing
  • US12579986B2 patent drawing

AI summary

Systems, devices, and methods for determining whether a segment of speech was generated by a human or by a machine, such as a robotic voice that is synthesized and used as part of an IVR system. The disclosed approach can be used to assist in implementing a process to automate the detection of the start and end of a hold time during a call to a call center and in response execute a desired action.