Speech Source Classification Using ASR and Sliding Window Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for distinguishing between human and machine-generated speech require extensive training data and are not suitable for real-time or almost real-time applications, limiting their feasibility in situations where rapid classification is necessary, such as automating hold times or screening spam calls.
Innovation Solution
A system that processes audio segments into text using ASR, applies a sliding window analysis with configurable size, and evaluates key features like keyword presence, speech speed, and silence patterns to classify speech as human or machine-generated, enabling near real-time detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional approaches use extensive training data and speaker identification methods, then accuracy in distinguishing human from machine-generated speech is improved, but the time required for determination increases making real-time classification infeasible
Solution Approach 1:
The patent extracts only the most critical acoustic features (pitch, formant frequencies, spectral characteristics) that are most indicative of human versus machine-generated speech, rather than analyzing the entire speech signal or using comprehensive training data. This selective extraction maintains classification accuracy while dramatically reducing processing time for real-time application.
Solution Approach 2:
The speech signal is segmented into short analysis windows (e.g., 20-50 milliseconds) that can be processed independently and rapidly. This segmentation allows the system to make classification decisions on small, manageable portions of speech in real-time, rather than requiring analysis of entire sentences or paragraphs.
2Reliability
If conventional methods analyze entire sections of text or audio, then comprehensive classification is achieved, but the approach cannot be applied in real-time situations
Solution Approach 1:
The system performs preliminary analysis by immediately extracting key acoustic features as speech arrives, rather than waiting to collect complete sentences or sections. This preliminary extraction of critical features enables real-time classification decisions to be made as speech is being spoken, maintaining both reliability and productivity.
Solution Approach 2:
The patent applies partial action by analyzing only the most discriminative acoustic features rather than performing complete speech analysis. By focusing on specific formant frequencies, pitch contours, and spectral characteristics that are most indicative of human versus machine speech, the system achieves reliable classification without the computational burden of comprehensive analysis.
Data Source
AI summary
Systems, devices, and methods for determining whether a segment of speech was generated by a human or by a machine, such as a robotic voice that is synthesized and used as part of an IVR system. The disclosed approach can be used to assist in implementing a process to automate the detection of the start and end of a hold time during a call to a call center and in response execute a desired action.


