Speech Detection via Time-Domain Amplitude Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech detection algorithms are either inaccurate or require complex analysis, being sensitive to audio quality changes and computational intensive, limiting their implementation in systems with limited computing power, and fail to quantify the amount of speech in an audio signal.
Innovation Solution
A method that calculates the audio signal speech grade by analyzing the dynamic behavior and ratio of high and low volume parts, using segment and block values to determine the presence and quantity of speech, providing a computationally lightweight solution that is amplitude and language independent.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If complex frequency analysis or phonetic comparison is used for speech detection, then measurement precision is improved, but use of energy increases and device complexity increases
Solution Approach 1:
The patent extracts only the essential dynamic characteristics (zero-crossing rate, peak detection, RMS variations) from the audio signal that are sufficient for speech detection, rather than performing complete frequency analysis or phonetic comparison. This selective extraction of critical features maintains detection accuracy while dramatically reducing computational requirements and energy consumption.
Solution Approach 2:
Instead of analyzing the audio signal in the frequency domain through complex transforms, the patent inverts the approach by working directly in the time domain with simple amplitude-based metrics. This inversion from frequency-domain complex analysis to time-domain simple measurement achieves comparable speech detection performance with minimal computational overhead.
2Measurement precision
If complex frequency analysis or phonetic comparison is used for speech detection, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent extracts only the essential dynamic characteristics (zero-crossing rate, peak detection, RMS variations) from the audio signal that are sufficient for speech detection, rather than performing complete frequency analysis or phonetic comparison. This selective extraction of critical features maintains detection accuracy while dramatically reducing computational requirements and energy consumption.
Solution Approach 2:
Instead of analyzing the audio signal in the frequency domain through complex transforms, the patent inverts the approach by working directly in the time domain with simple amplitude-based metrics. This inversion from frequency-domain complex analysis to time-domain simple measurement achieves comparable speech detection performance with minimal computational overhead.
3Use of energy by moving object
If simple speech detection algorithms are used, then use of energy is reduced, but measurement precision deteriorates
Solution Approach 1:
The patent creates simplified copies of speech signal characteristics by computing basic time-domain features (zero-crossing rate, peak amplitude, RMS) that replicate the essential dynamic behavior of speech. These simplified proxies capture speech presence accurately enough for detection purposes while requiring minimal computational resources compared to full spectral analysis.
Solution Approach 2:
The patent changes the measurement parameters from complex frequency-domain representations to simple time-domain amplitude characteristics. By monitoring how these basic parameters change over time (dynamic behavior), the system achieves speech detection accuracy comparable to complex algorithms while consuming minimal computational power.
4Device complexity
If simple speech detection algorithms are used, then device complexity is reduced, but measurement precision deteriorates
Solution Approach 1:
The patent creates simplified copies of speech signal characteristics by computing basic time-domain features (zero-crossing rate, peak amplitude, RMS) that replicate the essential dynamic behavior of speech. These simplified proxies capture speech presence accurately enough for detection purposes while requiring minimal computational resources compared to full spectral analysis.
Solution Approach 2:
The patent changes the measurement parameters from complex frequency-domain representations to simple time-domain amplitude characteristics. By monitoring how these basic parameters change over time (dynamic behavior), the system achieves speech detection accuracy comparable to complex algorithms while consuming minimal computational power.
5Loss of time
If speech detection is not performed, then processing time is reduced, but loss of information increases due to recording noise or silence unknowingly
Solution Approach 1:
The patent performs preliminary speech detection analysis using simple time-domain features before committing to full audio processing or recording. This preliminary check quickly identifies whether speech is present, preventing wasteful processing of silent or noisy segments while ensuring compliance-critical speech segments are captured and processed appropriately.
Solution Approach 2:
The patent creates simplified copies of speech signal characteristics by computing basic time-domain features (zero-crossing rate, peak amplitude, RMS) that replicate the essential dynamic behavior of speech. These simplified proxies capture speech presence accurately enough for detection purposes while requiring minimal computational resources compared to full spectral analysis.
Data Source
AI summary
A system and method for determining an amount of speech in an audio signal may include for example: obtaining segments of the audio signal, wherein the segments are grouped into blocks; for each one of the segments, calculating a segment value indicative of an amplitude of the audio signal of a respective segment; for each one of the blocks calculating a block value indicative of the amplitude of the audio signal of a respective block; and calculating an audio signal speech grade based on segment values and block values, wherein the audio signal speech grade is indicative of the amount of speech in the audio signal.


