Speech Rate Adjustment Using Syllable Timing and Smoothing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio playback systems struggle to maintain a consistent speech rate when dealing with multiple speakers or a single speaker who varies their speech rate, leading to difficulties in hearing and requiring manual adjustments by users with hearing or neurological disorders.
Innovation Solution
A method and system that applies syllable onset analysis to determine an average inter-syllable time, adjusts the buffer period based on a target speech rate, and uses smoothing filters and overlapping buffer periods to ensure a consistent speech rate without audio artifacts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If audio playback speed is increased to save time, then listening efficiency is improved, but speech intelligibility deteriorates when multiple speakers or varying speech rates are present
Solution Approach 1:
The system dynamically adjusts the playback speed in real-time based on detected speech characteristics. Instead of applying a fixed speed increase, the processor continuously monitors inter-syllable timing and adapts the playback rate to maintain both efficiency gains and speech intelligibility across different speaking styles and multiple speakers.
Solution Approach 2:
The system changes the timing parameter of audio playback based on detected speech patterns. By analyzing inter-syllable intervals and adjusting playback speed accordingly, the system optimizes both listening efficiency and intelligibility by matching the playback rate to the actual speech rate being delivered.
2Device complexity
If a fixed playback speed adjustment is applied, then processing simplicity is maintained, but adaptability to different speakers and speech patterns is reduced
Solution Approach 1:
The system performs self-adjustment by automatically detecting speech characteristics and modifying playback speed without user intervention. The processor continuously analyzes the audio stream, identifies speech patterns, and autonomously optimizes playback rate for different speakers and speech styles, eliminating the need for manual configuration.
Solution Approach 2:
The system implements a feedback loop where the detected inter-syllable timing information is used to continuously adjust playback speed. The processor monitors speech characteristics in real-time and uses this feedback to dynamically optimize playback rate, ensuring adaptability to varying speech patterns while maintaining relatively simple processing.
Data Source
AI summary
According to one embodiment, a method, computer system, and computer program product for adjusting speech rate for an audio input is provided. The present invention may include applying syllable onset analysis to speech input in a buffer period; determining an average inter-syllable time for the buffer period; determining a rate adjustment required for the average inter-syllable time of the buffer period to conform to a target speech rate; applying a smoothing filter to smooth the rate adjustments across multiple sequential buffer periods; and adjusting the buffer period based on the smoothed rate adjustment.


