Sound Rate Modification Using Phoneme Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional sound rate modification techniques often result in unnatural-sounding audio when applied to speech, as they alter both pitch and timing, leading users to avoid these methods.
Innovation Solution
The implementation of sound rate rules that reflect a natural sound model, allowing for differential rate modifications based on sound characteristics such as pauses and speech components, ensuring that the modified sound data maintains a natural tone.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If conventional sampling rate modification techniques are applied to speech, then the rate of sound output is modified, but the speech sounds unnatural due to simultaneous changes in pitch and timing
Solution Approach 1:
The patent segments the audio signal into individual phonemes and applies rate modification to each phoneme separately rather than to the entire audio signal uniformly. This segmentation allows pitch and timing to be modified independently for each phoneme, preventing the unnatural sounding effects that occur when the entire signal is modified at once.
Solution Approach 2:
The patent applies different modification rates to different portions of the sound data based on local characteristics. Specifically, different phonemes can be modified at different rates, and the timing adjustments are applied locally to each phoneme rather than uniformly across the entire signal, preserving natural speech characteristics while achieving the desired rate modification.
2Loss of time
If uniform rate modification is applied to all portions of sound data, then the overall playback time is adjusted, but the audio quality deteriorates due to unnatural pitch and timing alterations
Solution Approach 1:
The patent employs dynamic rate modification where the modification rate varies over time based on the specific phoneme being processed. Each phoneme can have its own optimal modification rate applied, allowing the system to adapt dynamically to the local characteristics of the speech signal rather than applying a static uniform rate throughout.
Solution Approach 2:
The patent changes multiple parameters (pitch, timing, duration) independently and selectively for each phoneme rather than changing a single sampling rate parameter uniformly. This allows precise control over which parameters are modified and to what extent, preserving audio quality while achieving the desired playback time adjustment.
Data Source
AI summary
Sound rate modification techniques are described. In one or more implementations, an indication is received of an amount that a rate of output of sound data is to be modified. One or more sound rate rules are applied to the sound data that, along with the received indication, are usable to calculate different rates at which different portions of the sound data are to be modified, respectively. The sound data is then output such that the calculated rates are applied.


