Sound Rate Modification Using Phoneme Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional sound rate modification techniques often result in unnatural-sounding audio when applied to speech, as they alter both pitch and timing, leading users to avoid these methods.

Innovation Solution

The implementation of sound rate rules that reflect a natural sound model, allowing for differential rate modifications based on sound characteristics such as pauses and speech components, ensuring that the modified sound data maintains a natural tone.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If conventional sampling rate modification techniques are applied to speech, then the rate of sound output is modified, but the speech sounds unnatural due to simultaneous changes in pitch and timing

Engineering Contradiction:
Improverate of sound outputVSAvoidnaturalness of speech
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent segments the audio signal into individual phonemes and applies rate modification to each phoneme separately rather than to the entire audio signal uniformly. This segmentation allows pitch and timing to be modified independently for each phoneme, preventing the unnatural sounding effects that occur when the entire signal is modified at once.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different modification rates to different portions of the sound data based on local characteristics. Specifically, different phonemes can be modified at different rates, and the timing adjustments are applied locally to each phoneme rather than uniformly across the entire signal, preserving natural speech characteristics while achieving the desired rate modification.

Inventive Principle:
Principle #3Local quality

2Loss of time

If uniform rate modification is applied to all portions of sound data, then the overall playback time is adjusted, but the audio quality deteriorates due to unnatural pitch and timing alterations

Engineering Contradiction:
Improveplayback timeVSAvoidaudio quality
Core Design Contradiction:
Loss of timeVSManufacturing precision

Solution Approach 1:

The patent employs dynamic rate modification where the modification rate varies over time based on the specific phoneme being processed. Each phoneme can have its own optimal modification rate applied, allowing the system to adapt dynamically to the local characteristics of the speech signal rather than applying a static uniform rate throughout.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes multiple parameters (pitch, timing, duration) independently and selectively for each phoneme rather than changing a single sampling rate parameter uniformly. This allows precise control over which parameters are modified and to what extent, preserving audio quality while achieving the desired playback time adjustment.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10249321B2Sound rate modification
Publication Date: 2019.04.02 ADOBE INC
  • US10249321B2 patent drawing
  • US10249321B2 patent drawing
  • US10249321B2 patent drawing

AI summary

Sound rate modification techniques are described. In one or more implementations, an indication is received of an amount that a rate of output of sound data is to be modified. One or more sound rate rules are applied to the sound data that, along with the received indication, are usable to calculate different rates at which different portions of the sound data are to be modified, respectively. The sound data is then output such that the calculated rates are applied.