Automatic Rate Control for Audio Time Scaling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing media playback technologies fail to adjust playing speeds of audio content in a pitch-correct manner, leading to inadequate audio intelligibility for users as they cannot account for varying speaking rates across different audio sources.

Innovation Solution

A system that divides media data into subsets based on acoustic analysis, allowing for variable playback speeds while maintaining pitch integrity, using a combination of time shift and automatic rate control units to adjust playback speeds according to user-preferred rates of audio utterance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If digital signal processing algorithms are used to play audio signal at selected fast forward speed, then playback speed is improved, but pitch correctness deteriorates

Engineering Contradiction:
Improveplayback speedVSAvoidpitch correctness
Core Design Contradiction:
SpeedVSManufacturing precision

Solution Approach 1:

The audio signal is divided into multiple frames, with at least two different playback speeds applied to different frames. This segmentation allows the system to vary playback speed across different time segments while maintaining pitch correctness within each segment, thereby resolving the contradiction between improved playback speed and pitch correctness.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If fixed playback speed is used for all audio content, then device complexity is reduced, but audio intelligibility deteriorates

Engineering Contradiction:
Improveplayback control simplicityVSAvoidaudio intelligibility
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The system dynamically adjusts playback speed based on acoustic analysis of the audio signal. The playback speed is not fixed but varies according to the detected rate of audio utterance, allowing the system to maintain audio intelligibility while adapting to different speaking rates in the content.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs automatic rate control by analyzing the audio signal itself to determine the appropriate playback speed. The acoustic analysis unit detects the rate of audio utterance directly from the audio content, enabling the system to self-adjust without external intervention, thereby maintaining intelligibility without requiring complex user input.

Inventive Principle:
Principle #25Self-service

3Productivity

If playback speed is increased to reduce listening time, then productivity is improved, but audio intelligibility deteriorates

Engineering Contradiction:
Improvelistening efficiencyVSAvoidaudio intelligibility
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system dynamically adjusts playback speed based on the detected rate of audio utterance. When the original audio is spoken at a slower rate, the system applies faster playback speed to match user preferences and improve listening efficiency. When the original audio is already at a faster rate, the system maintains or reduces playback speed to preserve intelligibility, thereby optimizing productivity without sacrificing understanding.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10354676B2Automatic rate control for improved audio time scaling
Publication Date: 2019.07.16 ADEIA MEDIA SOLUTIONS INC
  • US10354676B2 patent drawing
  • US10354676B2 patent drawing
  • US10354676B2 patent drawing

AI summary

Input media data with an input playing speed is received and divided into input media data subsets. A first rate of audio utterance is determined for a first input media data subset in the media data subsets. A second different rate of audio utterance is determined for a second input media data subset in the media data subsets. Audio output media data is generated with an output playing speed at which audio utterance in the audio output media data is played at a preferred rate of audio utterance. The audio output media data comprises (a) a first output audio media data subset generated based on the preferred rate, the first rate, and the first input media data subset and (b) a second output audio media data subset generated based on the preferred rate, the second rate, and the second input media data subset.