Audio Signal Time Stretching via Steady-Transient Feature Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal time stretching technologies fail to maintain natural sound quality when expanding or compressing transient sections with fluctuating acoustic characteristics, leading to unnatural sound impressions.
Innovation Solution
An audio processing method that extracts feature quantities from multiple periods of an audio signal, calculates similarity indices, and performs time axis expansion or compression based on these indices and transition costs to match similar sections, while excluding sections with dissimilar fluctuations, thereby generating a second audio signal that maintains auditory naturalness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If uniform time axis expansion/compression is applied to all audio signal sections, then processing simplicity is maintained, but sound quality deteriorates due to unnatural distortion of transient sections
Solution Approach 1:
The audio signal is divided into multiple periods, and each period is further segmented into steady sections and transient sections based on feature quantity analysis. This segmentation allows different processing strategies to be applied to different sections, resolving the contradiction between processing simplicity and sound quality by automating the differentiation and application of appropriate processing methods.
Solution Approach 2:
Different time axis expansion/compression processing is applied to different sections of the audio signal based on their local characteristics. Steady sections undergo uniform expansion/compression while transient sections are processed differently to preserve their natural sound quality. This local quality approach ensures that each section receives the most appropriate processing for its specific characteristics.
2Productivity
If transient sections are included in time axis expansion/compression, then complete signal processing is achieved, but unnatural sound impressions are created
Solution Approach 1:
Transient sections are identified and extracted from the audio signal through feature quantity analysis. By separating transient sections from steady sections, the invention can apply different processing strategies: steady sections are expanded/compressed while transient sections are excluded or processed differently, thereby maintaining sound naturalness while still achieving complete signal processing.
Solution Approach 2:
The processing approach dynamically adapts to the local characteristics of each audio section. The system automatically adjusts the processing method based on whether a section is steady or transient, using feature quantity analysis to guide the dynamic selection of processing parameters and methods, thus maintaining both completeness and naturalness.
3Manufacturing precision
If feature quantity analysis is performed for each period to distinguish steady and transient sections, then sound quality is improved, but processing complexity increases
Solution Approach 1:
The audio signal itself provides the information needed for processing through its feature quantities. By analyzing features such as pitch, energy, and spectral characteristics that are inherently present in the signal, the system can automatically distinguish steady from transient sections without requiring external intervention or complex manual analysis, thus improving sound quality while managing processing complexity.
Solution Approach 2:
The system monitors changes in acoustic parameters (pitch, energy, spectral features) across different periods to identify steady versus transient sections. By detecting parameter stability or fluctuation patterns, the system can automatically determine appropriate processing strategies, improving sound quality through parameter-based differentiation while keeping the processing framework relatively simple.
Data Source
AI summary
An audio processing device includes a feature extraction unit and signal generating unit. The feature extraction unit is configured to extract a feature quantity of a first audio signal for each of a plurality of periods. The signal generating unit is configured to for generate a second audio signal by time axis expanding/compressing either a section of the first audio signal in which the feature quantity is steadily maintained for a period time, or a section of the first audio signal in which a fluctuation of the feature quantity is repeated and excluding from the time axis expanding/compressing a section of the first audio signal in which a fluctuation of the feature quantity is not similar to that of other sections of the first audio signal.


