MDCT Window Zero Pad Region for Audio Signal Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech processing technologies face challenges in minimizing the amount of information transmitted while maintaining perceived quality, particularly in real-time voice communication systems, where constraints on frame delay and look-ahead length limit the effectiveness of modified discrete cosine transform (MDCT) coding.
Innovation Solution
The method involves applying a modified discrete cosine transform (MDCT) window function to frames associated with non-speech signals to generate zero pad regions, allowing for efficient encoding and perfect reconstruction of audio signals, even with less than 50% frame overlap, by selecting an appropriate window function based on constraints such as frame length and delay.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If MDCT coding is used to compress audio signals, then bandwidth requirements are reduced, but perfect reconstruction becomes difficult with less than 50% frame overlap
Solution Approach 1:
The audio signal is divided into frames with overlapping regions. The patent segments the signal processing into distinct operations: windowing the overlapping region, performing MDCT on segmented portions, and selectively zeroing specific coefficients. This segmentation allows the system to maintain perfect reconstruction capability while using less than 50% frame overlap, resolving the contradiction between compression efficiency and reconstruction reliability.
Solution Approach 2:
The patent applies different processing treatments to different regions of the signal. Specifically, the overlapping region is windowed differently from non-overlapping regions, and certain MDCT coefficients in the overlapping region are zeroed while others are preserved. This local differentiation enables perfect reconstruction in the overlapping region while maintaining compression efficiency in non-overlapping regions.
2Loss of time
If frame overlap is reduced to minimize delay, then real-time processing is improved, but encoding efficiency and reconstruction quality deteriorate
Solution Approach 1:
The patent applies a window function to the overlapping region before performing MDCT encoding. This preliminary action on the overlapping region prepares the signal in advance, ensuring that when frames are processed with less than 50% overlap, the reconstruction quality is maintained. The windowing operation is performed ahead of time on the overlapping portion, allowing reduced delay without sacrificing reconstruction precision.
3Productivity
If less than 50% frame overlap is used, then processing speed is increased, but traditional MDCT cannot achieve perfect reconstruction
Solution Approach 1:
The patent dynamically adjusts the processing approach based on the frame overlap condition. When less than 50% overlap is used, the system activates specific adaptations: windowing the overlapping region and selectively zeroing MDCT coefficients. This dynamic adjustment allows the system to maintain perfect reconstruction capability while operating at higher processing speeds with reduced frame overlap.
Solution Approach 2:
The patent modifies specific parameters of the MDCT process to enable perfect reconstruction with less than 50% frame overlap. The key parameter changes include: applying a window function to the overlapping region, and selectively zeroing certain MDCT coefficients in the overlapping region while preserving others. These parameter modifications allow the system to achieve both high processing speed and perfect reconstruction reliability.
Data Source
AI summary
A method for modifying a window with a frame associated with an audio signal is described. A signal is received. The signal is partitioned into a plurality of frames. A determination is made if a frame within the plurality of frames is associated with a non-speech signal. A modified discrete cosine transform (MDCT) window function is applied to the frame to generate a first zero pad region, where the region has a length of (M−L)/2, where L is an arbitrary value, and a second zero pad region if it was determined that the frame is associated with a non-speech signal. The frame is encoded. The decoder window is the same as the encoder window.


