Variable Frame Audio Encoding via Attack Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional codecs face challenges in achieving both high compression efficiency and sound quality when encoding and decoding audio and speech signals, as speech codecs compromise sound quality for audio signals and vice versa.
Innovation Solution
A method and apparatus that dynamically adjust frame lengths based on the position and intensity of an attack in the input signal, transforming frames into frequency domains and determining encoding domains for each sub-frequency band to optimize time and frequency resolution, allowing for encoding and decoding in either the time or frequency domain as needed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a speech codec is used to encode audio signals, then compression efficiency is improved, but sound quality deteriorates
Solution Approach 1:
The patent applies dynamics by making the frame length variable rather than fixed. The frame length is adjusted dynamically based on the signal characteristics, specifically the attack position detected in the signal. This allows the encoding system to adapt between time-domain and frequency-domain encoding approaches, thereby resolving the contradiction between compression efficiency and sound quality for different types of signals.
Solution Approach 2:
The patent applies local quality by differentiating the encoding approach for different portions of the signal. Specifically, the signal is divided into frames, and each frame is encoded differently based on its local characteristics (attack position). Some frames are encoded in the time domain while others are encoded in the frequency domain, allowing optimal quality and compression for each local segment.
2Manufacturing precision
If a speech codec is used to encode speech signals, then sound quality is improved, but compression efficiency deteriorates
Solution Approach 1:
The system dynamically selects between time-domain and frequency-domain encoding based on the detected attack position in the signal. For speech signals, the variable frame length allows the system to use time-domain encoding when appropriate (maintaining sound quality) while still achieving good compression efficiency through adaptive framing and frequency-domain encoding when beneficial.
Solution Approach 2:
The patent changes the parameter of frame length from fixed to variable based on signal characteristics. By detecting the attack position and adjusting the frame length accordingly, the system can optimize the balance between sound quality and compression efficiency for speech signals, using different encoding parameters for different signal segments.
3Ease of operation
If fixed frame length is used for encoding, then processing simplicity is improved, but encoding efficiency deteriorates
Solution Approach 1:
The patent implements variable frame length by detecting the attack position in the signal and adjusting the frame length accordingly. This dynamic approach improves encoding efficiency by adapting to signal characteristics while maintaining reasonable processing complexity through systematic detection and adjustment mechanisms.
Solution Approach 2:
The system performs preliminary detection of the attack position before encoding. By detecting the attack position in advance and determining the appropriate frame length before the actual encoding process, the system prepares the optimal encoding parameters beforehand, improving overall encoding efficiency without significantly increasing processing complexity during the main encoding operation.
Data Source
AI summary
Provided is a method of encoding an audio/speech signal, the method including determining a variable length of a frame, that is, a processing unit of an input signal in accordance with a position of an attack in the input signal; transforming each frame of the input signal to a frequency domain and dividing the frame into a plurality of sub frequency bands; and, if a signal of a sub frequency band is determined to be encoded in the frequency domain, encoding the signal of the sub frequency band in the frequency domain, and if the signal of the sub frequency band is determined to be encoded in a time domain, inverse transforming the signal of the sub frequency band to the time domain and encoding the inverse transformed signal in the time domain. According to the present invention, the audio/speech signal may be efficiently encoded by controlling time resolution and frequency resolution.


