Method and audio processing unit for high frequency reconstruction of an audio signal
The method improves spectral band replication by regenerating high-frequency audio components based on metadata, addressing the limitations of existing techniques for music content, enhancing audio quality and compression efficiency.
Patent Information
- Application Number
- JP2025167842
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-01-26
- Filing Date
- 2025-10-06
- Publication Date
- 2026-01-14
AI Technical Summary
Spectral band replication techniques in audio coding, such as SBR, are not ideal for certain audio types like music content with low crossover frequencies, necessitating improvements in spectral band replication processes.
The method involves decoding an encoded audio bitstream, filtering the low-band audio signal, and regenerating the high-band portion using high-frequency reconstruction metadata, with options for spectral transformation or harmonic transposition based on flags, and combining the signals to form a wideband audio signal.
Enhances audio quality by accurately reconstructing high-frequency components, improving compression efficiency and adaptability to different audio types, while maintaining spectral characteristics.
Smart Images

Figure 2026004507000001_ABST
Abstract
Description
[Background technology]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to the following application, which is incorporated herein by reference: U.S. Provisional Application No. 62 / 622,205, filed January 26, 2018.
[0002] Technical Field Embodiments relate to audio signal processing, and more particularly to encoding, decoding, or transcoding of audio bitstreams according to control data specifying that either a basic form of high frequency reconstruction (HFR) or an enhanced form of HFR should be performed on the audio data.
[0003] Background of the Invention A typical audio bitstream contains both audio data (e.g., coded audio data) that describes one or more channels of audio content, and metadata that describes at least one characteristic of the audio data or audio content. One well-known format for generating coded audio bitstreams is the MPEG-4 Advanced Audio Coding (AAC) format, which is described in the MPEG standard ISO / IEC 14496-3:2009. In the MPEG4 standard, AAC stands for "Advanced Audio Coding," and HE-AAC stands for "High Efficiency Advanced Audio Coding."
[0004] The MPEG-4 AAC standard defines several audio profiles that determine the objects and coding tools present in a compliant encoder or decoder. Three of these audio profiles are (1) the AAC profile, (2) the HE-AAC profile, and (3) the HE-AAC v2 profile. The AAC profile includes the AAC Low Complexity (or "AAC-LC") object type. The AAC-LC object type corresponds to the MPEG-2 AAC Low Complexity profile with some adjustments and does not include the Spectral Band Replication ("SBR") or Parametric Stereo ("PS") object types. The HE-AAC profile is a superset of the AAC profile and additionally includes the SBR object type. The HE-AACv2 profile is a superset of the HE-AAC profile and additionally includes the PS object type.
[0005] The SBR object type includes a spectral band replication tool, an important high-frequency reconstruction ("HFR") coding tool that significantly improves the compression efficiency of perceptual audio codecs. SBR reconstructs the high-frequency components of an audio signal at the receiver (e.g., at the decoder). Thus, the encoder only needs to encode and transmit the low-frequency components, enabling very high audio quality at low data rates. SBR is based on replicating a pre-truncated sequence of harmonics from control data obtained from the encoder and available bandwidth-limited signals to reduce the data rate. The ratio between tonal and noise components is maintained by adaptive inverse filtering in addition to the selective addition of noise and sinusoids. In the MPEG-4 AAC standard, the SBR tool performs a spectral patching process (also known as a linear or spectral transform), in which a number of successive quadrature mirror filter (QMF) subbands are copied (or "patched") from the transmitted low-band portion of the audio signal to the high-band portion of the audio signal generated at the decoder.
[0006] Spectral patching or linear transformation may not be ideal for certain audio types, such as music content, with relatively low crossover frequencies. Therefore, techniques to improve spectral band replication are needed. Summary of the Invention
[0007] According to a first class of embodiments, a method for decoding an encoded audio bitstream is disclosed. The method includes receiving the encoded audio bitstream and decoding the audio data to generate a decoded low-band audio signal. The method further includes extracting high-frequency reconstruction metadata and filtering the decoded low-band audio signal with an analysis filterbank to generate a filtered low-band audio signal. The method further includes extracting a flag indicating whether a spectral transformation or a harmonic transposition should be performed on the audio data, and regenerating a high-band portion of the audio signal using the high-frequency reconstruction metadata and the filtered low-band audio signal according to the flag. Finally, the method includes combining the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal.
[0008] A second class of embodiments relates to an audio decoder for decoding an encoded audio bitstream. The decoder includes an input interface for receiving the encoded audio bitstream, the encoded audio bitstream including audio data representing a low-band portion of an audio signal, and a core decoder for decoding the audio data to generate a decoded low-band audio signal. The decoder also includes a demultiplexer for extracting from the encoded audio bitstream high-frequency reconstruction metadata, the high-frequency reconstruction metadata including operating parameters for a high-frequency reconstruction process that linearly transforms a successive number of subbands from the low-band portion of the audio signal to a high-band portion of the audio signal, and an analysis filterbank for filtering the decoded low-band audio signal to generate a filtered low-band audio signal. The decoder further includes the demultiplexer for extracting from the encoded audio bitstream a flag indicating whether a linear transformation or a harmonic transposition should be performed on the audio data, and a high-frequency regenerator for regenerating the high-band portion of the audio signal using the high-frequency reconstruction metadata and the filtered low-band audio signal according to the flag. Finally, the decoder includes a synthesis filter bank for combining the filtered low-band audio signal with the regenerated high-band portion to form a wideband audio signal.
[0009] Another class of embodiments relates to encoding and transcoding audio bitstreams that include metadata that identifies whether enhanced spectral band replication (eSBR) processing should be performed. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a block diagram of an embodiment of a system that may be configured to implement embodiments of the method of the present invention. [Figure 2] FIG. 2 is a block diagram of an encoder that is an embodiment of the audio processing unit of the present invention. [Figure 3] 1 is a block diagram of a system including a decoder, an embodiment of an audio processing unit of the present invention, and optionally a post-processor coupled thereto. [Figure 4] FIG. 2 is a block diagram of a decoder that is an embodiment of the audio processing unit of the present invention. [Figure 5] FIG. 2 is a block diagram of a decoder, which is another embodiment of the audio processing unit of the present invention. [Figure 6] FIG. 2 is a block diagram of another embodiment of the audio processing unit of the present invention. [Figure 7] 1 shows a block diagram of an MPEG-4 AAC bitstream containing divided segments. DETAILED DESCRIPTION OF THE INVENTION
[0011] Notation and Terminology Throughout this disclosure, including the claims, the expression performing processing "on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used broadly to refer to performing processing directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has been subjected to preliminary filtering or pre-processing before performing that processing).
[0012] Throughout this disclosure, including the claims, the terms "audio processing unit" or "audio processor" are used broadly to refer to a system, device, or apparatus configured to process audio data. Examples of audio processing units include, but are not limited to, encoders, transcoders, decoders, codecs, pre-processing systems, post-processing systems, and bitstream processing systems (often referred to as bitstream processing tools). Virtually all consumer electronic products, such as mobile phones, televisions, laptops, and tablet computers, incorporate an audio processing unit or audio processor.
[0013] Throughout this disclosure, including the claims, the terms "couple" or "coupled" are used broadly to mean a direct or indirect connection. Thus, if a first device couples to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections. Additionally, components that are integrated within or with other components are also coupled to each other.
[0014] Detailed Description of Embodiments of the Invention The MPEG-4 AAC standard contemplates that an encoded MPEG-4 AAC bitstream includes metadata that indicates each type of High Frequency Reconstruction (HFR) processing to be applied (if any) by a decoder to decode the audio content of the bitstream, and / or that controls such HFR processing and / or that indicates at least one characteristic or parameter of at least one HFR tool to be used to decode the audio content of the bitstream. Herein, we use the expression "SBR metadata" to indicate this type of metadata described or referenced in the MPEG-4 AAC standard for use with Spectral Band Replication ("SBR"). As will be understood by those skilled in the art, SBR is a form of HFR.
[0015] SBR is preferably used as a dual-rate system, where the underlying codec operates at half the original sampling rate, while SBR operates at the original sampling rate. The SBR encoder runs in parallel with the underlying core codec, albeit at a higher sampling rate. While SBR is primarily a post-processing step in the decoder, key parameters are extracted in the encoder to ensure the highest accuracy of high-frequency reconstruction in the decoder. The encoder estimates the spectral envelope of the SBR range for a time and frequency range / resolution appropriate to the characteristics of the current input signal segment. The spectral envelope is estimated through a complex QMF analysis followed by an energy calculation. The time and frequency resolution of the spectral envelope can be chosen with a high degree of freedom to ensure the optimal time-frequency resolution for a given input segment. The envelope estimation must take into account that transients located in the original region, mainly in the high-frequency region (e.g., a high-hat), are slightly present in the SBR high band generated before envelope adjustment, because the high band in the decoder is based on the low band, and the transients are deemed to be much smaller than in the high band. This aspect imposes different requirements on the time-frequency resolution of the spectral envelope data compared to regular spectral envelope estimation as used in other audio coding algorithms.
[0016] Apart from the spectral envelope, several additional parameters are extracted that describe the spectral characteristics of the input signal for different time and frequency domains. Since the encoder naturally has access not only to the original signal but also to information about how the SBR unit in the decoder generates the high band, the system can handle situations where the low band comprises a strong harmonic sequence and the regenerated high band comprises primarily random signal components, as well as situations where there are strong tonal components in the original high band that have no counterpart in the low band on which the high band region is based. Furthermore, the SBR encoder works closely with the underlying core codec to evaluate which frequency range should be covered by SBR at a given time. In the case of stereo signals, SBR data is efficiently encoded before transmission by exploiting the channel dependence of the control data as well as entropy coding.
[0017] Control parameter extraction algorithms typically need to be carefully tuned to the underlying codec at a given bit rate and a given sampling rate due to the fact that lower bit rates usually exhibit a larger SBR range compared to higher bit rates, and different sampling rates correspond to different temporal resolutions of the SBR frames.
[0018] An SBR decoder typically includes several distinct parts, including a bitstream decoding module, a high-frequency reconstruction module (HFR), an additional high-frequency component module, and an envelope adjustment module. The system is based on a complex-valued QMF filter bank (for high-quality SBR) or a real-valued QMF filter bank (for low-power SBR). Embodiments of the present invention are applicable to both high-quality and low-power SBR. In the bitstream extraction module, control data is read from the bitstream and decoded. A time-frequency grid is obtained for the current frame before reading the envelope data from the bitstream. The underlying core decoder decodes the audio signal for the current frame (albeit at a lower sampling rate) to generate time-domain audio samples. The resulting frame of audio data is used for high-frequency reconstruction by the HFR module. The decoded low-band signal is then analyzed using a QMF filter bank. High-frequency reconstruction and envelope adjustment are then performed on the subband samples of the QMF filter bank. The high-frequency band is reconstructed from the low-band signal in a flexible manner based on given control parameters. Furthermore, the reconstructed high band is adaptively filtered on a subband-channel basis according to control data to ensure proper spectral characteristics in a given time / frequency domain.
[0019] The top level of an MPEG-4 AAC bitstream is a sequence of data blocks ('raw_data_block' elements), each of which is a segment of data (hereafter referred to as a 'block') containing audio data (typically spanning a duration of 1024 or 960 samples) and related information and / or other data. Herein, we use the term 'block' to indicate a segment of an MPEG-4 AAC bitstream containing audio data (and corresponding metadata and optionally other associated data) determined or indicated by one (but not more than one) 'raw_data_block' element.
[0020] Each block of an MPEG-4 AAC bitstream can contain multiple syntax elements (each of which also appears in the bitstream as a segment of data). Seven types of such syntax elements are defined in the MPEG-4 AAC standard. Each syntax element is identified by a different value of the data element "id_syn_element". Examples of syntax elements include "single_channel_element()", "channel_pair_element()", and "fill_element()". A single channel element is a container that contains audio data for a single audio channel (monaural audio signal). A channel pair element contains audio data for two audio channels (stereo audio signal).
[0021] A fill element is an information container that contains an identifier (e.g., the value of the id_syn_ele element above) followed by data, referred to as "fill data". Fill elements have historically been used to adjust the instantaneous bit rate of a bitstream transmitted over a constant rate channel. By adding the appropriate amount of fill data to each block, a constant data rate can be achieved.
[0022] According to embodiments of the present invention, fill data may include one or more extension payloads that extend the types of data (e.g., metadata) that can be transmitted in a bitstream. A decoder that receives a bitstream with fill data containing new types of data may optionally be used by the device (e.g., decoder) receiving the bitstream to extend the capabilities of the device. Thus, as will be appreciated by those skilled in the art, a fill element is a special type of data structure that is distinct from data structures typically used to transmit audio data (e.g., audio payloads containing channel data).
[0023] In some embodiments of the present invention, the identifier used to identify a fill element may consist of a three bit unsigned integer transmitted most significant bit first (uimsbf) with a value of 0x6. In one block, several instances of the same type of syntax element (e.g., multiple fill elements) may occur.
[0024] Another standard for encoding audio bitstreams is the MPEG (Unified Speech and Audio Coding) standard (ISO / IEC 23003-3:2012). The MPEG USAC standard describes the encoding and decoding of audio content using a spectral band duplication process (including the SBR process described in the MPEG-4 AAC standard, as well as other enhanced forms of the spectral band duplication process). This process applies spectral band duplication tools (often referred to herein as "extended SBR tools" or "eSBR tools") that are extended and enhanced versions of the set of SBR tools described in the MPEG-4 AAC standard. Thus, eSBR (as defined in the USAC standard) is an improvement over SBR (as defined in the MPEG-4 AAC standard).
[0025] The term "enhanced SBR processing" (or "eSBR processing") is used herein to refer to a spectral band duplication process that uses at least one eSBR tool not described or referenced in the MPEG-4 AAC standard (e.g., at least one eSBR tool described or referenced in the MPEG USAC standard). Examples of such eSBR tools are harmonic transposition and QMF patch processing, additional preprocessing or "pre-flattening."
[0026] An integer-order T harmonic transposer maps a sinusoid at frequency ω to a sinusoid at frequency Tω while preserving signal duration. Three orders, T=2, 3, and 4, are typically used in sequence to generate each portion of the desired output frequency range using the smallest possible transposition order. If an output above the fourth-order transposition range is required, it may be generated by frequency shifting. When possible, a near-critically sampled baseband time domain is created for processing to minimize computational complexity.
[0027] Harmonic transposers can be either QMF- or DFT-based. When using a QMF-based harmonic transposer, bandwidth expansion of the core coder time-domain signal is performed entirely in the QMF domain using a modified phase vocoder structure, performing time expansion after decimation for all QMF subbands. Transpositions using several transposition factors (e.g., T = 2, 3, 4) are performed in a common QMF analysis / synthesis transform stage. Since QMF-based harmonic transposers do not feature signal-adaptive frequency-domain oversampling, the corresponding flag in the bitstream (sbrOversamplingFlag[ch]) may be ignored.
[0028] When using DFT-based harmonic transposers, factor 3 and 4 transposers (cubic and quartic transposers) are preferably embedded in factor 2 transposers (quadratic transposers) by interpolation to reduce complexity. For each frame (corresponding to coreCoderFrameLength core coder samples), the nominal "full size" transform size of the transposer is initially determined by the signal-adaptive frequency-domain oversampling flag (sbrOverSamplingFlag[ch]) in the bitstream.
[0029] If sbrPatchingMode==1, it indicates that a linear transposition should be used to generate the high band, and an additional step may be introduced to avoid discontinuities in the shape of the spectral envelope of the high-frequency signal input to the subsequent envelope adjuster. This improves the operation of the next envelope adjustment stage, resulting in a high-band signal that is perceived as more stable. The additional preprocessing operation is beneficial for signal types where the coarse spectral envelope of the low-band signal used for high-frequency reconstruction exhibits large fluctuations in level. However, the value of the bitstream element may be determined within the encoder by applying any kind of signal-dependent classification. The additional preprocessing is preferably activated by the 1-bit bitstream element bs_sbr_preprocessing. If bs_sbr_preprocessing is set to 1, the additional processing is enabled. If bs_sbr_preprocessing is set to zero, the additional processing is disabled. The additional processing is performed on the low-band X for each patch. Low It is preferable to utilize a pre-gain curve used by the high frequency generator to scale . For example, the pre-gain curve may be calculated according to:
number
number
number
number
[0030] Bitstreams produced in accordance with the MPEG USAC standard (sometimes referred to herein as "USAC bitstreams") contain encoded audio content and typically include metadata that indicates each type of spectral band replication process that is applied by a decoder to decode the audio content of the USAC bitstream, and / or that indicates at least one characteristic or parameter of at least one SBR tool and / or eSBR tool used to control such spectral band replication process and / or decode the audio content of the USAC bitstream.
[0031] In this application, the term "enhanced SBR metadata" (or eSBR metadata) is used to refer to metadata that indicates each type of spectral band duplication process applied by a decoder to decode audio content of a coded audio bitstream (e.g., a USAC bitstream) and / or that controls such spectral band duplication process and / or that indicates at least one characteristic or parameter of at least one SBR tool and / or eSBR tool used to decode such audio content, but that is not described or mentioned in the MPEG-4 AAC standard. Specific examples of eSBR metadata include metadata (for indicating or controlling spectral band duplication processes) that is described or mentioned in the MPEG-4 USAC standard but not in the MPEG-4 AAC standard. Therefore, eSBR metadata in this application refers to metadata that is not SBR metadata, and SBR metadata in this application refers to metadata that is not eSBR metadata.
[0032] A USAC bitstream may contain both SBR and eSBR metadata. More specifically, a USAC bitstream may contain eSBR metadata that controls the performance of eSBR processing by a decoder, and SBR metadata that controls the performance of SBR processing by a decoder. According to an exemplary embodiment of the present invention, eSBR metadata (e.g., eSBR-specific configuration data) is included in an MPEG-4 AAC bitstream (e.g., in an sbr_extension() container at the end of the SBR payload) (according to the present invention).
[0033] During decoding of a bitstream encoded using an eSBR toolset (including at least one eSBR tool), the decoder performs eSBR processing to regenerate the high frequency band of the audio signal based on replicating the sequence of harmonics that were truncated during encoding. Such eSBR processing typically involves adjusting the spectral envelope of the generated high frequency band, applying inverse filtering, and adding noise and sinusoidal components to recreate the spectral characteristics of the original audio signal.
[0034] According to an exemplary embodiment of the present invention, eSBR metadata is included in one or more metadata segments (e.g., a small number of control bits that are eSBR metadata) of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that contains audio data encoded in other segments (audio data segments). Typically, at least one such metadata segment in each block of the bitstream is (or includes) a fill element (including an identifier indicating the start of the fill element), and the eSBR metadata is included in the fill element following the identifier.
[0035] 1 is a block diagram of an exemplary audio processing chain (audio data processing system), one or more elements of which may be configured in accordance with embodiments of the present invention. The system includes elements coupled together as shown: an encoder 1, a distribution subsystem 2, a decoder 3, and a post-processing unit 4. Variations on the illustrated system may omit one or more elements or include additional audio data processing units.
[0036] In some implementations, the encoder 1 (optionally including a preprocessing unit) is configured to accept as input PCM (time-domain) samples containing audio content and to output an encoded audio bitstream (having a format conforming to the MPEG-4 AAC standard) representing the audio content. The data in the bitstream representing the audio content is sometimes referred to herein as "audio data" or "encoded audio data." When the encoder is configured according to exemplary embodiments of the present invention, the audio bitstream output from the encoder includes eSBR metadata (and typically other metadata) as well as the audio data.
[0037] One or more encoded audio bitstreams output from Encoder 1 may be asserted to Encoded Audio Distribution Subsystem 2. Subsystem 2 is configured to store and / or distribute each encoded bitstream output from Encoder 1. The encoded audio bitstreams output from Encoder 1 may be stored by Subsystem 2 (e.g., in the form of a DVD or Blu-ray disc), transmitted by Subsystem 2 (which may implement a transmission link or network), or both stored and transmitted by Subsystem 2.
[0038] Decoder 3 is configured to decode the encoded MPEG-4 AAC audio bitstream (generated by Encoder 1) received via Subsystem 2. In some embodiments, decoder 3 is configured to extract eSBR metadata from each block of the bitstream, decode the bitstream (including performing eSBR processing using the extracted eSBR metadata), and generate decoded audio data (e.g., a stream of decoded PCM audio samples). In some embodiments, decoder 3 is configured to extract SBR metadata from the bitstream (but ignore the eSBR metadata included in the bitstream), decode the bitstream (including performing SBR processing using the extracted SBR metadata), and generate decoded audio data (e.g., a stream of decoded PCM audio samples). Typically, decoder 3 includes a buffer that stores (e.g., in a non-temporary manner) segments of the encoded audio bitstream received from Subsystem 2.
[0039] 1 is configured to accept a stream of decoded audio data (e.g., decoded PCM audio samples) from decoder 3 and perform post-processing thereon. The post-processing unit may also be configured to render the post-processed audio content (or decoded speech received from decoder 3) for playback over one or more speakers.
[0040] FIG. 2 is a block diagram of an encoder (100), an embodiment of an audio processing unit of the present invention. Any of the components or elements of encoder 100 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware and software. Encoder 100 includes encoder 105, stuffer / formatter stage 107, metadata generation stage 106, and buffer memory 109, connected as shown. Typically, encoder 100 also includes other processing elements (not shown). Encoder 100 is configured to convert an input audio bitstream into an encoded output MPEG-4 AAC bitstream.
[0041] Metadata generator 106 is configured and coupled to generate (and / or pass through) stage 107 metadata (including eSBR metadata and SBR metadata) to be included by stage 107 in the encoded bitstream to be output from encoder 100.
[0042] Encoder 105 is configured and coupled to encode input audio data (e.g., by performing compression thereon) and assert the resulting encoded audio to stage 107 for inclusion in an encoded bitstream to be output from stage 107.
[0043] Stage 107 is configured to multiplex the encoded audio from encoder 105 and the metadata (including eSBR metadata and SBR metadata) from generator 106 to generate an encoded bitstream that is output from stage 107, preferably configured so that the encoded bitstream has a format as specified by any of the embodiments of the present invention.
[0044] The buffer memory 109 is configured to store (e.g., in a non-transitory manner) at least one block of the encoded audio bitstream output from the stage 107, and then a sequence of blocks of the encoded audio bitstream is asserted from the buffer memory 109 as output from the encoder 100 to a distribution system.
[0045] 3 is a block diagram of a system including a decoder (200) and an optional post-processor (300) coupled thereto, which are embodiments of an audio processing unit of the present invention. Any of the components or elements of the decoder 200 and post-processor 300 may be implemented as one or more processes and / or one or more circuits (e.g., an ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware and software. The decoder 200 includes a buffer memory 201, a bitstream payload deformatter 205, an audio decoding subsystem 202 (sometimes referred to as the "core" decoding stage or "core" decoding subsystem), an eSBR processing stage 203, and a control bit generation stage 204, connected as shown. Typically, the decoder 200 also includes other processing elements (not shown).
[0046] Buffer memory (buffer) 201 stores (e.g., in a non-temporal format) at least one block of an encoded MPEG-4 AAC audio bitstream that is received by decoder 200. In operation of decoder 200, a sequence of blocks of the bitstream is asserted from buffer 201 to deformatter 205.
[0047] In a variation of the embodiment of Figure 3 (or the embodiment of Figure 4 described below), an APU that is not a decoder (e.g., APU 500 of Figure 6) includes a buffer memory (e.g., a buffer memory identical to buffer 201) that stores (e.g., in a non-temporary manner) at least one block of the same type of coded audio bitstream (e.g., an MPEG-4 AAC audio bitstream) received by buffer 201 of Figure 3 or Figure 4 (i.e., a coded audio bitstream that includes eSBR metadata).
[0048] 3, deformatter 205 is configured and coupled to demultiplex (or separate) each block of the bitstream, extract SBR and eSBR metadata (including quantized envelope data) (and typically other metadata) therefrom, assert at least the eSBR and SBR metadata to eSBR processing stage 203, and typically also assert other extracted metadata to decoding subsystem 202 (and optionally to control bit generator 204). Deformatter 205 is also configured and coupled to extract audio data from each block of the bitstream and assert the extracted audio data to decoding subsystem (decoding stage) 202.
[0049] 3 also includes an optional post-processor 300. The post-processor 300 includes a buffer memory (buffer) 301 and other processing elements (not shown), including at least one processing element coupled to the buffer 301. The buffer 301 stores (in a non-temporary manner) at least one block (or frame) of decoded audio data received by the post-processor 300 from the decoder 200. The processing elements of the post-processor 300 are configured and coupled to receive and adaptively process sequences of blocks of decoded audio output from the buffer 301 using metadata output from the decoding subsystem 202 (and / or the deformatter 205) and / or control bits output from the stage 204 of the decoder 200.
[0050] The audio decoding subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the parser 205 (such decoding may be referred to as the "core" decoding process) to generate decoded audio data and assert the decoded audio data to the eSBR processing stage 203. The decoding is performed in the frequency domain and typically includes inverse quantization followed by spectral processing. Typically, the final stage of processing in the subsystem 202 applies a frequency-domain to time-domain transformation to the decoded frequency-domain audio data, such that the output of the subsystem is time-domain decoded audio data. The stage 203 is configured to apply the SBR and eSBR tools indicated by the eSBR metadata and eSBR (extracted by the parser 205) to the decoded audio data (i.e., perform SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata) to generate fully decoded audio data that is output from the decoder 200 (e.g., to the post-processor 300). Typically, decoder 200 includes a memory (accessible by subsystem 202 and stage 203) that stores the metadata and deformatted audio data from deformatter 205, with stage 203 configured to access the audio data and metadata (including SBR and eSBR metadata) as needed during SBR and eSBR processing. The SBR and eSBR processing in stage 203 may be considered post-processing on the output of core decoding subsystem 202.Optionally, decoder 200 also includes a final upmixing subsystem (which may apply parametric stereo (PS) tools as defined in the MPEG-4 AAC standard using the PS metadata extracted by deformatter 205 and / or the control bits generated in subsystem 204) configured and coupled to perform an upmix on the output of stage 203 to generate fully decoded upmixed audio that is output from decoder 200. Alternatively, post-processor 300 is configured to perform upmixing on the output of decoder 200 (e.g., using the PS metadata extracted by deformatter 205 and / or the control bits generated in subsystem 204).
[0051] In response to the metadata extracted by the deformatter 205, the control bit generator 204 can generate control data that can be used within the decoder 200 (e.g., in a final upmixing subsystem) and / or asserted as an output of the decoder 200 (e.g., to a post-processor 300 for use in post-processing). In response to the metadata extracted from the input bitstream (and optionally in response to the control data), stage 204 can generate (and assert to the post-processor 300) a control bit that indicates that the decoded audio data output from the eSBR processing stage 203 should undergo a particular type of post-processing. In some implementations, the decoder 200 is configured to assert the metadata extracted by the deformatter 205 from the input bitstream to the post-processor 300, and the post-processor 300 is configured to use the metadata to perform post-processing on the decoded audio data output from the decoder 200.
[0052] FIG. 4 is a block diagram of an audio processing unit ("APU") 210, another embodiment of an audio processing unit of the present invention. The APU 210 is a legacy decoder not configured to perform eSBR processing. Any of the components or elements of the APU 210 may be implemented as one or more processes and / or one or more circuits (e.g., an ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware. The APU 210 includes a buffer memory 201, a bitstream payload deformatter (parser) 215, an audio decoding subsystem 202 (often referred to as the "core" decoding stage or "core" decoding subsystem), and an SBR processing stage 213, connected as shown. Typically, the APU 210 includes other processing elements (not shown). The APU 210 could represent, for example, an audio encoder, decoder, or transcoder.
[0053] Elements 201 and 202 of APU 210 are identical to similarly numbered elements of decoder 200 (of FIG. 3), and their above description will not be repeated. In operation of APU 210, a sequence of blocks of an encoded audio bitstream (MPEG-4 AAC bitstream) received by APU 210 is asserted from buffer 201 to deformatter 215.
[0054] The deformatter 215 is configured and coupled to demultiplex each block of the bitstream and extract SBR metadata (including quantization envelope data) and typically other metadata therefrom, while ignoring eSBR metadata that may be included in the bitstream in accordance with some embodiments of the present invention. The deformatter 215 is configured to assert at least the SBR metadata to the SBR processing stage 213. The deformatter 215 is also configured and coupled to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.
[0055] The audio decoding subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the deformatter 215 to generate decoded audio data (such decoding may be referred to as the "core" decoding process) and assert the decoded audio data to the SBR processing stage 213. The decoding is performed in the frequency domain. Typically, the final stage of processing in the subsystem 202 applies a frequency-domain to time-domain transformation to the decoded frequency-domain audio data, so that the output of the subsystem is time-domain decoded audio data. The stage 213 is configured to apply SBR tools (but not eSBR tools) indicated by the SBR metadata (extracted by the deformatter 215) to the decoded audio data (i.e., perform SBR processing on the output of the decoding subsystem 202 using the SBR metadata) to generate fully decoded audio data that is output from the APU 210 (e.g., to the post-processor 300). Typically, APU 210 includes memory (accessible by subsystem 202 and stage 213) for storing deformatted audio data and metadata output from deformatter 215, with stage 213 configured to access the audio data and metadata (including SBR metadata) as needed during SBR processing. The SBR processing in stage 213 may be considered post-processing on the output of core decoding subsystem 202. Optionally, APU 210 also includes a coupled final upmixing subsystem (which may apply parametric stereo (PS) tools defined in the MPEG-4 AAC standard using the PS metadata extracted by deformatter 215) configured to perform upmixing on the output of stage 213 to generate fully decoded upmixed audio output from APU 210.Alternatively, the post-processor may be configured to perform upmixing on the output of the APU 210 (e.g., using PS metadata extracted by the deformatter 215 and / or control bits generated in the APU 210).
[0056] Various implementations of the encoder 100, decoder 200, and APU 210 are configured to perform various embodiments of the method of the present invention.
[0057] According to some embodiments, eSBR metadata is included in an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) (e.g., a small number of control bits are included that are eSBR metadata), so that legacy decoders (not configured to parse the eSBR metadata or to use any eSBR tools related to the eSBR metadata) can ignore the eSBR metadata but can still decode the bitstream to the extent possible without utilizing any eSBR tools or eSBR metadata related to the eSBR metadata, typically without any significant penalty in decoded audio quality. However, eSBR decoders configured to parse the bitstream to identify eSBR metadata and use at least one eSBR tool in response to the eSBR metadata will enjoy the benefits of using at least one such eSBR tool. Thus, embodiments of the present invention provide a means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner.
[0058] Typically, the eSBR metadata in the bitstream indicates (e.g., indicates at least one characteristic or parameter of) one or more of the following eSBR tools (which are described in the MPEG USAC standard and may or may not be applied by the encoder during generation of the bitstream): Harmonic transposition; and QMF-patching additional pre-processing (pre-flattening).
[0059] For example, eSBR metadata included in the bitstream may indicate the following parameter values (as described in the MPEG USAC standard and this disclosure): sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch], and bs_sbr_preprocessing.
[0060] Here, if X is some parameter, the notation X[ch] indicates that the parameter relates to a channel ("ch") of the audio content of the coded bitstream to be decoded. For simplicity, we often omit the notation [ch] and assume that the relevant parameter relates to a channel of the audio content.
[0061] Here, the notation X[ch][env], where X is some parameter, indicates that the parameter relates to the SBR envelope ("env") of a channel ("ch") of the audio content of the coded bitstream to be decoded. For simplicity, we often omit the notations [env] and [ch] and assume that the associated parameter relates to the SBR envelope of a channel of the audio content.
[0062] During decoding of the encoded bitstream, the performance of harmonic transposition during the eSBR processing stage of decoding (for each channel "ch" of the audio content represented by the bitstream) is controlled by the following eSBR metadata parameters: sbrPatchingMode[ch]; sbrOversamplingFlag[ch]; sbrPitchInBinsFlag[ch]; and sbrPitchInBins[ch].
[0063] The value "sbrPatchingMode[ch]" indicates the type of transposer used in eSBR: sbrPatchingMode[ch]=1 indicates linear transposition patching as described in section 4.6.18 of the MPEG-4 AAC standard (as used in high-quality SBR or low-power SBR); sbrPatchingMode[ch]=0 indicates harmonic SBR patching as described in sections 7.5.3 or 7.5.4 of the MPEG-4 AAC standard.
[0064] The value "sbrOversamplingFlag[ch]" indicates the use of signal-adaptive frequency-domain oversampling in eSBR combined with DFT-based harmonic SBR patching, as described in Section 7.5.3 of the MPEG USAC Standard. This flag controls the size of the DFT used in the transposer: 1 indicates that signal-adaptive frequency-domain oversampling is enabled, as described in Section 7.5.3.1 of the MPEG USAC Standard; 0 indicates that signal-adaptive frequency-domain oversampling is disabled, as described in Section 7.5.3.1 of the MPEG USAC Standard.
[0065] The value "sbrPitchInBinsFlag[ch]" controls the interpretation of the sbrPitchInBins[ch] parameter: 1 indicates that the value of sbrPitchInBins[ch] is valid and greater than zero; 0 indicates that the value of sbrPitchInBins[ch] is set to zero.
[0066] The value "sbrPitchInBins[ch]" controls the addition of cross product terms in the SBR harmonic transposer. The value sbrPitchinBins[ch] is an integer value in the range [0,127] and represents the distance measured in frequency bins of a 1536-line DFT acting on the sampling frequency of the core coder.
[0067] If an MPEG-4 AAC bitstream indicates an SBR channel pair, but those channels are not combined (not a single SBR channel), the bitstream shall indicate two instances of the above syntax (for harmonic or non-harmonic transposition), one for each channel in sbr_channel_pair_element().
[0068] Harmonic transposition in eSBR tools typically improves the quality of decoded music signals at relatively low crossover frequencies. Non-harmonic transposition (i.e., legacy spectral patching) typically improves speech signals. Therefore, the starting point in determining which type of transposition is preferred for encoding specific audio content is to select a transposition method dependent on speech / music detection, with harmonic transposition being used for music content and spectral patching being used for speech content.
[0069] The performance of pre-flattening during eSBR processing is controlled by the value of a one-bit eSBR metadata parameter known as "bs_sbr_preprocessing," with pre-flattening either occurring or not occurring depending on the value of this single bit. When an SBR QMF patch processing algorithm such as that described in section 4.6.18.6.3 of the MPEG-4 AAC standard is used, a pre-flattening step (if indicated by the "bs_sbr_preprocessing" parameter) may be performed to avoid discontinuities in the shape of the spectral envelope of the high-frequency signal input to the subsequent envelope adjuster (which performs another stage of eSBR processing). Pre-flattening typically improves the operation of the subsequent envelope adjustment stage, resulting in a high-band signal that is perceived as more stable.
[0070] The overall bitrate requirement for including eSBR metadata indicating the eSBR tools described above (such as harmonic transposition and pre-flattening) in an MPEG-4 AAC bitstream is expected to be on the order of several hundred bits per second, because only the differential control data required to perform the eSBR processing is transmitted in accordance with some embodiments of the present invention. Legacy decoders can ignore this information because it is included in a backwards-compatible manner (as described below). Therefore, the negative bitrate impact associated with including eSBR metadata is negligible for a number of reasons, including: The bitrate penalty (due to including eSBR metadata) is only a small fraction of the total bitrate, because only the differential control data necessary to perform the eSBR processing is transmitted (not a simulcast of SBR control data); and · Adjustments to SBR-related control information are typically independent of the details of the transposition. Examples of cases where control data depends on the operation of the transposer are described later in this application.
[0071] Thus, embodiments of the present invention provide a means for efficiently transmitting enhanced spectrum band replication (eSBR) control data or metadata in a backward-compatible manner. This efficient transmission of eSBR control data reduces memory requirements in decoders, encoders, and transcoders employing aspects of the present invention while having no substantial adverse effect on bitrate. Furthermore, the complexity and processing requirements associated with implementing eSBR in accordance with embodiments of the present invention are reduced because the SBR data need only be processed once, rather than simulcast, as would be the case if eSBR were treated as an entirely separate object type in MPEG-4 AAC, rather than being integrated into the MPEG-4 AAC codec in a backward-compatible manner.
[0072] Elements of a block of an MPEG-4 AAC bitstream ("raw_data_block") containing eSBR metadata according to some embodiments of the present invention will now be described with reference to Figure 7. Figure 7 is a diagram of a block of an MPEG-4 AAC bitstream ("raw_data_block"), showing some segments of it.
[0073] A block of an MPEG-4 AAC bitstream may contain at least one "single_channel_element()" (e.g., the single channel element shown in FIG. 7) and / or at least one "channel_pair_element()" (not specifically shown in FIG. 7, but possibly present) containing audio data for an audio program. The block may also contain multiple "fill elements" (e.g., fill element 1 and / or fill element 2 in FIG. 7) containing data (e.g., metadata) related to the program. Each "single_channel_element()" contains an identifier (e.g., "ID1" in FIG. 7) indicating the start of a single channel element and may contain audio data representing different channels of a multi-channel audio program. Each "channel_pair_element()" contains an identifier (not shown in FIG. 7) indicating the start of a channel pair element and may contain audio data representing two channels of the program.
[0074] A "fill_element" in an MPEG-4 AAC bitstream (referred to herein as a fill element) includes an identifier ("ID2" in Figure 7) indicating the start of the fill element, followed by fill data. The identifier ID2 may consist of three unsigned integers ("uimsbf") transmitted most significant bit first with a value of 0x6. The fill data may include an "extension_payload()" element (referred to herein as an extension payload), whose syntax is shown in Table 4.57 of the MPEG-4 AAC standard. Several types of extension payloads exist, identified by the "extension_type" parameter, which is a 4-bit unsigned integer ("uimsbf") transmitted most significant bit first.
[0075] The fill data (e.g., its extension payload) may include a header or identifier (e.g., "Header 1" in Figure 7) that indicates the segment of the fill data that represents an SBR object (i.e., the header initializes an "SBR object" type called sbr_extension_data() in the MPEG-4 AAC standard). For example, a Spectral Band Replication (SBR) extension payload is identified with a value of '1101' or '1110' for the extension_type field in the header, where the identifier '1101' identifies the extension payload with SBR data and '1110' identifies the extension payload with SBR data together with a cyclic redundancy check (CRC) to verify the validity of the SBR data.
[0076] When a header (e.g., the extension_type field) initializes an SBR object type, SBR metadata (referred to herein as "spectral band replication data" and referred to as "sbr_data()" in the MPEG-4 AAC standard) follows the header, and at least one spectral band replication extension element (e.g., the "SBR extension element" in fill element 1 of Figure 7) can follow the SBR metadata. Such a spectral band replication extension element (a segment of a bitstream) is referred to as an "sbr_extension()" container in the MPEG-4 AAC standard. A spectral band replication extension element optionally includes a header (e.g., the "SBR extension header" in fill element 1 of Figure 7).
[0077] The MPEG-4 AAC standard assumes that the spectral band replication extension element can contain PS (parametric stereo) data of the program's audio data. The MPEG-4 AAC standard assumes that if a fill element's header (e.g., the header of its extension payload) initializes an SBR object type (such as "Header 1" in Figure 7) and the fill element's spectral band replication element contains PS data, the fill element (e.g., its extension payload) includes the spectral band replication data and a "bs_extension_id" parameter whose value (i.e., bs_extension_id=2) indicates that the PS data is included in the fill element's spectral band replication extension element.
[0078] According to some embodiments of the present invention, eSBR metadata (e.g., a flag indicating whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of the block) is included in the spectral band replication extension element of a fill element. For example, such a flag is shown in fill element 1 of FIG. 7, where the flag occurs after the header of the "SBR extension element" of fill element 1 ("SBR extension header" of fill element 1). Optionally, such a flag and additional eSBR metadata are included in the spectral band replication extension element after the header of the spectral band replication extension element (SBR extension element of fill element 1 of FIG. 7 after the SBR extension header). According to some embodiments of the present invention, a fill element that includes SBR metadata also includes a "bs_extension_id" parameter, the value of which (e.g., bs_extension_id = 3) indicates that eSBR metadata is included in the fill element and that eSBR processing should be performed on the audio content of the associated block.
[0079] According to some embodiments of the present invention, eSBR metadata is contained within a fill element (e.g., fill element 2 in FIG. 7) of an MPEG-4 AAC bitstream, rather than within the spectral band duplication extension element (SBR extension element) of the fill element. This is because a fill element containing an extension_payload() with SBR data or SBR data with a CRC does not contain any other extension payload of any other extension type. Therefore, in embodiments in which the eSBR metadata is stored in its own extension payload, a separate fill element is used to store the eSBR metadata. Such a fill element includes an identifier indicating the start of the fill element (e.g., "ID2" in FIG. 7) followed by the fill data. The fill data includes an extension_payload() element (sometimes referred to herein as the extension payload), the syntax of which is shown in Table 4.57 of the MPEG-4 AAC Standard. The fill data (e.g., its extension payload) includes a header (e.g., Header 2 of fill element 2 in FIG. 7) that indicates an eSBR object (i.e., the header initializes an enhanced spectral band replication (eSBR) object type), and the fill data (e.g., its extension payload) includes eSBR metadata after the header. For example, fill element 2 in FIG. 7 includes such a header ("Header 2"), which includes eSBR metadata after the header (i.e., "Flag" for fill element 2, which indicates whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of the block). Optionally, additional eSBR metadata is also included in the fill data of fill element 2 in FIG. 7 after Header 2. In the embodiment described in this paragraph, the header (e.g., Header 2 in FIG. 7) is not one of the conventional values specified in Table 4.57 of the MPEG-4 AAC Standard, but rather has an identifying value that indicates an eSBR extension payload (so that the extension_type field of the header indicates that the fill data includes eSBR metadata).
[0080] In a first class of embodiments, the present invention is an audio processing unit (e.g., a decoder) comprising: a memory (e.g., buffer 201 of FIG. 3 or 4) configured to store at least one block of an encoded audio bitstream (e.g., at least one block of an MPEG-4 AAC bitstream); a bitstream payload deformatter (e.g., element 205 of FIG. 3 or element 215 of FIG. 4 ), coupled to the memory and configured to demultiplex at least a portion of the blocks of the bitstream; and a decoding subsystem configured and coupled to decode at least a portion of the audio content of the block of the bitstream (e.g., elements 202 and 203 of FIG. 3 or elements 202 and 213 of FIG. 4); and the block contains: A fill element containing an identifier indicating the start of the fill element and the fill data following the identifier (e.g., the "id_syn_ele" identifier has the value 0x6 in Table 4.85 of the MPEG-4 AAC standard), where the fill data is: Contains at least one flag that identifies whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of the block (e.g., using eSBR metadata and spectral band replication data included in the block).
[0081] Flags are eSBR metadata; an example of a flag is sbrPatchingModeflag. Another example of a flag is the harmonicSBR flag. Both of these flags indicate whether a basic form of spectral band replication or an enhanced form of spectral replication is performed on the audio data of a block. A basic form of spectral replication is spectral patching, and an enhanced form of spectral band replication is harmonic transposition.
[0082] In some embodiments, the fill data also includes additional eSBR metadata (ie, eSBR metadata other than flags).
[0083] The memory may be a buffer memory (eg, an implementation of buffer 201 of FIG. 4) that stores (eg, in a non-transient manner) at least one block of the encoded audio bitstream.
[0084] The performance details of eSBR processing (using eSBR harmonic transposition and pre-flattening) by an eSBR decoder during decoding of an MPEG-4 AAC bitstream containing eSBR metadata (shown below as eSBR tools) would be as follows (for a typical decoding with the specified parameters): Harmonic transposition (16kbps, 14400 / 28800Hz) DFT-based: 3.68 WMOPS (weighted million operations per second) QMF base: 0.98 WMOPS QMF patch processing / pre-processing (pre-flattening): 0.1 WMOPS It is known that DFT-based transposition typically performs better than QMF-based transposition for transients.
[0085] According to some embodiments of the present invention, a fill element (of the encoded audio bitstream) containing eSBR metadata may also include a parameter (e.g., bs_extension_id parameter) whose value (e.g., bs_extension_id=3) indicates that eSBR metadata is included in the fill element and that eSBR processing should be performed on the audio content of the associated block, and / or It includes a parameter (eg, the same "bs_extension_id" parameter) whose value (eg, bs_extension_id=2) indicates that the fill element's sbr_extension() container contains PS data. For example, as shown in Table 1 below, a parameter with a value bs_extension_id=2 may indicate that the fill element's sbr_extension() container contains PS data, and a parameter with a value bs_extension_id=3 may indicate that the fill element's sbr_extension() container contains eSBR metadata. Table 1 [Table 1]
[0086] According to some embodiments of the present invention, the syntax of each spectral band replication extension element containing eSBR metadata and / or PS data is as shown in Table 2 below (where "sbr_extension()" indicates the container that is the spectral band replication extension element, "bs_extension_id" is as shown in Table 1 above, "ps_data" indicates the PS data, and "esbr_data" indicates the eSBR metadata). Table 2 [Table 2]
[0087] In an exemplary embodiment, the esbr_data() referenced in Table 2 above indicates values for the following metadata parameters: 1.1-bit metadata parameter "bs_sbr_preprocessing"; and 2. For each channel ("ch") of the audio content of the coded bitstream to be decoded, each of the parameters mentioned above: "sbrPatchingMode[ch]"; "sbrOversamplingFlag[ch]"; "sbrPitchInBinsFlag[ch]"; and "sbrPitchInBins[ch]". For example, in some embodiments, to indicate these metadata parameters, esbr_data() may have the syntax shown in Table 3: Table 3 [Table 3-1] [Table 3-2]
[0088] The above syntax allows for the efficient implementation of enhanced forms of spectral band duplication, such as harmonic transposition, as an extension to legacy decoders. Specifically, the eSBR data in Table 3 includes only those parameters required to perform the enhanced form of spectral band duplication that are not already supported in the bitstream and cannot be directly derived from parameters already supported in the bitstream. All other parameters and processing data required to perform the enhanced form of spectral band duplication are extracted from existing parameters at already-defined locations in the bitstream.
[0089] For example, a decoder compliant with MPEG-4 HE-AAC or HE-AACv2 may be extended to include an enhanced form of spectral band replication, such as harmonic transposition. This enhanced form of spectral band replication is in addition to the basic form of spectral band replication already supported by the decoder. In the case of a decoder compliant with MPEG-4 HE-AAC or HE-AACv2, the basic form of spectral band replication is the QMF spectral patching SBR tool defined in section 4.6.18 of the MPEG-4 AAC standard.
[0090] When performing the enhanced form of spectral band replication, the enhanced HE-AAC decoder can reuse many bitstream parameters already included in the SBR extension payload of the bitstream. Certain specific parameters that may be reused include, for example, various parameters that determine the master frequency band table. These parameters include bs_start_freq (a parameter that determines the start of the master frequency table parameters), bs_stop_freq (a parameter that determines the end of the master frequency table parameters), bs_freq_scale (a parameter that determines the number of frequency bands per octave), and bs_alter_scale (a parameter that alters the scale of the frequency bands). Potentially reused parameters include parameters that determine the noise band table (bs_noid_bands) and the limiter band table parameter (bs_limiter_bands). Therefore, in various embodiments, at least some equivalent parameters specified in the USAC standard are omitted from the bitstream, thereby reducing control overhead in the bitstream. Typically, when a parameter specified in the AAC standard has an equivalent parameter specified in the USAC standard, the equivalent parameter specified in the USAC standard has the same name as the parameter specified in the AAC standard, e.g., envelope scale factor E OrigMapped (the envelope scale factor EOrigMapped However, the equivalent parameters specified in the USAC standard usually have different values, which are "adjusted" for the enhanced SBR treatment specified in the USAC standard, but not for the SBR treatment specified in the AAC standard.
[0091] To improve subjective quality for audio content with harmonic frequency structure and strong tonal characteristics, especially at low bitrates, the activation of enhanced SBR is recommended. The value of the corresponding bitstream element (i.e., esbr_data()) may be determined in the encoder by controlling these tools and applying a signal-dependent classification mechanism. In general, using the harmonic patching method (sbrPatchingMode==1) is preferable when encoding music signals at very low bitrates, where the core codec may be severely limited in audio bandwidth. This is especially true if these signals contain significant harmonic structure. Conversely, using the regular SBR patching method is preferable for speech and mixed signals because it provides better preservation of the temporal structure in speech.
[0092] To improve the performance of the harmonic transposer, a preprocessing step can be activated (bs_sbr_preprocessing == 1) that seeks to avoid introducing spectral discontinuities into the signal going to the subsequent envelope adjuster. This operation is beneficial for signal types where the coarse spectral envelope of the low-band signal used for high-frequency reconstruction exhibits large fluctuations in level.
[0093] To improve the transient response of harmonic SBR patch processing, it is possible to apply signal-adaptive frequency-domain oversampling (sbrOversamplingFlag==1). Since signal-adaptive frequency-domain oversampling increases the computational complexity of the transponder but only benefits frames containing transients, the use of this tool is controlled by a bitstream element, which is transmitted once per frame and once per independent SBR channel.
[0094] Decoders operating in the proposed enhanced SBR mode typically need to be able to switch between legacy and enhanced SBR patch processing. Therefore, depending on the decoder configuration, a delay may be introduced that can be as long in duration as the duration of one core audio frame. Typically, the delay for both legacy and enhanced SBR patch processing will be similar.
[0095] In addition to a number of parameters, other data elements may also be reused by an enhanced HE-AAC decoder when performing enhanced spectral band replication according to embodiments of the present invention. For example, envelope data and noise floor data may also be extracted from the bs_data_env (envelope scalefactors) and bs_noid_env (noise floor scalefactors) data and used during enhanced spectral band replication.
[0096] Essentially, these embodiments utilize configuration parameters and envelope data already supported by legacy HE-AAC or HE-AACv2 decoders in the SBR extension payload to enable enhanced forms of spectral band replication that require as little extra transmission data as possible. While metadata was originally tailored to the base form of HFR (e.g., the spectral transformation process in SBR), in embodiments, it is used for enhanced forms of HFR (e.g., the harmonic transposition process in eSBR). As previously mentioned, the metadata generally expresses operational parameters (e.g., envelope scale factor, noise floor scale factor, time / frequency grid parameters, sinusoidal summation information, variable crossover frequencies / bands, inverse filtering mode, envelope resolution, smoothing mode, frequency interpolation mode) intended for use with the base form of HFR (e.g., the linear spectral transformation). However, this metadata, combined with additional metadata parameters specific to the HFR enhanced format (e.g., harmonic transposition), can be used to efficiently and effectively process audio data using the HFR enhanced format.
[0097] Therefore, enhanced decoders supporting the enhanced form of spectral band duplication can be created in a very efficient manner by relying on already defined bitstream elements (e.g., elements in the SBR extension payload) and adding only the parameters necessary to support the enhanced form of spectral band duplication (to the fill element extension payload). This data reduction property, combined with placing newly added parameters in reserved data fields such as extension containers, significantly reduces the barrier to creating decoders that support the enhanced form of spectral band duplication by ensuring that the bitstream is backward compatible with legacy decoders that do not support the enhanced form of spectral band duplication. It will be understood that reserved data fields are backward-compatible data fields, i.e., data fields already supported by earlier decoders, such as legacy HE-AAC or HE-AACv2 decoders. Similarly, extension containers are backward-compatible extension containers, i.e., extension containers already supported by earlier decoders, such as legacy HE-AAC or HE-AACv2 decoders. In Table 3, the numbers in the right column indicate the number of bits of the corresponding parameter in the left column.
[0098] In some embodiments, the SBR object type defined in MPEG-4 AAC is updated to include extended SBR (eSBR) tool features and SBR-Tool, as indicated by the SBR extension element (bs_extension_id==EXTENSION_ID_ESBR). When a decoder detects this SBR extension element, it uses the signaled features of the extended SBR tool.
[0099] In some embodiments, the present invention is a method that includes encoding audio data to generate an encoded bitstream (e.g., an MPEG-4 AAC bitstream), where at least one segment of at least one block of the encoded bitstream includes eSBR metadata and at least one other segment of the block includes audio data. In an exemplary embodiment, the method includes multiplexing the audio data and the eSBR metadata in each block of the encoded bitstream. In exemplary decoding of the encoded bitstream in an eSBR decoder, the decoder extracts the eSBR metadata from the bitstream (including separating and parsing the eSBR metadata and the audio data), processes the audio data using the eSBR metadata, and generates a stream of decoded audio data.
[0100] Another aspect of the present invention is an eSBR decoder configured to perform eSBR processing (e.g., using at least one of the eSBR tools known as harmonic transposition or pre-flattening) during decoding of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that does not include eSBR metadata. An example of such a decoder is described with reference to FIG. 5.
[0101] The eSBR decoder (400) of Figure 5 includes a buffer memory 201 (same as memory 201 of Figures 3 and 4), a bitstream payload deformatter 215 (same as deformatter 215 of Figure 4), an audio decoding subsystem 202 (sometimes referred to as the "core" decoding stage or "core" decoding subsystem, and same as core decoding subsystem 202 of Figure 3), an eSBR control data generation subsystem 401, and an eSBR processing stage 203 (same as stage 203 of Figure 3), connected as shown. Typically, decoder 400 also includes other processing elements (not shown).
[0102] In operation of the decoder 400 , a sequence of blocks of an encoded audio bitstream (MPEG-4 AAC bitstream) received by the decoder 400 is asserted from the buffer 201 to the deformatter 215 .
[0103] The deformatter 215 is configured and coupled to separate each block of the bitstream and extract SBR metadata (including quantized envelope data) and typically other metadata therefrom. The deformatter 215 is configured to assert at least the SBR metadata to the eSBR processing stage 203. The deformatter 215 is also configured and coupled to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.
[0104] The audio decoding subsystem 202 of the decoder 400 is configured to decode the audio data extracted by the deformatter 215 (such decoding may be referred to as the "core" decoding process) to generate decoded audio data and assert the decoded audio data to the eSBR processing stage 203. The decoding is performed in the frequency domain. Typically, the final stage of processing in the subsystem 202 applies a frequency-domain to time-domain transformation to the decoded frequency-domain audio data, such that the output of the subsystem is time-domain decoded audio data. The stage 203 is configured to apply the SBR tools (and eSBR tools) indicated by the SBR metadata (extracted by the deformatter 215) and the eSBR metadata generated by the subsystem 401 to the decoded audio data (i.e., perform SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata) to generate the fully decoded audio data that is output from the decoder 400. Typically, decoder 400 includes a memory (accessible by subsystem 202 and stage 203) that stores deformatted audio data and metadata output from deformatter 215 (and optionally subsystem 401), with stage 203 configured to access the audio data and metadata as needed during SBR and eSBR processing. The SBR processing in stage 203 may be considered post-processing on the output of core decoding subsystem 202. Optionally, decoder 400 also includes a final upmixing subsystem (which may apply parametric stereo (PS) tools defined in the MPEG-4 AAC standard using the PS metadata extracted by deformatter 215) configured and coupled to perform upmixing on the output of stage 203 and generate fully decoded upmixed audio that is output from APU 210.
[0105] Parametric stereo is a coding tool that represents stereo signals using a linear downmix of the left and right channels of the stereo signal and a set of spatial parameters that describe the stereo image. Parametric stereo typically uses three types of spatial parameters: (1) inter-channel intensity difference (IID), which describes the intensity difference between the channels; (2) inter-channel phase difference (IPD), which describes the phase difference between the channels; and (3) inter-channel coherence (ICC), which describes the coherence (or similarity) between the channels. Coherence may be measured as the maximum value of the cross-correlation as a function of time or phase. These three parameters generally enable high-quality reconstruction of the stereo image. However, the IPD parameter only specifies the relative phase differences between the channels of a stereo input signal and does not indicate the distribution of these phase differences across the left and right channels. Therefore, a fourth type of parameter, describing the overall phase offset or overall phase difference (OPD), may be additionally used. In the stereo reconstruction process, consecutive windowed segments of both the received downmix signal s[n] and the decorrelated version of the received downmix d[n] are processed together with the spatial parameters to obtain the left(l) as follows: k (n)) and right (r k Generate the reconstructed signal of (n): l k (n)=H 11 (k,n)s k (n)+H 21 (k,n)d k (n) r k (n)=H 12 (k,n)s k (n)+H 22 (k,n)d k (n) where H 11 , H 12 , H 21 and H 22 is defined by the stereo parameters. k (n) and signal r k(n) is finally transformed into the time domain by a frequency-to-time transformation.
[0106] 5 is configured and coupled to detect at least one characteristic of the encoded audio bitstream to be decoded and to generate eSBR control data (which may, in other embodiments of the present invention, be or include any type of eSBR metadata contained in the encoded audio bitstream) in response to at least one result of the detection step. Upon detection of a particular characteristic (or combination of characteristics) of the bitstream, the eSBR control data is asserted to stage 203 to trigger and / or control the application of individual eSBR tools or combinations of eSBR tools. For example, to control the performance of eSBR processing using harmonic transposition, some embodiments of the control data generation subsystem 401 would include: a music detector that sets the sbrPatchingMode[ch] parameter (and asserts the setting parameter to stage 203) in response to detecting whether the bitstream indicates music; a transient detector that sets the sbrOversamplingFlag[ch] parameter (and asserts the setting parameter to stage 203) in response to detecting the presence or absence of transients in the audio content indicated by the bitstream; and / or a pitch detector that sets the sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch] parameters (and asserts the setting parameter to stage 203) in response to detecting the pitch of the audio content indicated by the bitstream. Another aspect of the present invention is an audio bitstream decoding method performed by any of the embodiments of the decoder of the present invention described in this and the preceding paragraphs.
[0107] Aspects of the present invention include encoding or decoding methods of the type that any embodiment of an APU, system, or device of the present invention is configured (e.g., programmed) to perform. Other aspects of the present invention include systems or devices configured (e.g., programmed) to perform any embodiment of the method of the present invention, and computer-readable media (e.g., disks) storing (e.g., non-transitory) code for performing any embodiment of the method of the present invention or steps thereof. For example, a system of the present invention can be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform any of various operations on data (including embodiments of the method of the present invention or steps thereof). Such a general-purpose processor can be or include a computer system including input devices, memory, and processing circuitry (programmed to perform embodiments of the method of the present invention (or steps thereof) in response to data being asserted).
[0108] Embodiments of the present invention may be implemented in hardware, firmware, or software, or a combination of both (e.g., programmable logic arrays). Unless otherwise specified, the algorithms or processes included as part of the present invention are not inherently related to any particular computer or other apparatus. In particular, various general-purpose machines may be used with programs written in accordance with the teachings herein, or it may be more profitable to construct a more specialized apparatus (e.g., an integrated circuit) to perform the required method steps. Thus, the present invention may be implemented in one or more computer programs executing on one or more programmable computer systems (e.g., an implementation of any of the elements of FIG. 1, or the encoder 100 (or elements thereof) of FIG. 2, or the decoder 200 (or elements thereof) of FIG. 3, or the decoder 210 (or elements thereof) of FIG. 4, or the decoder 400 (or elements thereof) of FIG. 5), each of which includes at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and to generate output information that is applied to one or more output devices, in known fashion.
[0109] Each such program may be implemented in any desired computer language (including machine, assembly, or high-level procedural, logical, or object-oriented programming languages) to communicate with a computer system, and in any case, the language may be a compiled or interpreted language.
[0110] For example, when implemented by computer software instruction sequences, the various functions and steps of embodiments of the present invention may be realized by multi-threaded software instruction sequences run on suitable digital signal processing hardware, in which case the various devices, steps and functions of the embodiments may correspond to portions of the software instructions.
[0111] Each such computer program is preferably stored on or downloaded to a general-purpose or special-purpose programmable computer-readable storage medium or device (e.g., solid-state memory or media, or magnetic or optical media) to configure and operate the computer when the storage medium or device is read by the computer system to perform the procedures described herein. The system of the present invention can be implemented as a computer-readable storage medium configured with (i.e., storing) a computer program, which causes the computer system to operate in a particular, predetermined manner to perform the functions described herein.
[0112] Many embodiments of the present invention have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the present invention. Many modifications and variations of the present invention are possible in light of the above teachings. For example, phase shifting may be used in combination with complex QMF analysis and synthesis filter banks to facilitate efficient implementation. The analysis filter bank is responsible for filtering the time-domain low-band signal generated by the core decoder into multiple subbands (e.g., QMF subbands). The synthesis filter bank is responsible for combining the reconstructed high-band generated by the selected HFR technique with the decoded low-band (as indicated by the received sbrPatchingMode parameter) to generate a wideband output audio signal. However, a given filter bank implementation operating in a particular sample rate mode, such as normal dual-rate operation or down-sampled SBR mode, should not have bitstream-dependent phase shifting. The QMF bank used in SBR is a complex-exponential extension of cosine-modulated filter bank theory. It can be shown that extending a cosine-modulated filter bank with complex-exponential modulation no longer requires the alias cancellation constraint. Therefore, in the SBR QMF bank, the analysis filter h k (n) and synthesis filter f k Both (n) can be defined as follows:
number
[0113] The coefficients p0(n) of the prototype filter can be defined with a length L of 640 as shown in Table 4 below. Table 4 [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5] [Table 4-6] The prototype filter, p0(n), may also be derived from Table 4 by one or more mathematical operations such as rounding, subsampling, interpolation, and decimation.
[0114] While the adjustment of SBR-related control information typically does not depend on the details of the transposition (as described above), in some embodiments, certain elements of the control data may be simulcast in the eSBR extension container (bs_extension_id==EXTENSION_ID_ESBR) to improve the quality of the regenerated signal. Some of the simulcast elements may include noise floor data (e.g., noise floor scale factors and parameters indicating the direction in either the frequency or time direction for delta coding for each noise floor), inverse filtering data (e.g., parameters indicating the inverse filtering mode selected from no inverse filtering, low-level inverse filtering, medium-level inverse filtering, and strong-level inverse-inverse filtering), and missing harmonics data (e.g., parameters indicating whether a sinusoid should be added to a specific frequency band of the regenerated high band). All of these elements depend on the synthetic emulation of the decoder's transposer performed within the encoder and, therefore, if properly adjusted for the selected transposer, can increase the quality of the regenerated signal.
[0115] Specifically, in some embodiments, missing harmonic and inverse filtering control data are transmitted in the eSBR extension container (along with other bitstream parameters in Table 3) and adjusted for the eSBR harmonic transposer. For the eSBR harmonic transposer, the additional bitrate required to transmit these two classes of metadata is relatively small. Therefore, transmitting adjusted missing harmonic and / or inverse filtering control data in the eSBR extension container will increase the audio quality produced by the transposer with minimal impact on bitrate. To ensure backward compatibility with legacy decoders, adjusted parameters for the SBR spectral transformation process may be transmitted in the bitstream as part of the SBR control data using either implicit or explicit signaling.
[0116] It is to be understood that, within the scope of the appended claims, the invention may be practiced otherwise than as specifically described herein. Any reference numerals that may be included in the following claims are for illustrative purposes only and should not be used to interpret or limit the scope of the claims in any manner. Various aspects of the present disclosure will be understood from the following enumerated exemplary embodiments (EEE):
[0117] EEE1. 1. A method for performing high frequency reconstruction of an audio signal, comprising: receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata; and decoding the audio data to generate a decoded low-band audio signal; extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata comprising operational parameters of a high frequency reconstruction process, the operational parameters comprising a patch processing mode parameter located in an extension container of the encoded audio bitstream, the patch processing mode parameter having a first value indicating a spectral transformation and a second value indicating a harmonic transposition with phase vocoder frequency spreading; filtering the decoded low-band audio signal to generate a filtered low-band audio signal; regenerating a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein the regeneration includes a spectral transformation when the patch processing mode parameter is the first value, and wherein the regeneration includes harmonic transposition with phase vocoder frequency spreading when the patch processing mode parameter is the second value; and combining the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal. EEE2. The method of EEE1, wherein the extension container includes inverse filtering control data to be used when the patching mode parameter is equal to the second value. EEE3. The method of any one of EEE1-2, wherein the extension container further includes missing harmonic control data to be used when the patching mode parameter is equal to the second value. EEE4. A method as described in any preceding EEE, wherein the encoded audio bitstream includes a fill element (having an identifier indicating the beginning of the fill element) and fill data following the identifier, the fill data including the extension container. EEE5. The method of claim 8, wherein the identifier is a 3-bit unsigned integer transmitted most significant bit first and has a value of 0x6. EEE6. The fill data includes an extension payload, the extension payload including spectrum band duplication extension data, the extension payload being identified by a 4-bit unsigned integer transmitted most significant bit first and having a value of '1101' or '1110', and optionally the spectrum band duplication extension data being: Optional spectral band duplication header, the spectral band replication data after the header, and Spectral band replication extension element after said spectral band replication data and wherein the spectral band duplication extension element includes a flag. EEE7. The method of any one of EEE1 to EEE6, wherein the high-frequency reconstruction metadata includes parameters indicating an envelope scale factor, a noise floor scale factor, time / frequency grid information, or crossover frequencies. EEE8. The filtering is performed by the analysis filter h, which is a modulated version of the prototype filter p0(n). k (n) according to the following equation:
number
[0118] [Patent Document 1] US Patent Application Publication No. 2018 / 0025737
Claims
1. 1. A method for performing high frequency reconstruction of an audio signal, comprising: receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata; decoding the audio data to generate a decoded low-band audio signal; extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata including operational parameters of a high frequency reconstruction process, the operational parameters including a patch processing mode parameter located in a backward compatible extension container of the encoded audio bitstream, the patch processing mode parameter having a first value indicating a spectral transformation and a second value indicating a harmonic transposition with phase vocoder frequency spreading, the encoded audio bitstream further including a fill element having an identifier indicating a start of the fill element and fill data following the identifier, the fill data including the backward compatible extension container, the identifier being a 3-bit unsigned integer having a value of 0x6 transmitted most significant bit first; filtering the decoded low-band audio signal to generate a filtered low-band audio signal, the filtering being performed using a prototype filter p 0 The analysis filter h is a modulated version of (n). k (n) according to the following equation: [Equation 1] Here, p 0 (n) is a real-valued symmetric or asymmetric prototype filter, and M is the is the number of channels in the analysis filter bank, and N is the next is a number, and the prototype filter p 0 (n) is derived from the coefficients in Table 4 of the specification, step; regenerating a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein the regeneration includes a spectral transformation when the patch processing mode parameter is the first value, and wherein the regeneration includes harmonic transposition with phase vocoder frequency spreading when the patch processing mode parameter is the second value; wherein certain elements of control data are simulcast in the backward compatible extension container, the certain elements of control data including at least one of inverse filtering control data, noise floor control data, and missing harmonics control data for use in regenerating a signal.
2. The method of claim 1 , wherein the backward compatible extension container includes inverse filtering control data to be used when the patching mode parameter is equal to the second value.
3. The method of claim 1 , wherein the backward compatible extension container further comprises missing harmonic control data to be used when the patching mode parameter is equal to the second value.
4. 2. The method of claim 1, wherein a phase shift is added to the filtered low-band audio signal after the filtering and compensated for before synthesis to reduce the complexity of the method.
5. A non-transitory computer-readable storage medium containing instructions that, when executed by a processor, cause the method of claim 1 to be performed.
6. 1. An audio processing unit for performing high frequency reconstruction of an audio signal, comprising: an input interface for receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata; a core audio decoder for decoding the audio data to generate a decoded low-band audio signal; a deformatter that extracts the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata including operational parameters for a high frequency reconstruction process, the operational parameters including a fill element having an identifier indicating the beginning of the fill element and fill data following the identifier, the fill data including a backward compatible extension container including a patch processing mode parameter, the patch processing mode parameter having a first value indicating a spectral transformation and a second value indicating a harmonic transposition with phase vocoder frequency spreading, the identifier being a 3-bit unsigned integer having a value of 0x6 transmitted most significant bit first; an analysis filterbank for filtering the decoded low-band audio signal to generate a filtered low-band audio signal, the filtering being performed using a prototype filter p 0 The analysis filter h is a modulated version of (n). k (n) according to the following equation: [Equation 2] Here, p 0 (n) is a real-valued symmetric or asymmetric prototype filter, M is the number of channels in the analysis filterbank, N is the order of the prototype filter, and the prototype filter p 0 (n) is the analysis filter bank, derived from the coefficients in Table 4 of the specification; a high frequency regeneration unit for regenerating a high frequency portion of the audio signal using the filtered low frequency audio signal and the high frequency reconstruction metadata, the regeneration including a spectral transformation when the patch processing mode parameter is the first value, and the regeneration including a harmonic transposition with phase vocoder frequency spreading when the patch processing mode parameter is the second value; and wherein certain elements of control data are simulcast in the backward compatible extension container, the certain elements of control data including at least one of inverse filtering control data, noise floor control data, and missing harmonics control data for use in regenerating a signal.
Citation Information
Patent Citations
Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element
US20180025737A1