Integration of high frequency reconstruction techniques with reduced post-processing delay
By introducing enhanced spectrum band replication (eSBR) technology into audio coding, and utilizing harmonic transpose and QMF patching, the problem that spectrum patching is not suitable for low cross-frequency music is solved, thereby improving the high-frequency reconstruction quality and coding efficiency of audio signals.
Patent Information
- Application Number
- CN202111584446.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-04-25
- Filing Date
- 2019-04-25
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2039-04-25
AI Technical Summary
Existing audio coding techniques are unsuitable for spectrum patching methods during spectrum copying, especially for music content with relatively low cross frequencies, leading to a decline in audio quality.
The enhanced spectrum band replication (eSBR) technology is employed, which improves the spectrum band replication process by using harmonic transposition and QMF patching for additional preprocessing, combined with high-frequency reconstruction metadata, to achieve accurate reconstruction of the high-frequency band.
It improves the reconstruction quality of the high-frequency components of audio signals, especially the sound quality of low cross-frequency music content, and enhances encoding efficiency and decoding effect.
Smart Images

Figure CN114242086B_ABST
Abstract
Description
[0001] Related application data
[0002] This case is a divisional application. The parent of this divisional application is the patent application for “Integration of High Frequency Reconstruction Techniques with Reduced Post-Processing Latency” having an application date of April 25, 2019, application number 201980034811.4.
[0003] Cross reference to related applications
[0004] This application claims priority benefit of U.S. Provisional Patent Application No. 62 / 662,296, filed April 25, 2018, the entirety of which is incorporated herein by reference. TECHNICAL FIELD
[0005] Embodiments relate to audio signal processing, and more specifically, embodiments relate to encoding, decoding, or transcoding an audio bitstream using control data that specifies performing a basic form of high frequency reconstruction (“HFR”) or an enhanced form of HFR on audio data. BACKGROUND
[0006] A typical audio bitstream includes both audio data (e.g., encoded audio data) indicative of one or more channels of audio content and metadata indicative of at least one characteristic of the audio data or audio content. One well-known format for producing an encoded audio bitstream is the MPEG-4 Advanced Audio Coding (AAC) format described in the MPEG standard ISO / IEC 14496-3:2009. In the MPEG-4 standard, AAC stands for “Advanced Audio Coding” and HE-AAC stands for “High-Efficiency Advanced Audio Coding.”
[0007] The MPEG-4 AAC standard defines several audio profiles that determine which objects and encoding tools are present in a compatible encoder or decoder. Three of these audio profiles are (1) the AAC profile, (2) the HE-AAC profile, and (3) the HE-AAC v2 profile. The AAC profile includes the AAC Low Complexity (or “AAC-LC”) object type. The AAC-LC object is the counterpart of the MPEG-2 AAC Low Complexity profile with some adjustments and does not include both the Spectral Band Replication (“SBR”) object type and the Parametric Stereo (“PS”) object type. The HE-AAC profile is a superset of the AAC profile and additionally includes the SBR object type. The HE-AAC v2 profile is a superset of the HE-AAC profile and additionally includes the PS object type.
[0008] The SBR object type contains a spectral band replication tool, which is an important high frequency reconstruction ("HFR") coding tool that can significantly improve the compression efficiency of perceptual audio codecs. SBR reconstructs high frequency components of an audio signal on the receiver side (e.g., in a decoder). Thus, an encoder only needs to encode and transmit low frequency components to allow much higher audio quality at low data rates. SBR is based on the available bandwidth limited signal and control data obtained from the encoder to replicate the truncated harmonic sequence previously for reduced data rates. The ratio between tonal and noise-like components is maintained by an adaptive inverse filtering and optionally adding noise and sinusoids. In the MPEG-4 AAC standard, the SBR tool performs spectral patching (also known as linear panning or spectral panning), in which several consecutive Quadrature Mirror Filter (QMF) subbands are copied (or "patched") from the transmitted low frequency band portion of an audio signal to the high frequency band portion of the audio signal (which is generated in the decoder).
[0009] Spectral patching or linear panning can not be suitable for certain audio types (e.g., music content with relatively low cross-over frequencies). Thus, there is a need for techniques for improving spectral band replication. SUMMARY
[0010] A first class of embodiments relates to a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream and decoding the audio data to produce a decoded low-band audio signal. The method further includes extracting high frequency reconstruction metadata and filtering the decoded low-band audio signal using an analysis filter bank to produce a filtered low-band audio signal. The method further includes extracting a flag indicating whether spectral panning or harmonic transposition is performed on the audio data and regenerating a high-band portion of the audio signal using the filtered low-band audio signal and the high frequency reconstruction metadata according to the flag. Finally, the method includes combining the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal.
[0011] A second class of embodiments relates to an audio decoder for decoding an encoded audio bitstream. The decoder includes an input interface for receiving the encoded audio bitstream, wherein the encoded audio bitstream includes audio data representing a low-band portion of an audio signal, and a core decoder for decoding the audio data to generate a decoded low-band audio signal. The decoder also includes a demultiplexer for extracting high-frequency reconstruction metadata from the encoded audio bitstream, wherein the high-frequency reconstruction metadata includes operational parameters for a high-frequency reconstruction process that linearly translates a number of consecutive subbands from the low-band portion of the audio signal to a high-band portion of the audio signal, and an analysis filterbank for filtering the decoded low-band audio signal to generate a filtered low-band audio signal. The decoder further includes a demultiplexer for extracting a flag from the encoded audio bitstream that indicates whether to perform linear translation or harmonic transposition on the audio data, and a high-frequency regenerator for regenerating the high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata in accordance with the flag. Finally, the decoder includes a synthesis filterbank for combining the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal.
[0012] Other class of embodiments relates to encoding and transcoding audio bitstreams containing metadata identifying whether an enhanced spectral band replication (eSBR) processing is performed. BRIEF DESCRIPTION OF DRAWINGS
[0013] Figure 1 is a block diagram of an embodiment of a system that can be configured to perform embodiments of the inventive method.
[0014] Figure 2 is a block diagram of an encoder that is an embodiment of the inventive audio processing unit.
[0015] Figure 3 is a block diagram of a system that includes a decoder that is an embodiment of the inventive audio processing unit and also optionally includes a post-processor coupled to the decoder.
[0016] Figure 4 is a block diagram of a decoder that is an embodiment of the inventive audio processing unit.
[0017] Figure 5 is a block diagram of a decoder that is another embodiment of the inventive audio processing unit.
[0018] Figure 6 is a block diagram of another embodiment of the inventive audio processing unit.
[0019] Figure 7is a block diagram of an MPEG-4 AAC bitstream, containing several segments into which it is divided.
[0020] Symbols and terms
[0021] In this disclosure, including in the claims, the expression "performing an operation on" a signal or data (e.g., filtering, scaling, transforming the signal or data, or applying a gain to the signal or data) is used in a broad sense to mean performing the operation directly on the signal or data or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performing the operation thereon).
[0022] In this disclosure, including in the claims, the expression "audio processing unit" or "audio processor" is used in a broad sense to mean a system, device, or apparatus configured to process audio data. Examples of audio processing units include, but are not limited to, encoders, transcoders, decoders, codecs, pre-processing systems, post-processing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools). Almost all consumer electronics products (e.g., mobile phones, televisions, laptop computers, and tablet computers) contain an audio processing unit or audio processor.
[0023] In this disclosure, including in the claims, the term "coupled" or "coupling" is used in a broad sense to mean directly or indirectly connected. Thus, if a first device is coupled to a second device, the connection can be direct or indirect via another device and connections. Also, a component integrated into or with another component is coupled to the other component. DETAILED DESCRIPTION
[0024] The MPEG-4 AAC standard contemplates that an encoded MPEG-4 AAC bitstream contains metadata that indicates each type of high frequency reconstruction ("HFR") processing applied (if to be applied) by a decoder to decode the audio content of the bitstream, and / or controls such HFR processing, and / or indicates at least one characteristic or parameter of at least one HFR tool used to decode the audio content of the bitstream. In this document, we use the expression "SBR metadata" to mean this type of metadata for use with spectral band replication ("SBR"), as described or referred to in the MPEG-4 AAC standard. It will be appreciated by those skilled in the art that SBR is a form of HFR.
[0025] SBR is preferably used as a dual-rate system, where the base codec operates at half the original sampling rate and SBR operates at the original sampling rate. The SBR encoder works in parallel with the base core codec, although with a higher sampling rate. Although SBR is mainly a post-processing in the decoder, important parameters are extracted in the encoder to ensure the most accurate high frequency reconstruction in the decoder. The encoder estimates a spectral envelope of the SBR range that is appropriate for the time and frequency resolution characteristic of the current input signal segment. The spectral envelope is estimated by complex QMF analysis and subsequent energy calculation. The time and frequency resolution of the spectral envelope can be chosen highly freely to ensure the most appropriate time frequency resolution for a given input segment. The envelope estimation needs to take into account that transients of original origin that are mainly located in the high frequency region (e.g. a kick drum) will be present to a lesser extent in the high frequency band generated by SBR, because the high frequency band in the decoder is based on the low frequency band where the transients are much less prominent than in the high frequency band, before the envelope is adjusted. This aspect puts different requirements on the time frequency resolution of the spectral envelope data compared to the general spectral envelope estimation used in other audio coding algorithms.
[0026] In addition to the spectral envelope, several additional parameters are extracted that represent spectral characteristics of the input signal for different time and frequency regions. Since the encoder has naturally access to the original signal as well as information about how the SBR unit in the decoder will generate the high frequency band, the system can handle the situation where the low frequency band constitutes a strong harmonic series and the re-generated high frequency band mainly constitutes random signal components, as well as the situation where strong tonal components exist in the original high frequency band that have no counterpart in the low frequency band (the high frequency band region is based on this), given a certain set of control parameters. Furthermore, the SBR encoder works closely with the base core codec to assess which frequency range should be covered by SBR at a given time. In the case of a stereo signal, the SBR data is efficiently encoded by exploiting the channel dependency of the control data as well as entropy coding before transmission.
[0027] It is generally required to carefully tune the control parameter extraction algorithm according to the base codec at a given bit rate and a given sampling rate. This is due to the fact that a lower bit rate generally means a larger SBR range than a high bit rate and different sampling rates correspond to different time resolutions of the SBR frame.
[0028] An SBR decoder typically comprises several different parts. It includes a bitstream decoding module, a high frequency reconstruction (HFR) module, an additional high frequency component module, and an envelope adjuster module. The system is based on a complex-valued QMF filterbank (for high quality SBR) or a real-valued QMF filterbank (for low power SBR). Embodiments of the invention are applicable to both high quality SBR and low power SBR. In the bitstream extraction module, control data is read and decoded from the bitstream. Before envelope data is read from the bitstream, a time-frequency grid of the current frame is obtained. A basic core decoder decodes the audio signal of the current frame (albeit at a lower sampling rate) to produce time-domain audio samples. The resulting frame of audio data is used by the HFR module for high frequency reconstruction. Then, the decoded low frequency band signal is analyzed using a QMF filterbank. Subsequently, high frequency reconstruction and envelope adjustment is performed on the subband samples of the QMF filterbank. Based on given control parameters, the high frequencies are reconstructed from the low frequencies in a flexible way. Furthermore, based on the control data, the reconstructed high frequency band is adaptively filtered from the subband channels to ensure proper spectral characteristics for a given time / frequency region.
[0029] The top level of an MPEG-4 AAC bitstream is a sequence of data blocks ("raw_data_block" elements), each of which is a data section (referred to herein as a "block") containing audio data (typically over a period of 1024 or 960 samples) and related information and / or other data. In this document, we use the term "block" to mean a section of an MPEG-4 AAC bitstream that includes audio data (and corresponding metadata and optionally other related data), which determines or indicates one (but not more than one) "raw_data_block" element.
[0030] Each block of an MPEG-4 AAC bitstream can contain several syntax elements (each of which is also materialized as a data section in the bitstream). Seven types of these syntax elements are defined in the MPEG-4 AAC standard. Each syntax element is identified by a different value of the data element "id_syn_ele". Examples of syntax elements include "single_channel_element()", "channel_pair_element()", and "fill_element()". A single channel element is a container for audio data of a single audio channel (a mono audio signal). A channel pair element contains audio data of two audio channels (i.e., a stereo audio signal).
[0031] A fill element is an information container that contains an identifier (e.g., a value of the element "id_syn_ele" described above) and is followed by data, which is referred to as "fill data". Fill elements have historically been used to adjust the instantaneous bit rate of a bitstream that is to be transmitted over a constant rate channel. Constant data rate can be achieved by adding an appropriate amount of fill data to each block.
[0032] According to embodiments of the application, the padding data can include one or more extension payloads that extend the types of data (e.g., metadata) that can be transmitted in a bitstream. A decoder that receives a bitstream with padding data containing new data types can optionally be used by a device (e.g., a decoder) that receives the bitstream to extend the functionality of the device. Thus, those skilled in the art will appreciate that the padding element is a special type of data structure and is different from the data structures typically used to transmit audio data (e.g., an audio payload containing channel data).
[0033] In some embodiments of the application, the identifier used to identify the padding element can consist of a 3-bit unsigned integer ("uimsbf") with the most significant bit transmitted first having a value of 0x6. In one block, several instances of the same type of syntax element (e.g., several padding elements) can occur.
[0034] Another standard for encoding audio bitstreams is the MPEG Unified Speech and Audio Coding (USAC) standard (ISO / IEC 23003-3:2012). The MPEG USAC standard describes encoding and decoding audio content using a spectral band replication process that includes the SBR process described in the MPEG-4 AAC standard and also includes other enhanced forms of the spectral band replication process. This process applies a spectral band replication tool that is an extended and enhanced version of the SBR toolset described in the MPEG-4 AAC standard (sometimes referred to herein as an "enhanced SBR tool" or "eSBR tool"). Thus, eSBR (as defined in the USAC standard) is an improvement over SBR (as defined in the MPEG-4 AAC standard).
[0035] In this document, we use the expression "enhanced SBR process" (or "eSBR process") to mean a spectral band replication process that uses at least one eSBR tool that is not described or mentioned in the MPEG-4 AAC standard (e.g., at least one eSBR tool described or mentioned in the MPEG USAC standard). Examples of these eSBR tools are harmonic transposition and QMF patching pre-processing or "pre-flattening".
[0036] A harmonic transposer of integer order T maps a sinusoid with frequency ω to a sinusoid with frequency Tω, while preserving signal duration. Three orders T = 2, 3, 4 are typically used in sequence to produce each part of the desired output frequency range using the smallest possible transposition order. If outputs above the 4th order transposition range are needed, they can be produced by frequency shifting. The baseband time domain is used for processing whenever possible to produce near-critical samples to minimize computational complexity.
[0037] Harmonic transposers can be based on QMF or DFT. When using a QMF-based harmonic transposer, the bandwidth extension of the core encoder time domain signal is fully implemented in the QMF domain using a modified phase vocoder structure to perform decimation and then time stretching for each QMF subband. Transposition using several transposition factors (e.g., T = 2, 3, 4) is implemented in the common QMF analysis / synthesis transform stage. Since the QMF-based harmonic transposer does not have the feature of signal adaptive frequency domain oversampling, the corresponding flag (sbrOversamplingFlag[ch]) in the bitstream can be ignored.
[0038] When using a DFT-based harmonic transposer, factor 3 and 4 transposers (3rd and 4th order transposers) are preferably integrated by interpolation into the factor 2 transposer (2nd order transposer) to reduce complexity. For each frame (corresponding to coreCoderFrameLength core encoder samples), the nominal "full size" transform size of the transposer is first determined by the signal adaptive frequency domain oversampling flag (sbrOversamplingFlag[ch]) in the bitstream.
[0039] When sbrPatchingMode == 1 to indicate that linear transposition will be used to generate the high frequency band, an additional step can be introduced to avoid shape discontinuities of the spectral envelope of the high frequency signal to be input to the subsequent envelope adjuster. This improves the operation of the subsequent envelope adjustment stage to result in a high frequency band signal that is perceived as more stable. The operation of the additional pre-processing is beneficial for signal types where the coarse spectral envelope of the low frequency band signal used for high frequency reconstruction shows large levels of variation. However, the value of the bitstream element can be determined in the encoder by applying any kind of signal dependent classification. Preferably, the additional pre-processing is activated by a 1-bit bitstream element bs_sbr_preprocessing. When bs_sbr_preprocessing is set to 1, the additional processing is enabled. When bs_sbr_preprocessing is set to 0, the additional pre-processing is disabled. The additional processing preferably scales each patched low frequency band X Low by the pre-gain curve used by the high frequency generator. For example, the pre-gain curve can be computed according to the following equation:
[0040] preGain(k) = 10(meanNrg - lowEnvSlope(k) / 20, 0 < k < k0
[0041] where k0 is the first QMF subband in the primary band table and lowEnvSlope is computed using a function (e.g., polyfit()) that computes the coefficients of the best fitting polynomial (in the least squares sense). For example, a cubic polynomial (using a 3rd order polynomial) can be employed
[0042] polyfit(3, k0, x_lowband, lowEnv, lowEnvSlope);
[0043] and wherein
[0044]
[0045] where x_lowband(k) = [0...k0-1], numTimeSlot is the number of SBR envelope time slots present in a frame, RATE is a constant (e.g. 2) indicating the number of QMF subband samples per time slot, are linear prediction filter coefficients (obtainable from the covariance method) and wherein
[0046]
[0047] A bitstream generated according to the MPEG USAC standard (sometimes referred to herein as a "USAC bitstream") contains encoded audio content and typically contains metadata indicating each type of spectral band replication processing applied by a decoder to decode the audio content of the USAC bitstream and / or controlling such spectral band replication processing and / or metadata indicating at least one characteristic or parameter of at least one SBR tool and / or eSBR tool used to decode the audio content of the USAC bitstream.
[0048] In this document, we use the expression "enhanced SBR metadata" (or "eSBR metadata") to denote metadata indicating each type of spectral band replication processing applied by a decoder to decode the audio content of an encoded audio bitstream (e.g. a USAC bitstream) and / or controlling such spectral band replication processing and / or indicating at least one characteristic or parameter of at least one SBR tool and / or eSBR tool used to decode such audio content that is not described or mentioned in the MPEG-4 AAC standard. An example of eSBR metadata is metadata (indicating or used to control spectral band replication processing) that is described or mentioned in the MPEG USAC standard but not in the MPEG-4 AAC standard. Thus, eSBR metadata in this document denotes metadata that is not SBR metadata and SBR metadata in this document denotes metadata that is not eSBR metadata.
[0049] A USAC bitstream can contain both SBR metadata and eSBR metadata. More specifically, a USAC bitstream can contain eSBR metadata controlling the eSBR processing performed by a decoder and SBR metadata controlling the SBR processing performed by a decoder. According to typical embodiments of the invention, eSBR metadata (e.g. eSBR specific configuration data) is contained in an MPEG-4 AAC bitstream (e.g. in an sbr_extension() container at the end of the SBR payload) according to the invention.
[0050] During decoding of an encoded bitstream using an eSBR toolset (including at least one eSBR tool), eSBR processing is performed by the decoder to regenerate a high frequency band of an audio signal based on a copy of a harmonic sequence that was truncated during encoding. This eSBR processing typically adjusts the spectral envelope of the generated high frequency band and applies inverse filtering, and adds noise and sinusoidal components to reproduce spectral characteristics of the original audio signal.
[0051] According to typical embodiments of the present invention, eSBR metadata (e.g., a small number of control bits that are eSBR metadata) is included in one or more metadata sections of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream), which also includes encoded audio data in other sections (audio data sections). Typically, at least one such metadata section of each block of the bitstream is (or includes) a stuff element (including an identifier indicating the start of the stuff element), and the eSBR metadata is included in the stuff element following the identifier.
[0052] Figure 1 is a block diagram of an exemplary audio processing chain (audio data processing system), in which one or more elements of the system can be configured according to embodiments of the present invention. The system includes the following elements coupled together as shown: an encoder 1, a transport subsystem 2, a decoder 3, and a post-processing unit 4. In variations of the system shown, one or more elements are omitted, or additional audio data processing units are included.
[0053] In some implementations, the encoder 1 (which optionally includes a pre-processing unit) is configured to accept PCM (time-domain) samples comprising audio content as input and to output an encoded audio bitstream (having a format conforming to the MPEG-4 AAC standard) indicative of the audio content. The data of the bitstream indicative of the audio content is sometimes referred to herein as "audio data" or "encoded audio data". If the encoder is configured according to typical embodiments of the present invention, the audio bitstream output from the encoder includes eSBR metadata (and typically also other metadata) as well as the audio data.
[0054] One or more encoded audio bitstreams output from the encoder 1 can be asserted to an encoded audio transport subsystem 2. The subsystem 2 is configured to store and / or transport each encoded bitstream output from the encoder 1. The encoded audio bitstream output from the encoder 1 can be stored by the subsystem 2 (e.g., in the form of a DVD or Blu-ray disc), or transmitted by the subsystem 2 (which can implement a transmission link or network), or can be stored and transmitted by the subsystem 2.
[0055] Decoder 3 is configured to decode the encoded MPEG-4 AAC audio bitstream (produced by encoder 1) that it receives via subsystem 2. In some embodiments, decoder 3 is configured to extract eSBR metadata from each block of the bitstream and decode the bitstream (including by performing eSBR processing using the extracted eSBR metadata) to produce decoded audio data (e.g., a stream of decoded PCM audio samples). In some embodiments, decoder 3 is configured to extract SBR metadata from the bitstream (but ignore the eSBR metadata included in the bitstream) and decode the bitstream (including by performing SBR processing using the extracted SBR metadata) to produce decoded audio data (e.g., a stream of decoded PCM audio samples). Typically, decoder 3 includes a buffer that stores (e.g., in a non-transitory manner) segments of the encoded audio bitstream received from subsystem 2.
[0056] Figure 1 Post-processing unit 4 of FIG. 1 is configured to accept the stream of decoded audio data (e.g., decoded PCM audio samples) from decoder 3 and perform post-processing on it. Post-processing unit can also be configured to render the post-processed audio content (or decoded audio received from decoder 3) for playback by one or more loudspeakers.
[0057] Figure 2 is a block diagram of an encoder 100, which is an embodiment of the inventive audio processing unit. Any component or element of encoder 100 can be implemented in hardware, software, or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASICs, FPGAs, or other integrated circuits). Encoder 100 includes an encoder 105, a filler / formatting stage 107, a metadata generation stage 106, and a buffer memory 109, as shown connected. Typically, encoder 100 also includes other processing elements (not shown). Encoder 100 is configured to convert an input audio bitstream into an encoded output MPEG-4 AAC bitstream.
[0058] Metadata generator 106 is coupled and configured to generate metadata (including eSBR metadata and SBR metadata) (and / or pass to stage 107) for inclusion in the encoded bitstream output from encoder 100 by stage 107.
[0059] Encoder 105 is coupled and configured to encode input audio data (e.g., by performing compression on it) and assert the resulting encoded audio to stage 107 for inclusion in the encoded bitstream output from stage 107.
[0060] Stage 107 is configured to multiplex the encoded audio from encoder 105 and the metadata (including eSBR metadata and SBR metadata) from generator 106 to produce an encoded bitstream output from stage 107, preferably such that the encoded bitstream has a format specified by one embodiment of the invention.
[0061] Buffer memory 109 is configured to store (e.g., in a non-transitory manner) at least one block of the encoded audio bitstream output from stage 107, and then assert a sequence of blocks of the encoded audio bitstream from buffer memory 109 as output from encoder 100 to a transmission system.
[0062] Figure 3 is a block diagram of a system that includes a decoder 200, which is an embodiment of an inventive audio processing unit, and optionally also includes a post-processor 300 coupled to decoder 200. Any component or element of decoder 200 and post-processor 300 can be implemented in hardware, software, or a combination of hardware and software, as one or more processes and / or one or more circuits (e.g., ASICs, FPGAs, or other integrated circuits). Decoder 200 includes a buffer memory 201, a bitstream payload deformatter (parser) 205, an audio decoding subsystem 202 (sometimes referred to as a "core" decoding stage or "core" decoding subsystem), an eSBR processing stage 203, and a control bit generation stage 204, connected as shown. In general, decoder 200 also includes other processing elements (not shown).
[0063] Buffer memory (buffer) 201 stores (e.g., in a non-transitory manner) at least one block of an encoded MPEG-4 AAC audio bitstream received by decoder 200. In operation of decoder 200, a sequence of blocks of the bitstream is asserted from buffer 201 to deformatter 205.
[0064] In Figure 3 embodiments (or variants of embodiments to be described Figure 4 An APU (which is not a decoder), such as APU 500 of Figure 6 includes a buffer memory (e.g., a buffer memory identical to buffer 201) that stores (e.g., in a non-transitory manner) at least one block of an encoded audio bitstream (e.g., an MPEG-4 AAC audio bitstream) of the same type received by buffer 201 of Figure 3 or Figure 4 decoder 200.
[0065] Referring again to Figure 3The deformatter 205 is coupled and configured to demultiplex each block of the bitstream to extract therefrom SBR metadata (including quantized envelope data) and eSBR metadata (and typically also other metadata) to assert at least the eSBR metadata and the SBR metadata to the eSBR processing stage 203 and typically also other extracted metadata to the decoding subsystem 202 (and optionally also to the control bit generator 204). The deformatter 205 is also coupled and configured to extract audio data from each block of the bitstream and to assert the extracted audio data to the decoding subsystem (decoding stage) 202.
[0066] Figure 3 The system of FIG. 2 also optionally includes a post-processor 300. The post-processor 300 includes a buffer memory (buffer) 301 and other processing elements (not shown) including at least one processing element coupled to the buffer 301. The buffer 301 stores (e.g., in a non-transient manner) at least one block (or frame) of decoded audio data received by the post-processor 300 from the decoder 200. The processing elements of the post-processor 300 are coupled and configured to receive and use metadata output from the decoding subsystem 202 (and / or the deformatter 205) and / or control bits output from the stage 204 of the decoder 200 to adaptively process a sequence of blocks (or frames) of decoded audio output from the buffer 301.
[0067] The audio decoding subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the parser 205 (this decoding can be referred to as a "core" decoding operation) to generate decoded audio data and assert the decoded audio data to the eSBR processing stage 203. The decoding is performed in the frequency domain and typically includes inverse quantization and then spectral processing. Typically, the last processing stage in the subsystem 202 applies a frequency domain to time domain transform to the decoded frequency domain audio data so that the output of the subsystem is time domain decoded audio data. The stage 203 is configured to apply the SBR tool and the eSBR tool indicated by the eSBR metadata and the eSBR (extracted by the parser 205) to the decoded audio data (i.e., perform SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata) to generate fully decoded audio data that is output from the decoder 200 (e.g., to the post-processor 300). Typically, the decoder 200 includes a memory that stores the deformatted audio data and metadata output from the deformatter 205 (accessible by the subsystem 202 and the stage 203), and the stage 203 is configured to access the audio data and metadata (including the SBR metadata and the eSBR metadata) as needed during the SBR and eSBR processing. The SBR processing and the eSBR processing in the stage 203 can be viewed as post-processing of the output of the core decoding subsystem 202. The decoder 200 also optionally includes a last upmixing subsystem (which can use PS metadata extracted by the deformatter 205 and / or control bits generated in the subsystem 204 to apply the parametric stereo ("PS") tool defined in the MPEG-4 AAC standard) that is coupled and configured to perform upmixing on the output of the stage 203 to generate fully decoded upmixed audio that is output from the decoder 200. Alternatively, the post-processor 300 is configured to perform upmixing on the output of the decoder 200 (e.g., using PS metadata extracted by the deformatter 205 and / or control bits generated in the subsystem 204).
[0068] In response to the metadata extracted by the deformatter 205, the control bit generator 204 can generate control data, and the control data can be used within the decoder 200 (e.g., in the last upmixing subsystem) and / or asserted as output from the decoder 200 (e.g., to the post-processor 300 for post-processing). In response to the metadata extracted from the input bitstream (and optionally also in response to the control data), the stage 204 can generate control bits (and assert the control bits to the post-processor 300) to indicate that the decoded audio data output from the eSBR processing stage 203 should undergo a particular type of post-processing. In some implementations, the decoder 200 is configured to assert the metadata extracted by the deformatter 205 from the input bitstream to the post-processor 300, and the post-processor 300 is configured to perform post-processing on the decoded audio data output from the decoder 200 using the metadata.
[0069] Figure 4 This is a block diagram of an audio processing unit (“APU”) 210, which is another embodiment of the inventive audio processing unit. APU 210 is a conventional decoder not configured to perform eSBR processing. Any component or element of APU 210 may be implemented in hardware, software, or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuits). APU 210 includes a buffer memory 201, a bitstream payload deformatter (parser) 215, an audio decoding subsystem 202 (sometimes referred to as the “core” decoding stage or “core” decoding subsystem), and an SBR processing stage 213, connected as shown. Typically, APU 210 also includes other processing elements (not shown). APU 210 may represent, for example, an audio encoder, decoder, or transcoder.
[0070] Components 201 and 202 of APU 210 are the same as ( Figure 3 The same numbered elements of the decoder 200 will not be repeated in their above description. In the operation of the APU 210, the block sequence of the encoded audio bitstream (MPEG-4 AAC bitstream) received by the APU 210 is asserted from the buffer 201 to the deformatter 215.
[0071] Deformatter 215 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (including quantization envelope data) from it, and typically also extracts other metadata from it, but ignores eSBR metadata that may be included in the bitstream according to any embodiment of the invention. Deformatter 215 is configured to assert at least SBR metadata to SBR processing stage 213. Deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to decoding subsystem (decoding stage) 202.
[0072] The audio decoding subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the deformatter 215 (this decoding can be referred to as a "core" decoding operation) to generate decoded audio data and assert the decoded audio data to the SBR processing stage 213. The decoding is performed in the frequency domain. Typically, the last processing stage in the subsystem 202 applies a frequency-to-time domain transform to the decoded frequency domain audio data, so that the output of the subsystem is time domain decoded audio data. The stage 213 is configured to apply the SBR tool (but not the eSBR tool) indicated by the SBR metadata (extracted by the deformatter 215) to the decoded audio data (i.e., using the SBR metadata to perform SBR processing on the output of the decoding subsystem 202) to generate fully decoded audio data that is output from the APU 210 (e.g., to the post-processor 300). Typically, the APU 210 includes a memory that stores the deformatted audio data and metadata output from the deformatter 215 (accessible by the subsystem 202 and stage 213), and the stage 213 is configured to access the audio data and metadata (including the SBR metadata) as needed during the SBR processing. The SBR processing in the stage 213 can be viewed as post-processing of the output of the core decoding subsystem 202. The APU 210 also optionally includes a last upmixing subsystem (which can use the PS metadata extracted by the deformatter 215 to apply the parametric stereo "PS" tool defined in the MPEG-4 AAC standard) that is coupled and configured to perform upmixing on the output of the stage 213 to generate fully decoded upmixed audio that is output from the APU 210. Alternatively, the post-processor is configured to perform upmixing on the output of the APU 210 (e.g., using the PS metadata extracted by the deformatter 215 and / or control bits generated in the APU 210).
[0073] Various implementations of the encoder 100, decoder 200, and APU 210 are configured to perform different embodiments of the inventive method.
[0074] According to some embodiments, eSBR metadata (e.g., a small number of control bits that are eSBR metadata) is included in an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) so that a legacy decoder (that is not configured to parse eSBR metadata or use any eSBR tools related to the eSBR metadata) can ignore the eSBR metadata but still decode the bitstream as much as possible without using the eSBR metadata or any eSBR tools related to the eSBR metadata, typically without a significant loss in decoded audio quality. However, an eSBR decoder (that is configured to parse the bitstream to identify the eSBR metadata and use at least one eSBR tool in response to the eSBR metadata) will benefit from using at least one such eSBR tool. Thus, embodiments of the invention provide a method for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner.
[0075] Typically, the eSBR metadata in the bitstream indicates one or more of the following eSBR tools (e.g., at least one property or parameter indicative thereof) (the eSBR tools are described in the MPEG USAC standard, and can or can not be applied by the encoder during generation of the bitstream):
[0076] • Harmonic transposition; and
[0077] • QMF patching extra pre-processing (pre-flattening).
[0078] For example, the eSBR metadata included in the bitstream can indicate values of the parameters (as described in the MPEG USAC standard and in the present invention): sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch] and bs_sbr_preprocessing.
[0079] In this document, the notation X[ch] (where X is some parameter) indicates that the parameter pertains to a channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, we sometimes omit the expression [ch] and assume that the relevant parameter pertains to a channel of the audio content.
[0080] In this document, the notation X[ch][env] (where X is some parameter) indicates that the parameter pertains to an SBR envelope ("env") of a channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, we sometimes omit the expressions [env] and [ch] and assume that the relevant parameter pertains to an SBR envelope of a channel of the audio content.
[0081] During decoding of the encoded bitstream, the harmonic transposition (for each channel "ch" of the audio content indicated by the bitstream) is performed during the decoded eSBR processing stage is controlled by the following eSBR metadata parameters: sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch].
[0082] The value "sbrPatchingMode[ch]" indicates the type of transposer used in eSBR: sbrPatchingMode[ch] = 1 indicates the linear transposition patching described in section 4.6.18 of the MPEG-4 AAC standard (used with high-quality SBR or low-power SBR); sbrPatchingMode[ch] = 0 indicates the harmonic SBR patching described in section 7.5.3 or 7.5.4 of the MPEG USAC standard.
[0083] The value "sbrOversamplingFlag[ch]" indicates the combined use of signal adaptive frequency domain oversampling in eSBR with the DFT-based harmonic SBR patching described in section 7.5.3 of the MPEG USAC standard. This flag controls the size of the DFT used in the transposer: 1 indicates that signal adaptive frequency domain oversampling is enabled as described in section 7.5.3.1 of the MPEG USAC standard; 0 indicates that signal adaptive frequency domain oversampling is disabled as described in section 7.5.3.1 of the MPEG USAC standard.
[0084] The value "sbrPitchInBinsFlag[ch]" controls the interpretation of the sbrPitchInBins[ch] parameter: 1 indicates that the value of sbrPitchInBins[ch] is valid and greater than 0; 0 indicates that the value of sbrPitchInBins[ch] is set to 0.
[0085] The value "sbrPitchInBins[ch]" controls the addition of cross-product terms in the SBR harmonic transposer. The value sbrPitchinBins[ch] is an integer value in the range [0, 127] and represents the distance measured in frequency bins of a 1536-line DFT operating on the sampling frequency of the core encoder.
[0086] If the MPEG-4 AAC bitstream indicates a coupled SBR channel pair (instead of a single SBR channel), the bitstream indicates two instances of the above syntax (for harmonic or non-harmonic transposition), one for each channel of the channel pair, sbr_channel_pair_element().
[0087] The harmonic transposition of the eSBR tool generally improves the quality of the decoded music signal at relatively low cross-over frequencies. The non-harmonic transposition (i.e., the conventional spectral patching) generally improves speech signals. Thus, the starting point for deciding which type of transposition is preferred for encoding a particular audio content is to select the transposition method depending on a speech / music detection, where harmonic transposition is employed for music content and spectral patching is employed for speech content.
[0088] Performing pre-flattening during eSBR processing is controlled by the value of a 1-bit eSBR metadata parameter called "bs_sbr_preprocessing", in the sense that pre-flattening is performed or not depending on the value of this single bit. When the SBR QMF repair algorithm described in section 4.6.18.6.3 of the MPEG-4 AAC standard is used, the steps of pre-flattening (when indicated by the "bs_sbr_preprocessing" parameter) can be performed in an attempt to avoid shape discontinuities of the spectral envelope of the high frequency signal from entering the subsequent envelope adjuster (which performs another stage of the eSBR processing). Pre-flattening generally improves the operation of the subsequent envelope adjust stage, resulting in a high frequency band signal that is perceived as more stable.
[0089] According to some embodiments of the present application, the total bit rate requirement for inclusion in the MPEG-4 AAC bitstream eSBR metadata that indicates the above-mentioned eSBR tools (harmonic transposition and pre-flattening) is expected to be on the order of a few hundred bits per second, since only differential control data required to perform the eSBR processing is transmitted. Conventional decoders can ignore this information, since it is included in a backward compatible manner (as will be explained later). Thus, the adverse impact on bit rate associated with including the eSBR metadata can be negligible for several reasons, including:
[0090] • The bit rate loss (due to the inclusion of the eSBR metadata) is very small in terms of the total bit rate, since only differential control data required to perform the eSBR processing (and not a simulcast of the SBR control data) is transmitted; and
[0091] • The tuning of the SBR related control information generally does not depend on the details of the transposition. Examples of control data that depend on the operation of the transposer will be discussed later in this application.
[0092] Thus, embodiments of the present application provide a method for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward compatible manner. This efficient transmission of eSBR control data reduces the memory requirements in decoders, encoders and transcoders that employ aspects of the present application, while having no significant adverse impact on bit rate. Furthermore, the complexity and processing requirements associated with performing eSBR according to embodiments of the present application are also reduced, since the SBR data only needs to be processed once and not simulcasted, which is the case when eSBR is treated as a completely independent object type in MPEG-4 AAC, rather than integrated into the MPEG-4 AAC codec in a backward compatible manner.
[0093] Next, with reference to Figure 7 we describe the elements of a block ("raw_data_block") of an MPEG-4 AAC bitstream (in which eSBR metadata is included) according to some embodiments of the present application.Figure 7 is a diagram of a block ("raw_data_block") of an MPEG-4 AAC bitstream, which shows some sections of the MPEG-4 AAC bitstream.
[0094] A block of an MPEG-4 AAC bitstream can include at least one "single_channel_element()" (e.g., the single channel element shown in Figure 7 ) and / or at least one "channel_pair_element()" (e.g., the channel pair element shown in Figure 7 ), which includes audio data of an audio program. The block can also include several "fill_element" (e.g., fill element 1 and / or fill element 2 of Figure 7 ), which includes data (e.g., metadata) related to the program. Each "single_channel_element()" includes an identifier (e.g., "ID1" of Figure 7 ) indicating the beginning of a single channel element, and can include audio data indicating different channels of a multi-channel audio program. Each "channel_pair_element()" includes an identifier (not shown in Figure 7 ) indicating the beginning of a channel pair element, and can include audio data indicating two channels of a program.
[0095] A fill_element of an MPEG-4 AAC bitstream (referred to herein as a fill element) includes an identifier (e.g., "ID2" of Figure 7 ) indicating the beginning of the fill element and fill data following the identifier. The identifier ID2 can consist of a 3-bit unsigned integer ("uimsbf") with the most significant bit first having a value of 0x6. The fill data can include an extension_payload() element (referred to herein sometimes as an extension payload) whose syntax is shown in Table 4.57 of the MPEG-4 AAC standard. There are several types of extension payloads and are identified by an "extension_type" parameter, which is a 4-bit unsigned integer ("uimsbf") with the most significant bit first.
[0096] The fill data (e.g., the extension payload of Figure 7"SBR Object" type referred to in the MPEG-4 AAC standard as sbr_extension_data()). For example, the value of "1101" or "1110" for the extension_type field in the header is used to identify the Spectral Band Replication (SBR) extension payload, where the identifier "1101" identifies an extension payload with SBR data and "1110" identifies an extension payload containing SBR data with a Cyclic Redundancy Check (CRC) to verify the correctness of the SBR data.
[0097] When the header (e.g., the extension_type field) initializes the SBR Object type, SBR metadata (sometimes referred to herein as "spectral band replication data," and referred to in the MPEG-4 AAC standard as sbr_data()) follows the header, and at least one spectral band replication extension element (e.g., a "SBR extension element" of the fill element 1 of Figure 7 "SBR extension header") of the fill element 1 of Figure 7 "SBR extension header") of the fill element 1 of
[0098] The MPEG-4 AAC standard contemplates that the spectral band replication extension element can contain PS (Parametric Stereo) data for the audio data of the program. The MPEG-4 AAC standard contemplates that when the header of the fill element (e.g., the extension payload thereof) initializes the SBR Object type (as in "Header 1" of Figure 7 "SBR extension header") of the fill element 1 of
[0099] According to some embodiments of the present application, eSBR metadata (e.g., a flag indicating whether enhanced spectral band replication (eSBR) processing is performed on the audio content of the block) is included in the spectral band replication extension element of the fill element. For example, this flag is indicated in the fill element 1 of Figure 7 "SBR extension header") of the fill element 1 of Figure 7(In the SBR extension element of padding element 1 in the present invention). According to some embodiments of the present invention, the padding element containing eSBR metadata also includes a "bs_extension_id" parameter, the value of which (e.g., bs_extension_id = 3) indicates that the eSBR metadata is included in the padding element and that eSBR processing is performed on the audio content of the relevant block.
[0100] According to some embodiments of the present invention, eSBR metadata is included in the padding elements of the MPEG-4 AAC bitstream (e.g. Figure 7 The padding element 2) is in the spectral band copy extension element (SBR extension element) that is not a padding element. This is because the padding element containing extension_payload() (which has SBR data or SBR data with CRC) does not contain any other extension payload of any other extension type. Therefore, in embodiments where the eSBR metadata stores its own extension payload, a separate padding element is used to store the eSBR metadata. This padding element contains an identifier indicating the start of the padding element (e.g., ...). Figure 7 The padding data follows the identifier “ID2”. The padding data may contain the extension_payload() element (sometimes referred to herein as the extended payload) whose syntax is shown in Table 4.57 of the MPEG-4 AAC standard. The padding data (e.g., its extended payload) contains a header indicating the eSBR object (e.g., ...). Figure 7 The padding element 2 is “Header 2” (i.e., the header initializes the Enhanced Spectrum Band Replication (eSBR) object type), and the padding data (e.g., its extended payload) contains the eSBR metadata following the header. For example, Figure 7 Padding element 2 contains this header (“Header 2”) and also contains eSBR metadata following the header (i.e., the “tags” in padding element 2 that indicate whether Enhanced Spectral Band Replication (eSBR) processing is performed on the audio content of the block). Additional eSBR metadata may also be included after Header 2. Figure 3 The padding data in padding element 2. In the embodiments described in this paragraph, the header (e.g. Figure 3 The header 2) has an identification value that is not the regular value specified in Table 4.57 of the MPEG-4 AAC standard, but instead indicates the eSBR extension payload (so that the extension_type field of the header indicates that the padding data contains eSBR metadata).
[0101] In a first-type embodiment, the present invention is an audio processing unit (e.g., a decoder) comprising:
[0102] memory (e.g.) Figure 4Or a buffer 201 of 4, which is configured to store at least one block of an encoded audio bitstream (e.g., at least one block of an MPEG-4 AAC bitstream);
[0103] Bitstream payload deformatter (e.g.) Figure 3 Component 205 or Figure 4 Element 215), which is coupled to the memory and configured to demultiplex at least a portion of the block of the bit stream; and
[0104] Decoding subsystem (e.g.) Figure 4 Components 202 and 203 or bs_extension_id Elements 202 and 213), coupled and configured to decode at least a portion of the audio content of the block of the bitstream, wherein the block comprises:
[0105] A padding element comprising an identifier indicating the start of the padding element (e.g., an "id_syn_ele" identifier with value 0x6 from Table 4.85 of the MPEG-4 AAC standard) and padding data following the identifier, wherein the padding data comprises:
[0106] At least one flag identifies whether enhanced spectrum band copying (eSBR) processing is performed on the audio content of the block (e.g., using spectrum band copying data and eSBR metadata contained in the block).
[0107] The tag is eSBR metadata, and an instance of the tag is the sbrPatchingMode tag. Another instance of the tag is the harmonicSBR tag. Both of these tags indicate whether the basic form of spectral band copying or an enhanced form of spectral copying is performed on the audio data of the block. The basic form of spectral copying is spectral patching, and the enhanced form of spectral band copying is harmonic transpose.
[0108] In some embodiments, the padding data also includes additional eSBR metadata (i.e., eSBR metadata other than the tags).
[0109] The memory may be a buffer memory (e.g.) Meaning An embodiment of buffer 201 stores (e.g., in a non-transitory manner) at least one block of the encoded audio bitstream.
[0110] It is estimated that the complexity of performing eSBR processing (using eSBR harmonic transpose and pre-flattening) by the eSBR decoder during the decoding of an MPEG-4 AAC bitstream containing eSBR metadata (indicating these eSBR tools) will be as follows (for typical decoding with indicator parameters):
[0111] • Harmonic transposition (16 kbps, 14400 / 28800 Hz)
[0112] o DFT-based: 3.68 WMOPS (Weighted Million Operations Per Second);
[0113] o QMF-based: 0.98 WMOPS;
[0114] • QMF patch pre-processing (pre-flattening): 0.1 WMOPS.
[0115] It is well known that for transients, DFT-based transposition usually performs better than QMF-based transposition.
[0116] According to some embodiments of the application, a stuffing element (of an encoded audio bitstream) containing eSBR metadata also contains a parameter (e.g. the "bs_extension_id" parameter) and / or a value (e.g. bs_extension_id = 3) of which indicates that the eSBR metadata is contained in the stuffing element and that eSBR processing is performed on the audio content of the associated block, and / or a value (e.g. bs_extension_id = 2) of which indicates that the sbr_extension() container of the stuffing element contains PS data. For example, as indicated in Table 1 below, this parameter with value bs_extension_id = 2 can indicate that the sbr_extension() container of the stuffing element contains PS data, and this parameter with value bs_extension_id = 3 can indicate that the sbr_extension() container of the stuffing element contains eSBR metadata:
[0117] Table 1
[0118] Reserved Reserved 0 EXTENSION_ID_PS 1 EXTENSION_ID_ESBR 2 Figure 5 3 Figure 5
[0119] According to some embodiments of the application, the syntax of each spectral band replication extension element containing eSBR metadata and / or PS data is as indicated in Table 2 below (where "sbr_extension()" denotes a container of a spectral band replication extension element, "bs_extension_id" is as described in Table 1 above, "ps_data" denotes PS data, and "esbr_data" denotes eSBR metadata):
[0120] Table 2
[0121]
[0122]
[0123] In an illustrative embodiment, the esbr_data() referred to in Table 2 above indicates values for the following metadata parameters:
[0124] 1. a 1-bit metadata parameter "bs_sbr_preprocessing"; and
[0125] 2. for each channel ("ch") of the audio content of the encoded bitstream to be decoded, each of the above parameters is "sbrPatchingMode[ch]", "SbrOversamplingFlag[ch]", "SbrPitchInBinsFlag[ch]", and "sbrPitchInBins[ch]".
[0126] For example, in some embodiments, the esbr_data() can have the syntax indicated in Table 3 to indicate these metadata parameters:
[0127] Table 3
[0128]
[0129]
[0130]
[0131] The above syntax enables the efficient implementation of the enhanced form of spectral band replication, such as harmonic transposition, as an extension of a legacy decoder. In particular, the eSBR data of Table 3 contains only the parameters required to perform the enhanced form of spectral band replication, which are not already supported in the bitstream and cannot be directly derived from the already supported parameters in the bitstream. All other parameters and processing data required to perform the enhanced form of spectral band replication are extracted from the ready-made parameters in the already defined positions in the bitstream.
[0132] For example, an MPEG-4 HE-AAC or HE-AAC v2 compliant decoder can be extended to include an enhanced form of spectral band replication, such as harmonic transposition. This enhanced form of spectral band replication is an addition to the basic form of spectral band replication already supported by the decoder. In the context of an MPEG-4 HE-AAC or HE-AAC v2 compliant decoder, this basic form of spectral band replication is the QMF spectral patching SBR tool, as defined in section 4.6.18 of the MPEG-4 AAC standard.
[0133] When performing an enhanced form of spectral band replication, the extended HE-AAC decoder can re-use many bitstream parameters that are already contained in the SBR extension payload of the bitstream. Particular parameters that can be re-used include, for example, various parameters that determine the main band table. These parameters include, for example, bs_start_freq (a parameter that determines the start of the main frequency table parameter), bs_stop_freq (a parameter that determines the stop of the main frequency table), bs_freq_scale (a parameter that determines the number of frequency bands per octave), and bs_alter_scale (a parameter that alters the scale of the frequency bands). Parameters that can be re-used also include parameters that determine the noise band table (bs_noise_bands) and the limiter band table parameter (bs_limiter_bands). Thus, in various embodiments, at least some equivalent parameters specified in the USAC standard are omitted from the bitstream to thereby reduce the control burden of the bitstream. In general, when a parameter specified in the AAC standard has an equivalent parameter specified in the USAC standard, the equivalent parameter specified in the USAC standard has the same name as the parameter specified in the AAC standard, e.g., the envelope scale factor E OrigMapped However, the equivalent parameters specified in the USAC standard generally have different values that are "tuned" according to the enhanced SBR processing defined in the USAC standard rather than the SBR processing defined in the AAC standard.
[0134] The activation of the enhanced SBR is recommended to improve the subjective quality of audio content with harmonic frequency structure and strong tonal characteristics, especially at low bitrates. The values of the corresponding bitstream elements controlling these tools (i.e., esbr_data()) can be determined in the encoder by applying a signal dependent classification mechanism. In general, the use of the harmonic patching method (sbrPatchingMode == 1) is preferred for encoding music signals at very low bitrates, where the audio bandwidth of the core codec can be severely limited. This is particularly true when these signals contain a significant harmonic structure. Conversely, the use of the regular SBR patching method is preferred for speech and mixed signals, as it provides a better preservation of the temporal structure of speech.
[0135] To improve the performance of the harmonic transposer, a pre-processing step (bs_sbr_preprocessing == 1) can be activated that tries to avoid introducing spectral discontinuities of the signal to the subsequent envelope adjuster. The operation of the tool is beneficial for signal types where the coarse spectral envelope of the low band signal used for high frequency reconstruction shows a large level of variation.
[0136] To improve the transient response of the harmonic SBR patching, signal adaptive frequency domain oversampling (sbrOversamplingFlag == 1) can be applied. Since signal adaptive frequency domain oversampling increases the computational complexity of the transposer and only brings benefits for frames containing transients, the use of this tool is controlled by a bitstream element, which is transmitted once per frame and per independent SBR channel.
[0137] A decoder operating in the proposed enhanced SBR mode typically needs to be able to switch between legacy SBR patching and enhanced SBR patching. Therefore, a delay which can be as long as the duration of one core audio frame can be introduced depending on the decoder setup. Typically, the delay for both legacy SBR patching and enhanced SBR patching will be similar.
[0138] In addition to many parameters, other data elements can also be re-used by an extended HE-AAC decoder when performing the enhanced form of spectral band replication according to embodiments of the present application. For example, envelope data and noise floor data can also be extracted from the bs_data_env (envelope scale factors) and bs_noise_env (noise floor scale factors) data and used during the enhanced form of spectral band replication.
[0139] In essence, these embodiments exploit the configuration parameters and envelope data already supported by legacy HE-AAC or HE-AAC v2 decoders in the SBR extension payload to enable the enhanced form of spectral band replication, which requires as little additional transmitted data as possible. The metadata is initially tuned for the basic form of HFR, such as the spectral translation operation of SBR, but according to embodiments, for the enhanced form of HFR, such as the harmonic transposition of eSBR. As discussed previously, the metadata generally represents the operational parameters (e.g., envelope scale factors, noise floor scale factors, time / frequency grid parameters, sinusoidal addition information, variable crossover frequency / bands, inverse filtering mode, envelope resolution, smoothing mode, frequency interpolation mode) that are tuned and designed for use with the basic form of HFR, such as linear spectral translation. However, this metadata can be used in combination with additional metadata parameters that are specific to the enhanced form of HFR, such as harmonic transposition, to efficiently and effectively process audio data using the enhanced form of HFR.
[0140] Therefore, an extended decoder supporting spectral band replication can be generated very efficiently by relying on predefined bitstream elements (e.g., bitstream elements in the SBR extended payload) and adding only the parameters required to support the enhanced form (in the padding element extended payload). This data reduction feature, combined with placing the newly added parameters in reserved data fields (e.g., extended containers), largely reduces the barriers to generating a decoder that supports the enhanced form by ensuring backward compatibility of the bitstream with conventional decoders that do not support spectral band replication.
[0141] In Table 3, the numbers in the right row indicate the number of bits in the corresponding parameter in the left row.
[0142] In some embodiments, the SBR object type defined in MPEG-4 AAC is updated to include aspects of the SBR tool and the enhanced SBR (eSBR) tool, as indicated in the SBR extension element (bs_extension_id == EXTENSION_ID_ESBR). If the decoder detects and supports this SBR extension element, then the decoder adopts the indicated aspect of the enhanced SBR tool. The SBR object type updated in this manner is referred to as SBR enhancement.
[0143] In some embodiments, the present invention is a method comprising the steps of encoding audio data to produce an encoded bitstream (e.g., an MPEG-4 AAC bitstream), including including eSBR metadata in at least one segment of at least one block of the encoded bitstream and audio data in at least another segment of said block. In a typical embodiment, the method includes the step of multiplexing the audio data and eSBR metadata in each block of the encoded bitstream. In typical decoding of the encoded bitstream in an eSBR decoder, the decoder extracts eSBR metadata from the bitstream (including parsing and demultiplexing the eSBR metadata and audio data) and uses the eSBR metadata to process the audio data to produce a decoded audio data stream.
[0144] Another aspect of the invention is an eSBR decoder configured to perform eSBR processing (e.g., using at least one of eSBR tools called harmonic transpose or pre-flattening) during the decoding of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that does not contain eSBR metadata. (See also...) Figure 3 To describe an instance of this decoder.
[0145] Figure 4 The eSBR decoder 400 includes a buffer memory 201 connected as shown (which is identical to...) Figure 3 and 4 The memory 201), bit stream payload deformatter 215 (which is the same as)Figure 3 deformatter 215), an audio decoding sub-system 202 (sometimes referred to as a "core" decoding stage or "core" decoding sub-system, and identical to Figure 5 core decoding sub-system 202 of the MPEG-4 AAC decoder 100), an eSBR control data generation sub-system 401, and an eSBR processing stage 203 (which is identical to Figure 1 stage 203 of the MPEG-4 AAC decoder 100). In general, the decoder 400 also includes other processing elements (not shown).
[0146] In operation of the decoder 400, a sequence of blocks of an encoded audio bitstream (MPEG-4 AAC bitstream) received by the decoder 400 is asserted from the buffer 201 to the deformatter 215.
[0147] The deformatter 215 is coupled and configured to demultiplex each block of the bitstream to extract therefrom SBR metadata (including quantized envelope data) and, in general, other metadata therefrom as well. The deformatter 215 is configured to assert at least the SBR metadata to the eSBR processing stage 203. The deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding sub-system (decoding stage) 202.
[0148] The audio decoding subsystem 202 of the decoder 400 is configured to decode the audio data extracted by the deformatter 215 (this decoding can be referred to as a "core" decoding operation) to generate decoded audio data and assert the decoded audio data to the eSBR processing stage 203. The decoding is performed in the frequency domain. Typically, the last processing stage in the subsystem 202 applies a frequency domain to time domain transform to the decoded frequency domain audio data, so that the output of the subsystem is time domain decoded audio data. The stage 203 is configured to apply the SBR tool (and the eSBR tool) indicated by the SBR metadata (extracted by the deformatter 215) and the eSBR metadata generated in the subsystem 401 to the decoded audio data (i.e., using the SBR and eSBR metadata to perform SBR and ESBR processing on the output of the decoding subsystem 202) to generate fully decoded audio data that is output from the decoder 400. Typically, the decoder 400 includes a memory (accessible by the subsystem 202 and the stage 203) that stores the deformatted audio data and metadata output from the deformatter 215 (and, optionally, the subsystem 401), and the stage 203 is configured to access the audio data and metadata as needed during the SBR and eSBR processing. The SBR processing in the stage 203 can be viewed as a post-processing of the output of the core decoding subsystem 202. The decoder 400 also optionally includes a final upmixing subsystem (which can apply the parametric stereo "PS" tool defined in the MPEG-4 AAC standard using the PS metadata extracted by the deformatter 215) that is coupled and configured to perform upmixing on the output of the stage 203 to generate fully decoded upmixed audio that is output from the APU 210.
[0149] Parametric stereo is an encoding tool that represents a stereo signal using a linear downmix of the left and right channels of the stereo signal and a set of spatial parameters that describe the stereo image. Parametric stereo typically employs three types of spatial parameters: (1) an inter-channel intensity difference (IID), which describes the intensity difference between the channels; (2) an inter-channel phase difference (IPD), which describes the phase difference between the channels; and (3) an inter-channel coherence (ICC), which describes the coherence (or similarity) between the channels. Coherence can be measured as the maximum of the cross-correlation as a function of time or phase. These three parameters typically enable a high-quality reconstruction of the stereo image. However, the IPD parameter only specifies the relative phase difference between the channels of the stereo input signal and does not indicate the distribution of these phase differences over the left and right channels. Therefore, a fourth type of parameter describing the overall phase offset or overall phase difference (OPD) can additionally be used. In the stereo reconstruction process, consecutive window segments of both the received downmix signal s[n] and a decorrelated version d[n] of the received downmix are processed together with the spatial parameters to generate left (l k (n)) and right (r k (n)) reconstructed signals according to the following equations:
[0150] | k (n) = H 11 (k, n) S k (n) + H 21 (k, n) d k (n)
[0151] r k (n) = H 12 (k, n) S k (n) + H 22 (k, n) d k (n)
[0152] where H 11 , H 12 , H 21 and H 22 are defined by the stereo parameters. Finally, the signals l k (n) and r k (n) are transformed back into the time domain by a frequency-to-time transform.
[0153] Figure 2 The control data generation subsystem 401 is coupled and configured to detect at least one property of the encoded audio bitstream to be decoded and to generate eSBR control data (which can be or include any type of eSBR metadata included in the encoded audio bitstream according to other embodiments of the invention) in response to at least one result of the detection step. The eSBR control data is asserted to stage 203 to trigger the application of individual eSBR tools or combinations of eSBR tools and / or to control the application of these eSBR tools after detection of a particular property (or combination of properties) of the bitstream. For example, to control the execution of the eSBR processing using harmonic transposition, some embodiments of the control data generation subsystem 401 will include: a music detector (e.g. a simplified version of a conventional music detector) for setting the sbrPatchingMode[ch] parameter (and asserting the set parameter to stage 203) in response to detecting whether the bitstream indicates music; a transient detector for setting the sbrOversamplingFlag[ch] parameter (and asserting the set parameter to stage 203) in response to detecting the presence or absence of transients in the audio content indicated by the bitstream; and / or a pitch detector for setting the sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch] parameters (and asserting the set parameters to stage 203) in response to detecting the pitch of the audio content indicated by the bitstream. Other aspects of the invention are audio bitstream decoding methods performed by any of the embodiments of the inventive decoder described in this paragraph and the preceding paragraph.
[0154] Aspects of the disclosure include a method type of encoding or decoding configured (e.g., programmed) to be performed by any embodiment of the inventive APU, system, or apparatus. Other aspects of the disclosure include a system or apparatus configured (e.g., programmed) to perform any embodiment of the inventive method and a computer-readable medium (e.g., optical disc) storing (e.g., in a non-transitory manner) code for implementing any embodiment of the inventive method or steps thereof. For example, the inventive system can be or include a programmable general purpose processor, digital signal processor, or microprocessor, any of which are programmed and / or otherwise configured using software or firmware to perform various operations on data, including embodiments of the inventive method or steps thereof. Such a general purpose processor can be or include a computer system including an input device, a memory, and a processing circuitry programmed (and / or otherwise configured) to perform embodiments of the inventive method (or steps thereof) in response to data asserted thereto.
[0155] Embodiments of the disclosure can be implemented in hardware, firmware, or software, or a combination of two or more of them, e.g., as a programmable logic array. Unless otherwise specified, the algorithms or processes included as part of the disclosure are not inherently related to any particular computer or other apparatus. In particular, various Figure 3 elements of the encoder 100 (or elements thereof), or Figure 4 elements of the decoder 200 (or elements thereof), or Figure 5 elements of the decoder 210 (or elements thereof), or elements of the decoder 400 (or elements thereof), or one or more programmable computer systems each including at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. The program code applies to input data to perform the functionality described herein and generate output information. The output information is applied to one or more output devices in a known manner.
[0156] Each such program can be implemented in any desired computer language (including machine, assembly, or high-level procedural, logical, or objected-oriented programming languages) to communicate with a computer system. In any case, the language can be a compiled or interpreted language.
[0157] For example, various functions and steps of embodiments of the application can be implemented by multi-threaded software instruction sequences running in suitable digital signal processing hardware when implemented by sequences of computer software instructions, in which case various means, steps and functions of embodiments can correspond to portions of software instructions.
[0158] Each such computer program is preferably stored upon or downloaded to a storage media or device (e.g., solid state memory or media or magnetic or optical media) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer system to perform the procedures described herein. The inventive system can also be implemented as a computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer system to operate in a specific and predefined manner to perform the functions described herein.
[0159] A number of embodiments of the application have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the application. Numerous modifications and variations of the present application can be made in light of the above teachings. For example, to facilitate efficient implementation, phase shifts can be used in combination with complex QMF analysis and synthesis filter banks. The analysis filter bank is responsible for filtering the time-domain low-band signal produced by the core decoder into a plurality of sub-bands (e.g., QMF sub-bands). The synthesis filter bank is responsible for combining the regenerated high-band produced by the selected HFR technique (as indicated by the received sbrPatchingMode parameter) with the decoded low-band to produce a wideband output audio signal. However, a given filter bank implementation operating in a certain sampling rate mode (e.g., normal double-rate operation or down-sampling SBR mode) should not have phase shifts that are dependent on the bitstream. The QMF bank used in SBR is a complex exponential extension of the theory of cosine modulation filter banks. It can be shown that when using complex exponential modulation to extend cosine modulation filter banks, the aliasing cancellation constraint becomes obsolete. Thus, for the SBR QMF bank, both the analysis filter h k (n) and the synthesis filter f k (n) can be defined by the following equations:
[0160]
[0161] where p0(n) is a real-valued symmetric or asymmetric prototype filter (typically a low-pass prototype filter), M denotes the number of frequency channels, and N is the prototype filter order. The number of frequency channels used in the analysis filter bank can be different from the number of frequency channels used in the synthesis filter bank. For example, the analysis filter bank can have 32 frequency channels and the synthesis filter bank can have 64 frequency channels. When the synthesis filter bank is operated in a down-sampling mode, the synthesis filter bank can have only 32 frequency channels. Since the sub-band samples from the filter bank are complex-valued, an additive feasible channel-dependent phase shift step can be added to the analysis filter bank. These additional phase shifts need to be compensated before the synthesis filter bank. Although the phase shift terms can in principle have any values without breaking the operation of the QMF analysis / synthesis chain, they can also be constrained to certain values for consistency verification. The SBR signal is affected by the choice of phase factors, while the low-pass signal from the core decoder is not. The audio quality of the output signal is not affected.
[0162] The coefficients p0(n) of the prototype filter can be defined as length L of 640, as shown in Table 4 below.
[0163] Table 4
[0164]
[0165]
[0166]
[0167]
[0168]
[0169]
[0170]
[0171]
[0172] The prototype filter p0(n) can also be derived from Table 4 by one or more mathematical operations such as rounding, sub-sampling, interpolation, and decimation.
[0173] Although the tuning of SBR related control information generally does not depend on the details of the transposition (as discussed previously), in some embodiments certain elements of the control data can be side- broadcast in the eSBR extension container (bs_extension_id == EXTENSION_ID_ESBR) to improve the quality of the reproduced signal. Some side-broadcast elements can include noise floor data (e.g. a noise floor scale factor and a parameter indicating the direction (frequency or time direction) of the delta encoding of each noise floor), inverse filtering data (e.g. a parameter indicating the inverse filtering mode selected from no inverse filtering, low inverse filtering degree, moderate inverse filtering degree, and strong inverse filtering degree), and missing harmonics data (e.g. a parameter indicating whether sinusoids should be added to specific frequency bands of the reproduced high frequency band). All of these elements rely on the synthesis model of the transposer of the decoder performed in the encoder and thus can improve the quality of the reproduced signal after being properly tuned according to the selected transposer.
[0174] In particular, in some embodiments, the missing harmonics and inverse filtering control data (along with the other bitstream parameters of Table 3) are transmitted in the eSBR extension container and tuned according to the harmonics transposer of eSBR. The additional bit rate required to transmit these two types of metadata for the harmonics transposer of eSBR is relatively low. Thus, sending tuned missing harmonics and / or inverse filtering control data in the eSBR extension container will improve the quality of the audio produced by the transposer while only slightly affecting the bit rate. To ensure backward compatibility with legacy decoders, the parameters tuned for the spectral translation operation of SBR can also be sent as part of the SBR control data using implicit or explicit signaling in the bitstream.
[0175] The complexity of the decoders with SBR enhancement described in this application must be limited so as not to significantly increase the overall computational complexity of the implementation. Preferably, the PCU of SBR object type is equal to or lower than 4.5 when using the eSBR tool, and the RCU of SBR object type is equal to or lower than 3 when using the eSBR tool. The approximate processing power is given in Processor Complexity Units (PCU), specified by an integer number of MOPS. The approximate RAM usage is given in RAM Complexity Units (RCU), specified by an integer number of kWords (1000 words). The RCU number does not include working buffers that can be shared between different objects and / or channels. Furthermore, the PCU is proportional to the sampling frequency. The PCU value is given in MOPS (Million Operations Per Second) per channel and the RCU value is given in kWords per channel.
[0176] Special attention needs to be paid to compressed data such as HE-AAC encoded audio which can be decoded by different decoder configurations. In this case, decoding can be done in a backward compatible way (AAC only) as well as in an enhanced way (AAC + SBR). If the compressed data allows both backward compatible and enhanced decoding and if the decoder operates in enhanced mode so that it uses a post-processor which inserts some additional delay (e.g. the SBR post-processor in HE-AAC), it has to be ensured that this additional time delay caused with respect to the backward compatible mode is taken into account when presenting the composition unit, as described by the corresponding value n. To ensure the correct handling of the composition timestamps (so that the audio stays synchronized with other media), the additional delay introduced by the post-processing given in the number of samples at the output sample rate (per audio channel) is 3010 when the decoder operating mode includes the SBR enhancement described in this application, including eSBR. Therefore, for the audio composition unit, the composition time applies to the 3011th audio sample within the composition unit when the decoder operating mode includes the SBR enhancement described in this application.
[0177] SBR enhancement should be activated to improve the subjective quality of audio content with harmonic frequency structure and strong tonal characteristics, especially at low bitrates. The values of the corresponding bitstream elements controlling these tools (i.e. esbr_data()) can be determined in the encoder by applying a signal dependent classification mechanism.
[0178] In general, the use of the harmonic patching method (sbrPatchingMode == 0) is preferred for music signals encoded at very low bitrates, where the audio bandwidth of the core codec is severely limited. This is particularly true when these signals contain a significant harmonic structure. Conversely, the use of the regular SBR patching method is preferred for speech and mixed signals, as it provides a better preservation of the temporal structure of speech.
[0179] To improve the performance of the MPEG-4 SBR transposer, a pre-processing step (bs_sbr_preprocessing == 1) can be activated, which avoids introducing spectral discontinuities of the signal to the subsequent envelope adjuster. The operation of the tool is beneficial for signal types where the coarse spectral envelope of the low-band signal used for high-frequency reconstruction shows a large level of variation.
[0180] To improve the transient response of the harmonic SBR patching (sbrPatchingMode == 0), a signal adaptive frequency domain oversampling (sbrOversamplingFlag == 1) can be applied. Since signal adaptive frequency domain oversampling increases the computational complexity of the transposer, but only brings benefits for frames containing transients, the use of this tool is controlled by a bitstream element, which is transmitted once per frame and per independent SBR channel.
[0181] Typical bitrate settings for HE-AAC v2 with SBR enhancement (i.e. a harmonic transposer with eSBR tool enabled) are recommended to correspond to 20 to 32 kbps for stereo audio content at a sampling rate of 44.1 kHz or 48 kHz. The relative subjective quality gain of SBR enhancement increases towards the lower bitrate boundary, and a properly configured encoder allows to extend this range to even lower bitrates. The bitrates provided above are recommendations only and can apply to specific service requirements.
[0182] A decoder operating in the proposed enhanced SBR mode typically needs to be able to switch between legacy SBR and enhanced SBR patching. Thus, a delay which can be as long as the duration of one core audio frame can be introduced depending on the decoder setup. Typically, the delay for both legacy SBR and enhanced SBR patching will be similar.
[0183] It is understood that the present application can be practiced with other than the specifically described methods, compositions, and apparatuses. Any element of the following claims can be expressed using the alternative language of the claims.
[0184] Various aspects of the present application can be appreciated from the following enumerated example embodiments (EEEs):
[0185] EEE 1. A method for performing high frequency reconstruction of an audio signal, the method comprising:
[0186] receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low frequency band portion of the audio signal and high frequency reconstruction metadata;
[0187] decoding the audio data to produce a decoded low frequency band audio signal;
[0188] extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata including an operation parameter of a high frequency reconstruction process, the operation parameter including a patching mode parameter positioned in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patching mode parameter indicates spectral translation and a second value of the patching mode parameter indicates harmonic transposition by phase vocoder frequency spreading;
[0189] filtering the decoded low frequency band audio signal to produce a filtered low frequency band audio signal;
[0190] regenerating a high frequency band portion of the audio signal using the filtered low frequency band audio signal and the high frequency reconstruction metadata, wherein if the patching mode parameter is the first value, the regenerating includes spectral translation, and if the patching mode parameter is the second value, the regenerating includes harmonic transposition by phase vocoder frequency spreading; and
[0191] combining the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal,
[0192] wherein the filtering, regenerating and combining are performed as a post-processing operation having a delay of 3010 samples or less per audio channel, and wherein the spectral translation comprises maintaining a ratio between tonal components and noise-like components by adaptive inverse filtering.
[0193] EEE 2. The method according to EEE 1, wherein the encoded audio bitstream further includes a padding element having an identifier indicating a start of the padding element and padding data following the identifier, wherein the padding data includes the backward-compatible extension container.
[0194] EEE 3. The method according to EEE 2, wherein the identifier is a 3-bit unsigned integer transmitted most significant bit first and having a value of 0x6.
[0195] EEE 4. The method according to EEE 2 or EEE 3, wherein the padding data includes an extension payload including spectral band replication extension data, and the extension payload is identified by a 4-bit unsigned integer transmitted most significant bit first and having a value of "1101" or "1110", and optionally,
[0196] wherein the spectral band replication extension data includes:
[0197] an optional spectral band replication header,
[0198] spectral band replication data following the header, and
[0199] a spectral band replication extension element following the spectral band replication data, and wherein the flag is included in the spectral band replication extension element.
[0200] EEE 5. The method according to any of EEEs 1 to 4, wherein the high-frequency reconstruction metadata includes an envelope scale factor, a noise floor scale factor, time / frequency grid information, or a parameter indicating a cross-over frequency.
[0201] EEE 6. The method according to any of EEEs 1 to 5, wherein the backward-compatible extension container further includes a flag indicating whether an additional pre-processing is used to avoid a shape discontinuity of a spectral envelope of the high-band portion when the patching mode parameter is equal to the first value, wherein a first value of the flag enables the additional pre-processing and a second value of the flag disables the additional pre-processing.
[0202] EEE 7. The method according to EEE 6, wherein the additional pre-processing includes using linear prediction filter coefficients to compute a pre-gain curve.
[0203] EEE 8. The method according to any of EEEs 1 to 5, wherein the backward compatible extension container further includes a flag indicating whether to apply signal adaptive frequency domain oversampling when the patch mode parameter equals the second value, wherein a first value of the flag enables the signal adaptive frequency domain oversampling and a second value of the flag disables the signal adaptive frequency domain oversampling.
[0204] EEE 9. The method according to EEE 8, wherein the signal adaptive frequency domain oversampling is applied only to frames containing transients.
[0205] EEE 10. The method of any of the preceding EEEs, wherein the harmonic transposition by phase vocoder frequency spreading is performed with an estimated complexity equal to or lower than 4.5 million operations per second and 3 kilo-words of memory.
[0206] EEE 11. A non-transitory computer-readable medium containing instructions that, when executed by a processor, perform the method according to any of EEEs 1 to 10.
[0207] EEE 12. A computer program product having instructions that, when executed by a computing device or system, cause the computing device or system to perform the method according to any of EEEs 1 to 10.
[0208] EEE 13. An audio processing unit for performing high frequency reconstruction of an audio signal, the audio processing unit comprising:
[0209] an input interface for receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low frequency band portion of the audio signal and high frequency reconstruction metadata;
[0210] a core audio decoder for decoding the audio data to produce a decoded low frequency band audio signal;
[0211] a deformatter for extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata including operational parameters for a high frequency reconstruction process, the operational parameters including a patch mode parameter positioned in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates spectral translation and a second value of the patch mode parameter indicates harmonic transposition by phase vocoder frequency spreading;
[0212] an analysis filterbank for filtering the decoded low frequency band audio signal to produce a filtered low frequency band audio signal;
[0213] a high-frequency regenerator for reconstructing a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein if the patching mode parameter is the first value, the reconstruction includes spectral translation, and if the patching mode parameter is the second value, the reconstruction includes harmonic transposition by phase vocoder frequency spreading; and
[0214] a synthesis filter bank for combining the filtered low-band audio signal and the reconstructed high-band portion to form a wide-band audio signal,
[0215] wherein the analysis filter bank, the high-frequency regenerator, and the synthesis filter bank are executed in a post-processor having a delay of 3010 samples per audio channel or less, and wherein the spectral translation comprises maintaining a ratio between tonal components and noise-like components by adaptive inverse filtering.
[0216] EEE 14. The audio processing unit of EEE 13, wherein the harmonic transposition by phase vocoder frequency spreading is executed with an estimated complexity equal to or below 4.5 million operations per second and 3 kilo-words of memory.
Claims
1. A method for performing high frequency reconstruction of an audio signal, the method comprising: receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low frequency band portion of the audio signal and high frequency reconstruction metadata; decoding the audio data to produce a decoded low frequency band audio signal; extracting the high frequency reconstruction metadata and a padding element from the encoded audio bitstream, the high frequency reconstruction metadata including an operation parameter of a high frequency reconstruction process, the operation parameter including a patch mode parameter positioned in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates spectral warping and a second value of the patch mode parameter indicates harmonic transposition by phase vocoder frequency spreading, wherein the padding element has an identifier indicating a start of the padding element and padding data following the identifier, wherein the padding data comprises the backward compatible extension container, wherein the identifier is a 3-bit unsigned integer transmitted with most significant bit first and having a value of 0x6; filtering the decoded low-band audio signal to produce a filtered low-band audio signal, wherein the filtering is performed by an analysis filterbank, the analysis filterbank comprising analysis filters h k (n), the analysis filters h k (n) being modulated versions of a prototype filter p0(n) according to the following equation: , , where p0(n) is a real-valued symmetric or asymmetric prototype filter, M is the number of frequency channels in the analysis filter bank, and N is the order of the prototype filter; and using the filtered low frequency band audio signal and the high frequency reconstruction metadata to reproduce a high frequency band portion of the audio signal; where the filtering and reproducing are performed as a post-processing operation with a delay of 3010 samples per audio frequency channel, and where the spectral warping comprises maintaining a ratio between tonal components and noise-like components by adaptive inverse filtering.
2. The method of claim 1, wherein the backward compatible extension container further includes a flag indicating whether to use additional pre-processing to avoid shape discontinuities of a spectral envelope of the high frequency band portion when the patch mode parameter is equal to the first value, wherein a first value of the flag enables the additional pre-processing and a second value of the flag disables the additional pre-processing.
3. The method of claim 2, wherein the additional pre-processing includes calculating a pre-gain curve using linear prediction filter coefficients.
4. The method of claim 1, wherein the backward compatible extension container further includes a flag indicating whether to apply signal adaptive frequency domain oversampling when the patch mode parameter is equal to the second value, wherein a first value of the flag enables the signal adaptive frequency domain oversampling and a second value of the flag disables the signal adaptive frequency domain oversampling.
5. The method of claim 4, wherein the signal adaptive frequency domain oversampling is applied only to frames containing transients.
6. The method of claim 1, wherein the harmonic transposition by phase vocoder frequency spreading is performed with an estimated complexity equal to or below 45 million operations per second and 3 kilo words of memory.
7. The method of claim 1, wherein filtering the decoded low frequency band audio signal to produce a filtered low frequency band audio signal includes filtering the decoded low frequency band audio signal into a plurality of sub-bands using a complex QMF analysis filter bank; and The method further comprises combining the filtered low-band audio signal with the regenerated high-band portion using a complex QMF synthesis filter to form a wideband audio signal.
8. An audio processing unit for performing high-frequency reconstruction of an audio signal, the audio processing unit comprising: an input interface for receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata; a core audio decoder for decoding the audio data to produce a decoded low-band audio signal; a de-formatter for extracting the high-frequency reconstruction metadata and a padding element from the encoded audio bitstream, the high-frequency reconstruction metadata including operational parameters for a high-frequency reconstruction process, the operational parameters including a patch mode parameter positioned in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates spectral warping and a second value of the patch mode parameter indicates harmonic transposition by phase vocoder frequency spreading, wherein the padding element has an identifier indicating a start of the padding element and padding data following the identifier, wherein the padding data comprises the backward compatible extension container, wherein the identifier is a 3-bit unsigned integer transmitted with most significant bit first and having a value of 0x6; an analysis filterbank for filtering the decoded low-band audio signal to generate a filtered low-band audio signal, wherein the analysis filterbank comprises analysis filters h k (n), the analysis filters h k (n) are modulated versions of a prototype filter p0(n) according to the following equation: , , where p0(n) is a real-valued symmetric or asymmetric prototype filter, M is the number of frequency channels in the analysis filter bank, and N is the order of the prototype filter; and a high-frequency regenerator for reconstructing a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata; wherein the analysis filter bank and the high-frequency regenerator are performed in a post-processor having a delay of 3010 samples per audio frequency channel, and wherein the spectral warping includes maintaining a ratio between tonal components and noise-like components by adaptive inverse filtering.
Citation Information
Patent Citations
Integration of high-frequency audio reconstruction technology
CN112189231B
Integration of high-frequency reconstruction techniques with reduced post-processing latency
CN112204659B
Integration of high frequency reconstruction techniques with reduced post-processing delay
CN114242088A
Integration of high frequency reconstruction techniques with reduced post-processing delay
CN114242089A
Integration of high frequency reconstruction techniques with reduced post-processing delay
CN114242090A