Backtracking compatible integration of high frequency reconstruction techniques for audio signals

By extracting high-frequency reconstruction metadata and using analytical filter banks to filter low-frequency audio signals, the high-frequency band portion is regenerated. Combined with the enhanced spectrum band copying tool (eSBR), the problem of suboptimal spectrum band copying in existing technologies is solved, achieving efficient encoding and high-quality audio reconstruction of low crossover frequency audio content.

CN120808802APending Publication Date: 2025-10-17DOLBY INTERNATIONAL AB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511109716.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-01-26
Filing Date
2019-01-28
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies are not ideal for certain audio types (such as music content with relatively low crossover frequencies) in spectrum band copying, and there is a need to improve spectrum band copying technology to enhance coding efficiency and audio quality.

Method used

By extracting high-frequency reconstruction metadata, filtering the low-frequency audio signal using an analytical filter bank, and regenerating the high-frequency band portion based on the flags and high-frequency reconstruction metadata, the high-frequency reconstruction of the audio signal is performed in conjunction with the enhanced spectrum band copying tool (eSBR).

Benefits of technology

It achieves efficient encoding and high-quality audio reconstruction of low crossover frequency audio content, improving encoding efficiency and audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808802A_ABST
    Figure CN120808802A_ABST
Patent Text Reader

Abstract

The invention relates to backtracking compatible integration of high frequency reconstruction techniques for audio signals. A method for decoding an encoded audio bitstream is disclosed. The method includes receiving the encoded audio bitstream and decoding the audio data to produce a decoded low-band audio signal. The method further includes extracting high frequency reconstruction metadata and filtering the decoded low band audio signal using an analysis filter bank to produce a filtered low band audio signal. The method also includes extracting a flag indicating performing spectral translation or harmonic transpose on the audio data and regenerating a high frequency band portion of the audio signal using the filtered low frequency band audio signal and the high frequency reconstruction metadata in accordance with the flag.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Information about divisional applications

[0002] This application is a divisional application. The parent application is invention patent application No. 202111240006.2, filed on January 28, 2019, entitled “Retrospective Compatible Integration of High-Frequency Reconstruction Technology for Audio Signals.”

[0003] Cross-reference to related applications

[0004] This application claims priority to U.S. Provisional Application No. 62 / 622,205, filed January 26, 2018, which is hereby incorporated by reference herein. Technical Field

[0005] Embodiments relate to audio signal processing, and more particularly, to encoding, decoding, or transcoding of an audio bitstream wherein control data indicates that a base form of high frequency reconstruction ("HFR") or an enhanced form of HFR be performed on the audio data. Background Art

[0006] A typical information bitstream includes both audio data (e.g., encoded audio data) indicating one or more channels of audio content and metadata indicating at least one characteristic of the audio data or audio content. One well-known format for generating encoded audio bitstreams is the MPEG-4 Advanced Audio Coding (AAC) format, described in MPEG standard ISO / IEC 14496-3:2009. In the MPEG-4 standard, AAC stands for "Advanced Audio Coding" and HE-AAC stands for "High Efficiency Advanced Audio Coding."

[0007] The MPEG-4 AAC standard defines several audio profiles that determine which objects and coding tools are present in a compliant encoder or decoder. Three of these audio profiles are (1) AAC profile, (2) HE-AAC profile, and (3) HE-AAC v2 profile. The AAC profile includes an AAC Low Complexity (or "AAC-LC") object type. The AAC-LC object is a counterpart to the MPEG-2 AAC Low Complexity profile with some adjustments and does not include a Spectral Band Replication ("SBR") object type or a Parametric Stereo ("PS") object type. The HE-AAC profile is a superset of the AAC profile and additionally includes an SBR object type. The HE-AAC v2 profile is a superset of the HE-AAC profile and additionally includes a PS object type.

[0008] The SBR object type contains a spectral band replication tool, which is an important high frequency reconstruction ("HFR") coding tool that significantly improves the compression efficiency of perceptual audio codecs. SBR reconstructs (e.g., in a decoder) high frequency components of an audio signal on the receiver side. Thus, an encoder only needs to encode and transmit low frequency components, allowing for higher audio quality at low data rates. SBR is based on the copying of previously truncated harmonic sequences in order to reduce the data rate from the available bandwidth limited signal and control data obtained from the encoder. The ratio between tonal and noise-like components is maintained by adaptive inverse filtering and selective addition of noise and sinusoids. In the MPEG-4 AAC standard, the SBR tool performs spectral patching (also referred to as linear or spectral panning), in which several consecutive Quadrature Mirror Filter (QMF) subbands are copied (or "patched") from a transmitted low frequency portion of an audio signal to a high frequency band portion of the audio signal (which is generated in a decoder).

[0009] For certain audio types (e.g., music content with a relatively low crossover frequency), spectral patching or linear panning can not be ideal. Thus, there is a need for techniques for improving spectral band replication. SUMMARY

[0010] A first category of embodiments is disclosed that relate to a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream and decoding the audio data to produce a decoded low band audio signal. The method further includes extracting high frequency reconstruction metadata and filtering the decoded low band audio signal using an analysis filter bank to produce a filtered low band audio signal. The method further includes extracting a flag indicating that spectral panning or harmonic transposition is performed on the audio data and regenerating a high band portion of the audio signal using the filtered low band audio signal and the high frequency reconstruction metadata according to the flag. Finally, the method includes combining the filtered low band audio signal with the regenerated high band portion to form a wide band audio signal.

[0011] A second category of embodiments relates to an audio decoder for decoding an encoded audio bitstream. The decoder includes an input interface for receiving the encoded audio bitstream, wherein the encoded audio bitstream includes audio data representing a low-band portion of an audio signal, and a core decoder for decoding the audio data to generate a decoded low-band audio signal. The decoder also includes a demultiplexer for extracting high-frequency reconstruction metadata from the encoded audio bitstream, wherein the high-frequency reconstruction metadata includes operational parameters for a high-frequency reconstruction process that linearly translates a contiguous number of subbands from the low-band portion of the audio signal to a high-band portion of the audio signal, and an analysis filterbank for filtering the decoded low-band audio signal to generate a filtered low-band audio signal. The decoder further includes a demultiplexer for extracting a flag indicating that a linear translation or a harmonic transposition is performed on the audio data from the encoded audio bitstream, and a high-frequency regenerator for regenerating the high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata in accordance with the flag. Finally, the decoder includes a synthesis filterbank for combining the filtered low-band audio signal with the regenerated high-band portion to form a wide-band audio signal.

[0012] Other categories of embodiments relate to encoding and transcoding an audio bitstream containing metadata identifying whether an enhanced spectral band replication (eSBR) process is performed. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 is a block diagram of an embodiment of a system that can be configured to perform embodiments of the inventive method.

[0014] Figure 2 is a block diagram of an encoder that is an embodiment of the inventive audio processing unit.

[0015] Figure 3 is a block diagram of a system that includes a decoder that is an embodiment of the inventive audio processing unit and optionally also a post-processor coupled to the decoder.

[0016] Figure 4 is a block diagram of a decoder that is an embodiment of the inventive audio processing unit.

[0017] Figure 5 is a block diagram of a decoder that is another embodiment of the inventive audio processing unit.

[0018] Figure 6 is a block diagram of another embodiment of the inventive audio processing unit.

[0019] Figure 7is a diagram of a block of an MPEG-4 AAC bitstream, the block comprising segments into which it is partitioned.

[0020] Notations and Nomenclature

[0021] Throughout the disclosure (including in the claims), the expression "performing an operation on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used broadly to mean performing the operation directly on the signal or data or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performing the operation thereon).

[0022] Throughout the disclosure (including in the claims), the expression "audio processing unit" or "audio processor" is used broadly to mean a system, device, or apparatus configured to process audio data. Examples of audio processing units include, but are not limited to, encoders, codecs, decoders, transcoders, pre-processing systems, post-processing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools). Essentially all consumer electronics devices (e.g., mobile phones, televisions, laptop computers, and tablet computers) contain an audio processing unit or audio processor.

[0023] Throughout the disclosure (including in the claims), the term "coupled" or "coupling" is used broadly and intended to connote either a direct connection between two devices or an indirect connection through one or more additional devices and connections. In addition, a device or element coupled to another device or element functions as a coupling device or element. DETAILED DESCRIPTION

[0024] The MPEG-4 AAC standard contemplates that an encoded MPEG-4 AAC bitstream includes metadata that indicates each type of high frequency reconstruction ("HFR") processing to be applied, if any, by a decoder to decode the audio content of the bitstream and / or controls this HFR processing and / or indicates at least one characteristic or parameter of at least one HFR tool to be used to decode the audio content of the bitstream. In this document, we use the expression "SBR metadata" to mean this type of metadata described or referred to in the MPEG-4 AAC standard for use with spectral band replication ("SBR"). As will be appreciated by those skilled in the art, SBR is a form of HFR.

[0025] SBR is preferably used as a dual-rate system, where the underlying codec operates at half the original sampling rate, while SBR operates at the original sampling rate. The SBR encoder works in parallel to the underlying core codec, but at the higher sampling rate. While SBR is mainly a post-process in the decoder, important parameters are extracted in the encoder to ensure the most accurate high frequency reconstruction in the decoder. The encoder estimates a spectral envelope for the SBR range, which is adapted to the characteristics of the current input signal segment in terms of time and frequency range / resolution. The spectral envelope is estimated by complex QMF analysis and subsequent energy computation. The time and frequency resolution of the spectral envelope can be chosen highly freely to ensure the most appropriate time frequency resolution for a given input segment. The envelope estimation needs to take into account that transients in the original, which are mainly located in the high frequency region (e.g. hat region), will be present slightly in the SBR generated high frequency band before envelope adjustment, since the high frequency band in the decoder is based on the low frequency band, where the transients are less prominent than in the high frequency band. This aspect puts different requirements on the time frequency resolution of the spectral envelope data compared to a normal spectral envelope as used in other audio coding algorithms.

[0026] In addition to the spectral envelope, several additional parameters are extracted, which represent spectral characteristics of the input signal for different time and frequency regions. Since the encoder has natural access to the original signal as well as information about how the SBR unit in the decoder will generate the high frequency band, the system can handle situations where the low frequency band constitutes a strong harmonic series and the high frequency band to be reproduced mainly constitutes a random signal component, as well as situations where a strong tonal component exists in the original high frequency band without a counterpart in the low frequency band (from which the high frequency band region is based). Furthermore, the SBR encoder works closely with the underlying core codec to assess which frequency range should be covered by SBR at a given time. In the case of a stereo signal, the SBR data is efficiently encoded before transmission by exploiting the channel dependency of the control data as well as entropy coding.

[0027] It is generally required to carefully tune the control parameter extraction algorithm to the underlying codec at a given bit rate and a given sampling rate. This is due to the fact that lower bit rates generally imply larger SBR ranges compared to high bit rates, and different sampling rates correspond to different time resolutions of the SBR frames.

[0028] An SBR decoder typically comprises several different parts. The SBR decoder comprises a bitstream decoding module, a high frequency reconstruction (HFR) module, an additional high frequency component module and an envelope adjuster module. The system is based on a complex-valued QMF filter bank (for high quality SBR) or a real-valued QMF filter bank (for low power SBR). Embodiments of the invention can be applied to both high quality SBR and low power SBR. In the bitstream extraction module, control data is read from the bitstream and decoded. A time-frequency grid is obtained for the current frame before envelope data is read from the bitstream. A lower layer core decoder decodes the audio signal of the current frame (although at a lower sampling rate) to produce time-domain audio samples. The resulting frame of audio data is used for high frequency reconstruction by the HFR module. The decoded low frequency band signal is then analyzed using a QMF filter bank. High frequency reconstruction and envelope adjustment is then performed on the subband samples of the QMF filter bank. The high frequencies are reconstructed from the low frequencies in a flexible way based on given control parameters. Furthermore, the reconstructed high frequency band is adaptively filtered on a subband channel basis according to the control data to ensure proper spectral characteristics for a given time / frequency region.

[0029] The top layer of an MPEG-4 AAC bitstream is a sequence of data blocks ("raw_data_block" elements), each of which is a segment of data (referred to herein as a "block") containing audio data (typically for a time period of 1024 or 960 samples) and related information and / or other data. In this document, we use the term "block" to refer to a segment of an MPEG-4 AAC bitstream that includes audio data (and corresponding metadata and optionally also other related data) of one (but not more than one) "raw_data_block" element.

[0030] Each block of an MPEG-4 AAC bitstream can contain several syntax elements (each of which is also materialized as a segment of data in the bitstream). Several types of these syntax elements are defined in the MPEG-4 AAC standard. Each syntax element is identified by a different value of the data element "id_syn_ele". Examples of syntax elements include "single_channel_element()", "channel_pair_element()" and "fill_element()". A single channel element is a container for audio data of a single audio channel (a mono audio signal). A channel pair element contains audio data of two audio channels (i.e., a stereo audio signal).

[0031] A stuff element is a container that contains an identifier (e.g., the value of the element "id_syn_ele" mentioned above), followed by information for data, which is referred to as "stuff data." Stuff elements have historically been used to adjust the instantaneous bit rate of a bitstream to be transmitted via a constant rate channel. By adding an appropriate amount of stuff data to each block, a constant data rate can be achieved.

[0032] According to embodiments of the present application, the stuff data can include one or more extension payloads that extend the type of data (e.g., metadata) that can be transmitted in the bitstream. The decoder receiving the bitstream with the stuff data containing the new type of data can optionally be used by a device (e.g., a decoder) receiving the bitstream to extend the functionality of the device. Thus, one of skill in the art will appreciate that a stuff element is a particular type of data structure and is different from the data structures typically used to transmit audio data (e.g., an audio payload containing channel data).

[0033] In some embodiments of the present application, the identifier used to identify a stuff element can consist of a three-bit unsigned magnitude first transmitted most significant bit ("uimsbf") having a value of 0x6. In one block, several instances of the same type of syntax element (e.g., several stuff elements) can occur.

[0034] Another standard for encoding audio bitstreams is the MPEG Unified Speech and Audio Coding (USAC) standard (ISO / IEC 23003-3:2012). The MPEG USAC standard describes encoding and decoding audio content using spectral band replication processing (including SBR processing as described in the MPEG-4 AAC standard, and also including other enhanced forms of spectral band replication processing). This processing applies a spectral band replication tool (sometimes referred to herein as an "enhanced SBR tool" or "eSBR tool") that is an extended and enhanced version of the SBR toolset described in the MPEG-4 AAC standard. Thus, eSBR (as defined in the USAC standard) is an improvement over SBR (as defined in the MPEG-4 AAC standard).

[0035] In this document, we use the expression "enhanced SBR processing" (or "eSBR processing") to mean spectral band replication processing that uses at least one eSBR tool that is not described or referred to in the MPEG-4 AAC standard (e.g., at least one eSBR tool described or referred to in the MPEG USAC standard). Examples of such eSBR tools are harmonic transposition and QMF-patching pre-processing or "pre-flattening."

[0036] An integer order T harmonic transposer maps a sinusoid with frequency ω to a sinusoid with frequency Tω while preserving signal duration. Typically three orders T = 2, 3, 4 are used in sequence to produce each part of the desired output frequency range using the smallest possible transposition order. If outputs above the fourth order transposition range are needed, they can be produced by frequency shifting. When feasible, the baseband time domain is produced with near-critical sampling for processing to minimize computational complexity.

[0037] The harmonic transposer can be based on QMF or DFT. When using a QMF based harmonic transposer, the bandwidth extension of the core coder time domain signal is carried out entirely in the QMF domain using a modified phase vocoder structure which performs decimation followed by time stretching for each QMF subband. Transposition using several transposition factors (e.g. T = 2, 3, 4) is carried out in the common QMF analysis / synthesis transform stage. Since the QMF based harmonic transposer does not feature signal adaptive frequency domain oversampling, the corresponding flag (sbrOversamplingFlag[ch]) in the bitstream can be ignored.

[0038] When using a DFT based harmonic transposer, factor 3 and 4 transposers (3rd and 4th order transposers) are preferably integrated into the factor 2 transposer (2nd order transposer) by interpolation to reduce complexity. The nominal "full size" transform size of the transposer is first determined by the signal adaptive frequency domain oversampling flag (sbrOversamplingFlag[ch]) in the bitstream for each frame (which corresponds to coreCoderFrameLength core coder samples).

[0039] When sbrPatchingMode == 1 (which indicates that linear transposition is to be used to produce the high band), an additional step can be introduced to avoid discontinuities in the shape of the spectral envelope of the high frequency signal input to the subsequent envelope adjuster. This improves the operation of the subsequent envelope adjustment stage, resulting in a high band signal which is perceived as more stable. The operation of the additional pre-processing is beneficial for signal types where the coarse spectral envelope of the low band signal used for high frequency reconstruction shows large level variations. However, the value of the bitstream element can be determined in the encoder by applying any kind of signal dependent classification. The additional pre-processing is preferably activated by the unitary bitstream element bs_sbr_preprocessing. When bs_sbr_preprocessing is set to 1, the additional pre-processing is enabled. When bs_sbr_preprocessing is set to 0, the additional pre-processing is disabled. The additional processing preferably makes use of a pre-gain curve which is used by the high frequency generator to scale the low band X Low For example, the pre-gain curve can be computed according to:

[0040] preGain(k) = 10 (meanNrg-lowEnvSlope(k)) / 20 , 0 < k < k0

[0041] where k0 is the first QMF subband in the primary band table and lowEnvSlope is computed using a function (e.g., polyfit()) that computes the best polynomial fit coefficients (in the least squares sense). For example,

[0042] polyfit(3, k0, x_lowband, lowEnv, lowEnvSlope);

[0043] may be employed (using a cubic polynomial) and where

[0044]

[0045] where x_lowband(k) = [0... k 0-1 ], numTimeSlot is the number of SBR envelope time slots present within a frame, RATE is a constant (e.g., 2) indicating the number of QMF subband samples per time slot, are linear prediction filter coefficients (potentially obtained from a covariance method) and where

[0046]

[0047] A bitstream generated according to the MPEG USAC standard (sometimes referred to herein as a "USAC bitstream") includes encoded audio content and typically includes metadata indicating each type of spectral band replication processing to be applied by a decoder to decode the audio content of the USAC bitstream, and / or controlling such spectral band replication processing, and / or at least one characteristic or parameter of at least one SBR tool and / or eSBR tool to be used to decode the audio content of the USAC bitstream.

[0048] In this document, we use the expression "enhanced SBR metadata" (or "eSBR metadata") to denote metadata that indicates each type of spectral band replication processing to be applied by a decoder to decode the audio content of an encoded audio bitstream (e.g. a USAC bitstream) and / or controls this spectral band replication processing and / or indicates at least one characteristic or parameter of at least one SBR tool and / or eSBR tool to be used to decode this audio content, but is not described or mentioned in the MPEG-4 AAC standard. An example of eSBR metadata is metadata (indicating, or used to control spectral band replication processing) that is described or mentioned in the MPEG USAC standard, but is not described or mentioned in the MPEG-4 AAC standard. Thus, eSBR metadata in this document denotes metadata that is not SBR metadata, and SBR metadata in this document denotes metadata that is not eSBR metadata.

[0049] A USAC bitstream can contain both SBR metadata and eSBR metadata. More specifically, a USAC bitstream can contain eSBR metadata that controls the execution of eSBR processing by a decoder, and SBR metadata that controls the execution of SBR processing by a decoder. According to typical embodiments of the invention, eSBR metadata (e.g. eSBR specific configuration data) is included in an MPEG-4 AAC bitstream (e.g. in an sbr_extension() container at the end of the SBR payload) according to the invention.

[0050] During the decoding of an encoded bitstream using an eSBR toolset (which includes at least one eSBR tool), the eSBR processing performed by the decoder regenerates the high frequency band of the audio signal based on the replication of the truncated harmonic series during encoding. This eSBR processing typically adjusts the spectral envelope of the generated high frequency band and applies inverse filtering, and adds noise and sinusoidal components to regenerate the spectral characteristics of the original audio signal.

[0051] According to typical embodiments of the invention, eSBR metadata (e.g. a small number of control bits that are eSBR metadata) is included in one or more of the metadata segments of an encoded audio bitstream (e.g. an MPEG-4 AAC bitstream) that also contains encoded audio data in other segments (audio data segments). Typically, at least one such metadata segment of each block of the bitstream is (or contains) a padding element (which contains an identifier that indicates the start of the padding element), and the eSBR metadata is included in the padding element after the identifier.

[0052] Figure 1is a block diagram of an exemplary audio processing chain (audio data processing system) in which one or more elements of the system can be configured in accordance with embodiments of the present disclosure. The system includes the following elements coupled together as shown: an encoder 1, a delivery subsystem 2, a decoder 3, and a post-processing unit 4. In variants of the shown system, one or more elements are omitted, or additional audio data processing units are included.

[0053] In some implementations, the encoder 1 (which optionally includes a pre-processing unit) is configured to accept as input PCM (time-domain) samples comprising audio content, and to output an encoded audio bitstream indicative of the audio content (which has a format compatible with the MPEG-4 AAC standard). The data of the bitstream indicative of the audio content is sometimes referred to herein as "audio data" or "encoded audio data". If the encoder is configured in accordance with typical embodiments of the present disclosure, the audio bitstream output from the encoder includes eSBR metadata (and typically also other metadata) as well as the audio data.

[0054] The one or more encoded audio bitstreams output from the encoder 1 can be asserted to an encoded audio delivery subsystem 2. The subsystem 2 is configured to store and / or deliver each encoded bitstream output from the encoder 1. The encoded audio bitstream output from the encoder 1 can be stored by the subsystem 2 (e.g., in the form of a DVD or Blu-ray disc), or transmitted by the subsystem 2 (which can implement a transmission link or network), or both can be stored and transmitted by the subsystem 2.

[0055] The decoder 3 is configured to decode the encoded MPEG-4 AAC audio bitstream (which was produced by the encoder 1) that it receives via the subsystem 2. In some embodiments, the decoder 3 is configured to extract eSBR metadata from each block of the bitstream, and to decode the bitstream (including by performing eSBR processing using the extracted eSBR metadata) to produce decoded audio data (e.g., a stream of decoded PCM audio samples). In some embodiments, the decoder 3 is configured to extract SBR metadata from the bitstream (but ignoring the eSBR metadata included in the bitstream), and to decode the bitstream (including by performing SBR processing using the extracted SBR metadata) to produce decoded audio data (e.g., a stream of decoded PCM audio samples). In general, the decoder 3 includes a buffer (e.g., in a non-transitory manner) that stores segments of the encoded audio bitstream received from the subsystem 2.

[0056] Figure 1 The post-processing unit 4 is configured to accept a stream of decoded audio data (e.g., decoded PCM audio samples) from the decoder 3 and to perform post-processing on it. The post-processing unit can also be configured to present the post-processed audio content (or decoded audio received from the decoder 3) for playback by one or more loudspeakers.

[0057] Figure 2is a block diagram of an encoder 100 that is an embodiment of an inventive audio processing unit. Any component or element of the encoder 100 can be implemented as one or more processes and / or one or more circuits (e.g., ASICs, FPGAs, or other integrated circuits) in hardware, software, or a combination of hardware and software. The encoder 100 includes an encoder 105, a filler / formatting stage 107, a metadata generation stage 106, and a buffer memory 109, connected as shown. In general, the encoder 100 also includes other processing elements (not shown). The encoder 100 is configured to convert an input audio bitstream to an encoded output MPEG-4 AAC bitstream.

[0058] The metadata generator 106 is coupled and configured to generate metadata (including eSBR metadata and SBR metadata) (and / or pass the metadata to the stage 107) to be included in the encoded bitstream to be output from the encoder 100 by the stage 107.

[0059] The encoder 105 is coupled and configured to encode input audio data (e.g., by performing compression on it), and to assert the resulting encoded audio to the stage 107 to be included in the encoded bitstream to be output from the stage 107.

[0060] The stage 107 is configured to multiplex the encoded audio from the encoder 105 and the metadata (including eSBR metadata and SBR metadata) from the generator 106 to generate an encoded bitstream to be output from the stage 107, preferably such that the encoded bitstream has a format as specified by one of the embodiments of the invention.

[0061] The buffer memory 109 is configured to store (e.g., in a non-transitory manner) at least one block of the encoded audio bitstream output from the stage 107, and then to assert a sequence of blocks of the encoded audio bitstream from the buffer memory 109 as output from the encoder 100 to a delivery system.

[0062] Figure 3 is a block diagram of a system that includes a decoder 200 that is an embodiment of an inventive audio processing unit, and optionally also a post-processor 300 coupled to the decoder 200. Any component or element of the decoder 200 and the post-processor 300 can be implemented as one or more processes and / or one or more circuits (e.g., ASICs, FPGAs, or other integrated circuits) in hardware, software, or a combination of hardware and software. The decoder 200 includes a buffer memory 201, a bitstream payload de-formatter (parser) 205, an audio decoding subsystem 202 (sometimes referred to as a "core" decoding stage or "core" decoding subsystem), an eSBR processing stage 203, and a control bit generation stage 204, connected as shown. In general, the decoder 200 also includes other processing elements (not shown).

[0063] A buffer memory (buffer) 201 stores (e.g., in a non-transitory manner) at least one block of an encoded MPEG-4 AAC audio bitstream received by the decoder 200. In operation of the decoder 200, sequential blocks of the bitstream are asserted from the buffer 201 to the deformatter 205.

[0064] In Figure 3 embodiments (or to be described Figure 4 embodiments) of the decoder 200, the non-decoder APU (e.g., the APU 500 of Figure 6 ) includes a buffer memory (e.g., the same buffer memory as the buffer 201) that stores (e.g., in a non-transitory manner) at least one block of the same type of encoded audio bitstream (e.g., MPEG-4 AAC audio bitstream) (i.e., an encoded audio bitstream that includes eSBR metadata) received by the buffer 201 of the decoder 200. Figure 3 or Figure 4 In

[0065] embodiments (or to be described Figure 3 embodiments) of the decoder 200, the non-decoder APU (e.g., the APU 500 of ) includes a buffer memory (e.g., the same buffer memory as the buffer 201) that stores (e.g., in a non-transitory manner) at least one block of the same type of encoded audio bitstream (e.g., MPEG-4 AAC audio bitstream) (i.e., an encoded audio bitstream that includes eSBR metadata) received by the buffer 201 of the decoder 200.

[0066] Figure 3 In embodiments (or to be described

[0067] embodiments) of the decoder 200, the non-decoder APU (e.g., the APU 500 of ) includes a buffer memory (e.g., the same buffer memory as the buffer 201) that stores (e.g., in a non-transitory manner) at least one block of the same type of encoded audio bitstream (e.g., MPEG-4 AAC audio bitstream) (i.e., an encoded audio bitstream that includes eSBR metadata) received by the buffer 201 of the decoder 200.The audio decoding subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the parser 205 (this decoding can be referred to as a "core" decoding operation) to generate decoded audio data, and to assert the decoded audio data to the eSBR processing stage 203. The decoding is performed in the frequency domain and typically includes inverse quantization followed by spectral processing. Generally, the last stage of processing in the subsystem 202 applies a frequency-to-time domain transform to the decoded frequency domain audio data, such that the output of the subsystem is time domain, decoded audio data. The stage 203 is configured to apply the SBR tool and the eSBR tool indicated by the SBR metadata and the eSBR (extracted by the parser 205) to the decoded audio data (i.e., to perform SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata) to generate fully decoded audio data that is output from the decoder 200 (e.g., to the post-processor 300). Generally, the decoder 200 includes a memory that stores the deformatted audio data and metadata output from the deformatter 205 (accessible by the subsystem 202 and the stage 203), and the stage 203 is configured to access the audio data and metadata (including the SBR metadata and the eSBR metadata) as needed during the SBR and eSBR processing. The SBR processing and the eSBR processing in the stage 203 can be viewed as post-processing of the output of the core decoding subsystem 202. Optionally, the decoder 200 also includes a final upmixing subsystem (which can apply the parametric stereo ("PS") tool defined in the MPEG-4 AAC standard, using PS metadata extracted by the deformatter 205 and / or control bits generated in the subsystem 204) that is coupled and configured to perform upmixing on the output of the stage 203 to generate fully decoded, upmixed audio that is output from the decoder 200. Alternatively, the post-processor 300 is configured to perform upmixing on the output of the decoder 200 (e.g., using PS metadata extracted by the deformatter 205 and / or control bits generated in the subsystem 204).

[0068] In response to the metadata extracted by the deformatter 205, the control bit generator 204 may generate control data, and the control data may be used within the decoder 200 (e.g., in a final upmix system) and / or asserted as an output of the decoder 200 (e.g., to the post-processor 300 for use in post-processing). In response to the metadata extracted from the input bitstream (and optionally also in response to the control data), the stage 204 may generate (and assert to the post-processor 300) a control bit indicating that the decoded audio data output from the eSBR processing stage 203 should undergo a particular type of post-processing. In some implementations, the decoder 200 is configured to assert the metadata extracted from the input bitstream by the deformatter 205 to the post-processor 300, and the post-processor 300 is configured to use the metadata to perform post-processing on the decoded audio data output from the decoder 200.

[0069] Figure 4 2 is a block diagram of an audio processing unit ("APU") (210), which is another embodiment of an audio processing unit of the present invention. APU 210 is a legacy decoder that is not configured to perform eSBR processing. Any component or element of APU 210 may be implemented as one or more processes and / or one or more circuits (e.g., an ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware and software. APU 210 includes a buffer memory 201, a bitstream payload deformatter (parser) 215, an audio decoding subsystem 202 (sometimes referred to as the "core" decoding stage or "core" decoding subsystem), and an SBR processing stage 213, connected as shown. In general, APU 210 also includes other processing elements (not shown). APU 210 may represent, for example, an audio encoder, decoder, or codec.

[0070] The components 201 and 202 of the APU 210 are Figure 3 Like numbered elements of the decoder 200 are identical and their above description will not be repeated. In operation of the APU 210, a sequence of blocks of an encoded audio bitstream (MPEG-4 AAC bitstream) received by the APU 210 is asserted from the buffer 201 to the deformatter 215.

[0071] According to any embodiment of the present invention, the deformatter 215 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (including quantization envelope data) and typically also extract other metadata therefrom, but ignores eSBR metadata that may be included in the bitstream. The deformatter 215 is configured to assert at least the SBR metadata to the SBR processing stage 213. The deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.

[0072] The audio decoding subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the deformatter 215 (this decoding can be referred to as a "core" decoding operation) to generate decoded audio data, and to assert the decoded audio data to the SBR processing stage 213. The decoding is performed in the time domain. In general, the last stage of processing in the subsystem 202 applies a frequency-to-time domain transform to the decoded frequency domain audio data, such that the output of the subsystem is time domain, decoded audio data. The stage 213 is configured to apply SBR tools (but not eSBR tools) indicated by the SBR metadata (extracted by the parser 215) to the decoded audio data (i.e., SBR processing is performed on the output of the decoding subsystem 202 using the SBR metadata) to generate fully decoded audio data that is output from the APU 210 (e.g., to the post-processor 300). In general, the APU 210 includes a memory that stores the deformatted audio data and metadata output from the deformatter 215 (which is accessible by the subsystem 202 and the stage 213), and the stage 213 is configured to access the audio data and metadata (including the SBR metadata) as needed during SBR processing. The SBR processing in the stage 213 can be viewed as post-processing of the output of the core decoding subsystem 202. Optionally, the APU 210 also includes a final upmixing subsystem (which can apply the parametric stereo ("PS") tools defined in the MPEG-4 AAC standard, using PS metadata extracted by the deformatter 215) that is coupled and configured to perform upmixing on the output of the stage 213 to generate fully decoded, upmixed audio that is output from the APU 210. Alternatively, the post-processor is configured to perform upmixing on the output of the APU 210 (e.g., using PS metadata extracted by the deformatter 215 and / or control bits generated in the APU 210).

[0073] Various implementations of the encoder 100, the decoder 200, and the APU 210 are configured to perform different embodiments of the inventive method.

[0074] According to some embodiments, eSBR metadata (e.g., comprising a small number of control bits that are eSBR metadata) is included in an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) such that legacy decoders (which are not configured to parse eSBR metadata or use any eSBR tools related to eSBR metadata) can ignore the eSBR metadata, but still decode the bitstream as much as possible without using the eSBR metadata or any eSBR tools related to the eSBR metadata, typically without any significant loss in decoded audio quality. However, eSBR decoders configured to parse the bitstream to identify eSBR metadata and use at least one eSBR tool in response to the eSBR metadata will enjoy the benefits of using at least one such eSBR tool. Thus, embodiments of the invention provide a means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backwards-compatible manner.

[0075] In general, the eSBR metadata in the bitstream indicates one or more of the following eSBR tools (which are described in the MPEG USAC standard and which can or can not be applied by an encoder during generation of the bitstream) (e.g., indicates at least one characteristic or parameter of one or more of the following eSBR tools):

[0076] • Harmonic transposition; and

[0077] • QMF-patching additional pre-processing (pre-flattening).

[0078] For example, the eSBR metadata included in the bitstream can indicate values of the parameters (described in the MPEG USAC standard and in this invention) sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch], and bs_sbr_preprocessing.

[0079] In this document, the notation X[ch] (where X is some parameter) means that the parameter is related to a channel ("ch") of the audio content of the encoded bitstream to be decoded. For brevity, we sometimes omit the expression [ch] and assume that the relevant parameter is related to a channel of the audio content.

[0080] In this document, the notation X[ch][env] (where X is some parameter) means that the parameter is related to an SBR envelope ("env") of a channel ("ch") of the audio content of the encoded bitstream to be decoded. For brevity, we sometimes omit the expressions [env] and [ch] and assume that the relevant parameter is related to an SBR envelope of a channel of the audio content.

[0081] During decoding of the encoded bitstream, the execution of the harmonic transposition during the eSBR processing stage of the decoding (for each channel "ch" of the audio content indicated by the bitstream) is controlled by the following eSBR metadata parameters: sbrPatchingMode[ch]: sbrOversamplingFlag[ch]; sbrPitchInBinsFlag[ch]; and sbrPitchInBins[ch].

[0082] The value "sbrPatchingMode[ch]" indicates the type of transposer used in eSBR: sbrPatchingMode[ch]=l indicates linear transposition patching as described in section 4.6.18 of the MPEG-4 AAC standard (as used with High-Quality SBR or Low-Power SBR); sbrPatchingMode[ch]=0 indicates harmonic SBR patching as described in section 7.5.3 or 7.5.4 of the MPEG USAC standard.

[0083] The value "sbrOversamplingFlag[ch]" indicates the use of signal-adaptive frequency-domain oversampling in eSBR in combination with DFT-based harmonic SBR patching as described in section 7.5.3 of the MPEG USAC standard. This flag controls the size of the DFT utilized in the transposer: 1 indicates that signal-adaptive frequency-domain oversampling is enabled as described in section 7.5.3.1 of the MPEG USAC standard; 0 indicates that signal-adaptive frequency-domain oversampling is disabled as described in section 7.5.3.1 of the MPEG USAC standard.

[0084] The value "sbrPitchInBinsFlag[ch]" controls the interpretation of the sbrPitchInBins[ch] parameter: 1 indicates that the value in sbrPitchInBins[ch] is valid and greater than zero; 0 indicates that the value of sbrPitchInBins[ch] is set to zero.

[0085] The value "sbrPitchInBins[ch]" controls the addition of cross-product terms in the SBR harmonic transposer. The value sbrPitchinBins[ch] is an integer value in the range [0, 127] and represents the distance measured in frequency bins for a 1536-line DFT operating at the sampling frequency of the core encoder.

[0086] In case the MPEG-4 AAC bitstream indicates that the SBR channels of a channel pair are not coupled (instead of a single SBR channel), the bitstream indicates two instances of the above syntax (for harmonic or non-harmonic transform), one for each channel of the sbr_channel_pair_element().

[0087] Harmonic transposition of the eSBR tool generally improves the quality of decoded music signals at relatively low crossover frequencies. Non-harmonic transposition (i.e., old-style spectral patching) generally improves speech signals. Thus, a starting point in deciding which type of transposition is preferred for encoding a particular audio content is to select a transposition method depending on speech / music detection, where harmonic transposition is employed for music content and spectral patching for speech content.

[0088] The execution of pre-emphasis during eSBR processing is controlled by the value of a unit eSBR metadata parameter called "bs_sbr_preprocessing", which means that pre-emphasis is performed or not depending on the value of this single bit. When the SBR QMF-patching algorithm as described in section 4.6.18.6.3 of the MPEG-4 AAC standard is used, the steps of pre-emphasis can be performed (when indicated by the "bs_sbr_preprocessing" parameter) in an effort to avoid discontinuities in the shape of the spectral envelope of the high frequency signal input to the subsequent envelope adjuster, which performs another stage of the eSBR processing. Pre-emphasis generally improves the operation of the subsequent envelope adjustment stage, resulting in a high frequency band signal that is perceived as more stable.

[0089] It is expected that the overall bit rate requirement for including eSBR metadata in the MPEG-4 AAC bitstream indicating the above-mentioned eSBR tools (harmonic transposition and pre-emphasis) is on the order of a few hundred bits per second, since according to some embodiments of the present application only the differential control data required to perform the eSBR processing is transmitted. Old-style decoders can ignore this information, since it is included in a backwards compatible way (as will be explained later). Thus, the adverse impact of the bit rate associated with including the eSBR metadata can be negligible for several reasons, including the following:

[0090] • The bit rate loss (due to including the eSBR metadata) is a very small fraction of the total bit rate, since only the differential control data required to perform the eSBR processing (and not a simulcast of the SBR control data) is transmitted; and

[0091] • The tuning of the SBR-related control information generally does not depend on the details of the transposition. Examples of when the control data does depend on the operation of the transposer are discussed later in this application.

[0092] Accordingly, embodiments of the present application provide a means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backwards compatible manner. This efficient transmission of eSBR control data reduces memory requirements in decoders, encoders and codecs employing aspects of the present application, while having no tangible adverse impact on bit rate. Furthermore, complexity and processing requirements associated with performing eSBR according to embodiments of the present application are also reduced, as SBR data need only be processed once and not simulcast, as would be the case if eSBR were treated as a completely independent object type in MPEG-4 AAC rather than integrated in a backwards compatible manner into the MPEG-4 AAC codec.

[0093] Next, with reference to Figure 7 , we describe elements of a block ("raw_data_block") of an MPEG-4 AAC bitstream containing eSBR metadata according to some embodiments of the present application. Figure 7 is a diagram of a block ("raw_data_block") of an MPEG-4 AAC bitstream showing some segments of the block.

[0094] A block of an MPEG-4 AAC bitstream can contain at least one "single_channel_element()" (e.g., a single channel element shown in Figure 7 ) and / or at least one "channel_pair_element()" (although it can exist, it is not specifically shown in Figure 7 ), containing audio data for an audio program. The block can also contain a number of "fill_elements" (e.g., fill element 1 and / or fill element 2 of Figure 7 ), containing data (e.g., metadata) related to the program. Each "single_channel_element()" contains an identifier (e.g., "ID1" of Figure 7 ) indicating the start of a single channel element, and can contain audio data for a different channel of a multi-channel audio program. Each "channel_pair_element()" contains an identifier (not shown in Figure 7 ) indicating the start of a channel pair element, and can contain audio data for two channels of a program.

[0095] A fill_element (referred to herein as a fill element) of an MPEG-4 AAC bitstream contains an identifier (e.g., "ID2" of Figure 7ID2) and padding data following the identifier. The identifier ID2 can consist of a three-bit unsigned magnitude first transfer most significant bit ("uimsbf") with a value of 0x6. The padding data can include an extension_payload() element (sometimes referred to herein as an extension payload), the syntax of which is shown in Table 4.57 of the MPEG-4 AAC standard. Several types of extension payloads exist and are identified by an "extension_type" parameter, which is a four-bit unsigned magnitude first transfer most significant bit ("uimsbf").

[0096] The padding data (e.g., extension payload thereof) can include a header or identifier (e.g., Figure 7 "Header 1") that indicates a segment of padding data that indicates an SBR object (i.e., the header initializes the "SBR object" type, which is referred to in the MPEG-4 AAC standard as sbr_extension_data()). For example, a spectral band replication (SBR) extension payload is identified with a value of '1101' or '1110' for the extension_type field in the header, where the identifier '1101' identifies an extension payload with SBR data and '1110' identifies an extension payload with SBR data having a cyclic redundancy check (CRC) to verify the correctness of the SBR data.

[0097] When the header (e.g., extension_type field) initializes the SBR object type, SBR metadata (sometimes referred to herein as "spectral band replication data" and referred to in the MPEG-4 AAC standard as sbr_data()) follows the header, and at least one spectral band replication extension element (e.g., Figure 7 " SBR extension element" of padding element 1) can follow the SBR metadata. This spectral band replication extension element (a segment of the bitstream) is referred to in the MPEG-4 AAC standard as an "sbr_extension()" container. The spectral band replication extension element optionally includes a header (e.g., Figure 7 "SBR extension header" of padding element 1).

[0098] The MPEG-4 AAC standard contemplates that the spectral band replication extension element can include PS (parametric stereo) data for the audio data of the program. The MPEG-4 AAC standard contemplates that when the header (e.g., of the extension payload thereof) of the padding element initializes the SBR object type (as Figure 7"bs_extension_id" parameter whose value (i.e., bs_extension_id = 2) indicates that PS data is included in the spectral band replication extension element of the padding element.

[0099] According to some embodiments of the application, eSBR metadata (e.g., a flag indicating whether to perform enhanced spectral band replication (eSBR) processing on the audio content of a block) is included in the spectral band replication extension element of a padding element. For example, this flag is indicated in padding element 1 of Figure 7 , where the flag appears after the header of the "SBR extension element" of padding element 1 (the "SBR extension header" of padding element 1). Optionally, this flag and additional eSBR metadata are included in the spectral band replication extension element after the header of the spectral band replication extension element (e.g., in the SBR extension element of padding element 1 in Figure 7 , after the SBR extension header). According to some embodiments of the application, the padding element including eSBR metadata also includes a "bs_extension_id" parameter whose value (e.g., bs_extension_id = 3) indicates that eSBR metadata is included in the padding element and that eSBR processing is to be performed on the audio content of the associated block.

[0100] According to some embodiments of the application, eSBR metadata is included in a padding element (e.g., padding element 2 of Figure 7 ) of an MPEG-4 AAC bitstream, rather than in the spectral band replication extension element (SBR extension element) of a padding element. This is because a padding element containing an extension_payload() with SBR data or SBR data with CRC does not contain any other extension payload of any other extension type. Thus, in embodiments in which eSBR metadata stores its own extension payload, a separate padding element is used to store the eSBR metadata. This padding element includes an identifier indicating the start of the padding element (e.g., "ID2" of Figure 7 ) and padding data following the identifier. The padding data can include an extension_payload() element (sometimes referred to herein as an extension payload), the syntax of which is shown in Table 4.57 of the MPEG-4 AAC standard. The padding data (e.g., its extension payload) includes a header (e.g., "header" of Figure 7The header 2 of the filler element 2 is "header 2" (i.e., the header initializes the enhanced spectral band replication (eSBR) object type), and the filler data (e.g., its extended payload) contains the eSBR metadata after the header. For example, Figure 7 The filler element 2 of contains this header ("Header 2") and also contains eSBR metadata after the header (i.e., a "flag" in filler element 2 that indicates whether enhanced spectral band replication (eSBR) processing was performed on the audio content of the block). Optionally, Figure 7 The padding data of the padding element 2 of the header 2 also includes additional eSBR metadata. In the embodiment described in this paragraph, the header (e.g. Figure 7 The header 2) has an identification value that is not one of the conventional values ​​specified in Table 4.57 of the MPEG-4 AAC standard and instead indicates an eSBR extended payload (such that the extension_type field of the header indicates that the padding data includes eSBR metadata).

[0101] In a first category of embodiments, the present invention is an audio processing unit (e.g., a decoder) comprising:

[0102] Memory (e.g. Figure 3 or 4) configured to store at least one block of an encoded bitstream (e.g., at least one block of an MPEG-4 AAC bitstream);

[0103] Bitstream payload deformatter (e.g. Figure 3 Component 205 or Figure 4 215) coupled to the memory and configured to demultiplex at least a portion of the block of the bitstream; and

[0104] Decoding subsystem (e.g. Figure 3 Elements 202 and 203 or Figure 4 202 and 213 ) coupled and configured to decode at least a portion of the audio content of the block of the bitstream, wherein the block comprises:

[0105] A filler element comprising an identifier indicating the start of the filler element (e.g., the "id_syn_ele" identifier with a value of 0x6 of Table 4.85 of the MPEG-4 AAC standard) and filler data following the identifier, wherein the filler data comprises:

[0106] At least one flag identifying whether to perform enhanced spectral band replication (eSBR) processing on the audio content of the block (eg, using spectral band replication data and eSBR metadata contained in the block).

[0107] The flag is an eSBR metadata, and an instance of the flag is the sbrPatchingMode flag. Another instance of the flag is the harmonicSBR flag. Both of these flags indicate whether to perform a basic form of spectral band replication or an enhanced form of spectral replication on the audio data of the block. The basic form of spectral replication is spectral patching, and the enhanced form of spectral band replication is harmonic transposition.

[0108] In some embodiments, the padding data further includes additional eSBR metadata (i.e., eSBR metadata other than flags).

[0109] The memory can be a buffer memory (e.g., in a non-transitory manner) storing the at least one block of the encoded audio bitstream. Figure 4 embodiment of the buffer 201 of FIG. 1).

[0110] It is estimated that the complexity of performing eSBR processing (using eSBR harmonic transposition and pre- flattening) by an eSBR decoder during the decoding of an MPEG-4 AAC bitstream that includes eSBR metadata (indicating these eSBR tools) would be as follows (for typical decoding using the indicated parameters):

[0111] • Harmonic transposition (16 kbps, 14400 / 28800 Hz)

[0112] o DFT-based: 3.68 WMOPS (Weighted Million Operations Per Second);

[0113] o QMF-based: 0.98 WMOPS;

[0114] • QMF patching pre-processing (pre-flattening): 0.1 WMOPS.

[0115] It is known that DFT-based transposition generally performs better for transient than QMF-based transposition.

[0116] According to some embodiments of the present application, the stuffing element (of the encoded audio bitstream) containing the eSBR metadata also contains a parameter (e.g., the "bs_extension_id" parameter) and / or a value thereof (e.g., bs_extension_id = 3) signaling that the eSBR metadata is contained in the stuffing element and to be performed on the audio content of the associated block, and / or a parameter (e.g., the same "bs_extension_id" parameter) and / or a value thereof (e.g., bs_extension_id = 2) signaling that the sbr_extension() container of the stuffing element contains PS data. For example, as indicated in Table 1 below, this parameter with value bs_extension_id = 2 can signal that the sbr_extension() container of the stuffing element contains PS data, and this parameter with value bs_extension_id = 3 can signal that the sbr_extension() container of the stuffing element contains eSBR metadata:

[0117] Table 1

[0118]

[0119]

[0120] According to some embodiments of the present application, the syntax of each spectral band replication extension element containing eSBR metadata and / or PS data is as indicated in Table 2 below (where "sbr_extension()" denotes the container of the spectral band replication extension element, "bs_extension_id" is as described in Table 1 above, "ps_data" denotes PS data, and "esbr_data" denotes eSBR metadata):

[0121] Table 2

[0122]

[0123] In the exemplary embodiments, the esbr_data() mentioned in the above Figure 2 indicates the values of the following metadata parameters:

[0124] 1. the bit data parameter "bs_sbr_preprocessing"; and

[0125] 2. for each channel ("ch") of the audio content of the encoded bitstream to be decoded, each of the above described parameters: "sbrPatchingMode[ch]"; "sbrOversamplingFlag[ch]"; "sbrPitchInBinsFlag[ch]"; and "sbrPitchInBins[ch]".

[0126] For example, in some embodiments, esbr_data() may have the syntax indicated in Table 3 to indicate these metadata parameters:

[0127] Table 3

[0128]

[0129]

[0130] The above syntax enables efficient implementation of enhanced forms of spectral band replication (e.g., harmonic transposition) as an extension to legacy decoders. Specifically, the eSBR data of Table 3 only contains those parameters required to perform the enhanced form of spectral band replication that are not already supported in the bitstream or are not directly derivable from parameters already supported in the bitstream. All other parameters and processing data required to perform the enhanced form of spectral band replication are extracted from pre-existing parameters in defined locations in the bitstream.

[0131] For example, an MPEG-4 HE-AAC or HE-AAC v2 compliant decoder can be extended to include an enhanced form of spectral band replication, such as harmonic transposition. This enhanced form of spectral band replication is in addition to the basic form of spectral band replication already supported by the decoder. In the context of an MPEG-4 HE-AAC or HE-AAC v2 compliant decoder, this basic form of spectral band replication is the QMF spectral patching (SBR) tool as defined in section 4.6.18 of the MPEG-4 AAC standard.

[0132] When performing an enhanced form of spectral band replication, the extended HE-AAC decoder can reuse many bitstream parameters already included in the SBR extension payload of the bitstream. Specific parameters that can be reused include, for example, various parameters that determine the master band table. These parameters include bs_start_freq (a parameter that determines the start of the master frequency table parameters), bs_stop_freq (a parameter that determines the stop of the master frequency table), bs_freq_scale (a parameter that determines the number of frequency bands per octave), and bs_alter_scale (a parameter that changes the scale of the frequency bands). Reusable parameters also include parameters that determine the noise band table (bs_noise_bands) and the limiter band table parameters (bs_limiter_bands). Therefore, in various embodiments, at least some equivalent parameters specified in the USAC standard are omitted from the bitstream, thereby reducing control overhead in the bitstream. Generally speaking, where a parameter specified in the AAC standard has an equivalent parameter specified in the USAC standard, the equivalent parameter specified in the USAC standard has the same name as the parameter specified in the AAC standard, for example, the envelope scaling factor EOrigMapped However, the equivalent parameters specified in the USAC standard generally have different values, being "tuned" for the enhanced SBR process defined in the USAC standard rather than for the SBR process defined in the AAC standard.

[0133] In order to improve the subjective quality of audio content with harmonic frequency structure and strong tonal characteristics, especially at low bit rates, it is recommended to activate enhanced SBR. The values ​​of the corresponding bitstream elements (i.e., esbr_data()) that control these tools can be determined in the encoder by applying a signal-dependent classification mechanism. In general, the use of the harmonic patching method (sbrPatchingMode == 1) is preferred for encoding music signals at very low bit rates, where the core codec may be significantly limited in terms of audio bandwidth. This is especially true if these signals contain a significant harmonic structure. In contrast, the use of the conventional SBR patching method is preferred for speech and mixed signals because it provides a better preservation of the temporal structure in speech.

[0134] In order to improve the performance of the harmonic transposer, a preprocessing step can be enabled (bs_sbr_preprocessing == 1) which strives to avoid introducing spectral discontinuities into the signal entering the subsequent envelope adjuster. The operation of the tool is beneficial for signal types in which the coarse spectral envelope of the low-band signal used for high-frequency reconstruction shows large level variations.

[0135] To improve the transient response of the harmonic SBR patching, signal-adaptive frequency-domain oversampling can be applied (sbrOversamplingFlag == 1). Since signal-adaptive frequency-domain oversampling increases the computational complexity of the transposer but only benefits frames containing transients, its use is controlled by a bitstream element that is transmitted once per frame and per independent SBR channel.

[0136] Decoders operating in the proposed enhanced SBR mode typically need to be able to switch between legacy SBR patching and enhanced SBR patching. Consequently, a delay may be introduced that, depending on the decoder settings, may be as long as the duration of one Core Audio frame. Generally speaking, the delay for legacy SBR patching and for enhanced SBR patching will be similar.

[0137] In addition to many parameters, the extended HE-AAC decoder may also reuse other data elements when performing an enhanced form of spectral band replication according to embodiments of the present invention. For example, envelope data and noise data may also be extracted from the bs_data_env (envelope scale factor) and bs_noise_env (noise floor scale factor) data and used during the enhanced form of spectral band replication.

[0138] In essence, these embodiments utilize configuration parameters and envelope data already supported by legacy HE-AAC or HE-AAC v2 decoders in the SBR extension payload to implement an enhanced form of spectral band replication that requires as little additional transmitted data as possible. The metadata is originally tuned for the base form of HFR (e.g., the spectral translation operation of SBR), but according to embodiments, the metadata is used for the enhanced form of HFR (e.g., the harmonic transposition of eSBR). As previously discussed, the metadata generally represents operation parameters (e.g., envelope scale factor, noise floor scale factor, time / frequency grid parameters, sinusoidal addition information, variable crossover frequency / bands, inverse filtering mode, envelope resolution, smoothing mode, frequency interpolation mode) that are tuned and intended for use with the base form of HFR (e.g., linear spectral translation). However, this metadata in combination with additional metadata parameters specific to the enhanced form of HFR (e.g., harmonic transposition) can be used to efficiently and effectively process audio data using the enhanced form of HFR.

[0139] Thus, an extended decoder that supports the enhanced form of spectral band replication can be produced in a very efficient manner by relying on already defined bitstream elements (e.g., bitstream elements in the SBR extension payload) and only adding (in the padding element extension payload) the parameters needed to support the enhanced form of spectral band replication. This data reduction feature, combined with placing the newly added parameters in a reserved data field (e.g., extension container), substantially reduces the barrier to producing decoders that support the enhanced form of spectral band replication by ensuring that the bitstream is backwards compatible with legacy decoders that do not support the enhanced form of spectral band replication. It will be appreciated that the reserved data field is a backwards compatible data field that is a data field already supported by earlier decoders (e.g., legacy HE-AAC or HE-AAC v2 decoders). Similarly, the extension container is backwards compatible that is an extension container already supported by earlier decoders (e.g., legacy HE-AAC or HE-AAC v2 decoders).

[0140] In Table 3, the numbers in the right column indicate the number of bits for the corresponding parameter in the left column.

[0141] In some embodiments, the SBR object type defined in MPEG-4 AAC is updated to contain aspects of the SBR tool and the enhanced SBR (eSBR) tool as signaled in the SBR extension element (bs_extension_id == EXTENSION_ID_ESBR). If a decoder detects this SBR extension element, the decoder employs the signaled aspects of the enhanced SBR tool.

[0142] In some embodiments, the present invention is a method comprising the steps of encoding audio data to produce an encoded bitstream (e.g., an MPEG-4 AAC bitstream), including by including eSBR metadata in at least one fragment of at least one block of the encoded bitstream and including audio data in at least another fragment of the block. In a typical embodiment, the method comprises the steps of multiplexing the audio data with the eSBR metadata in each block of the encoded bitstream. In typical decoding of the encoded bitstream in an eSBR decoder, the decoder extracts the eSBR metadata from the bitstream (including by parsing and demultiplexing the eSBR metadata and audio data) and uses the eSBR metadata to process the audio data to produce a decoded audio data stream.

[0143] Another aspect of the present invention is an eSBR decoder configured to perform eSBR processing (e.g., using at least one of the eSBR tools known as harmonic transposition or pre-flattening) during decoding of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that does not include eSBR metadata. Figure 5 Describes an instance of this decoder.

[0144] Figure 5 The eSBR decoder (400) comprises a buffer memory 201 connected as shown (which is connected to Figure 3 and 4 201), a bitstream payload deformatter 215 (which is the same as Figure 4 The deformatter 215 is the same as the audio decoding subsystem 202 (which is sometimes called the "core" decoding stage or the "core" decoding subsystem and is the same as the audio decoding subsystem 202). Figure 3 core decoding subsystem 202), eSBR control data generation subsystem 401 and eSBR processing stage 203 (which is the same as Figure 3 Generally, decoder 400 also includes other processing elements (not shown).

[0145] In operation of the decoder 400 , a sequence of blocks of an encoded audio bitstream (MPEG-4 AAC bitstream) received by the decoder 400 is asserted from the buffer 201 to the deformatter 215 .

[0146] The deformatter 215 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (including quantization envelope data) and typically other metadata therefrom. The deformatter 215 is also configured to forward at least the SBR metadata to the eSBR processing stage 203. The deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and forward the extracted audio data to the decoding subsystem (decoding stage) 202.

[0147] The audio decoding subsystem 202 of the decoder 400 is configured to decode the audio data extracted by the deformatter 215 (this decoding can be referred to as a "core" decoding operation) to generate decoded audio data, and to assert the decoded audio data to the eSBR processing stage 203. The decoding is performed in the frequency domain. In general, the last stage of processing in the subsystem 202 applies a frequency-to-time domain transform to the decoded frequency domain audio data, so that the output of the subsystem is time domain, decoded audio data. The stage 203 is configured to apply SBR tools (and eSBR tools) indicated by the SBR metadata (extracted by the deformatter 215) and by the eSBR metadata generated in the subsystem 401 to the decoded audio data (i.e., to perform SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata) to generate fully decoded audio data that is output from the decoder 400. In general, the decoder 400 includes a memory that stores the deformatted audio data and metadata output from the deformatter 215 (and optionally also from the subsystem 401), which is accessible by the subsystem 202 and the stage 203, and the stage 203 is configured to access the audio data and metadata as needed during the SBR and eSBR processing. The SBR processing in the stage 203 can be viewed as post-processing of the output of the core decoding subsystem 202. Optionally, the decoder 400 also includes a final upmixing subsystem (which can apply the parametric stereo ("PS") tools defined in the MPEG-4 AAC standard, using PS metadata extracted by the deformatter 215), which is coupled and configured to perform upmixing on the output of the stage 203 to generate fully decoded, upmixed audio that is output from the APU 210.

[0148] Parametric stereo is an encoding tool that represents a stereo signal using a downmix of the left and right channels of the stereo signal and a set of spatial parameters that describe the stereo image. Parametric stereo typically employs three types of spatial parameters: (1) an inter-channel intensity difference (IID) that describes the intensity difference between the channels; (2) an inter-channel phase difference (IPD) that describes the phase difference between the channels; and (3) an inter-channel coherence (ICC) that describes the coherence (or similarity) between the channels. Coherence can be measured as the maximum of the cross-correlation as a function of time or phase. These three parameters typically enable a high quality reconstruction of the stereo image. However, the IPD parameter only specifies the relative phase difference between the channels of the stereo input signal and does not indicate the distribution of these phase differences within the left and right channels. Therefore, a fourth type of parameter that describes the overall phase shift or overall phase difference (OPD) can additionally be used. In a stereo reconstruction procedure, consecutive windowed segments of both the received downmix signal s[n] and an uncorrelated version of the received downmix d[n] are processed together with the spatial parameters to generate left (l k (n)) and right (rk (n)) the reconstructed signal:

[0149] l k (n) = H 11 (k, n)s k (n) + H 21 (k, n)d k (n)

[0150] r k (n) = H 12 (k, n)s k (n) + H 22 (k, n)d k (n)

[0151] where H 11 , H 12 , H 21 and H 22 are defined by the stereo parameters. Finally the signals l k (n) and r k (n) are transformed back to the time domain by a frequency-to-time transform.

[0152] Figure 5 The control data generation subsystem 401 is coupled and configured to detect at least one property of the encoded audio bitstream to be decoded, and in response to at least one result of the detection step to generate eSBR control data (which can be or include any type of eSBR metadata included in the encoded audio bitstream according to other embodiments of the invention). The eSBR control data is asserted to stage 203 to trigger the application of individual eSBR tools or combinations of eSBR tools and / or to control the application of these eSBR tools after detecting a particular property (or combination of properties) of the bitstream. For example, in order to control the eSBR processing performed using harmonic transposition, some embodiments of the control data generation subsystem 401 will include: a music detector (e.g. a simplified version of a conventional music detector) for setting the sbrPatchingMode[ch] parameter (and asserting the setting of the parameter to stage 203) in response to detecting that the bitstream indicates or does not indicate music; a transient detector for setting the sbrOversamplingFlag[ch] parameter (and asserting the setting of the parameter to stage 203) in response to detecting the presence or lack of transients in the audio content indicated by the bitstream; and / or a pitch detector for setting the sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch] parameters (and asserting the setting of the parameters to stage 203) in response to detecting the pitch of the audio content indicated by the bitstream. Other aspects of the invention are audio bitstream decoding methods performed by any of the embodiments of the inventive decoder described in this paragraph and in the previous paragraph.

[0153] Aspects of the disclosure include methods of encoding or decoding of the type configured (e.g., programmed) to be performed by any embodiment of the inventive APU, system, or device. Other aspects of the disclosure include systems or devices configured (e.g., programmed) to perform any embodiment of the inventive method, and computer-readable media (e.g., magnetic disks) storing (e.g., in a non-transitory manner) program code for implementing any embodiment of the inventive method or steps thereof. For example, the inventive system can be or include any programmable general purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or otherwise configured to perform various operations on data, including embodiments of the inventive method or steps thereof. Such a general purpose processor can be or include a computer system including an input device, a memory, and processing circuitry programmed (and / or otherwise configured) to perform embodiments of the inventive method (or steps thereof) in response to being asserted data.

[0154] Embodiments of the disclosure can be implemented in hardware, firmware, or software, or a combination of the three (e.g., as programmable logic arrays). Unless otherwise indicated, algorithms or programs that are part of the disclosure are not inherently related to any particular computer or other apparatus. In particular, various general purpose machines can be used with programs written in accordance with the teachings herein, or it can be more convenient to construct more specialized apparatus (e.g., integrated circuits) to perform the required method steps. Thus, the disclosure can be implemented in an electronic digital computer, such as those of known design, including a working memory, a processor, at least one input device, and at least one output device, or Figure 1 any element of Figure 2 the encoder 100 (or elements thereof) of Figure 3 the decoder 200 (or elements thereof) of Figure 4 the decoder 210 (or elements thereof) of Figure 5 the decoder 400 (or elements thereof) of one or more computer programs executing on any element of the one or more programmable computer systems each comprising at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. The program code applies to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in the known fashion.

[0155]

[0155] Each such program can be implemented in any desired computer language (including machine, assembly, or high-level procedural, logical, or object-oriented programming languages) to communicate with a computer system. In any case, the language can be a compiled or interpreted language.

[0156] For example, when implemented by a sequence of computer software instructions, various functions and steps of embodiments of the application can be implemented by a multithreaded sequence of software instructions running in suitable digital signal processing hardware, in which case various apparatuses, steps, and functions of embodiments can correspond to portions of the software instructions.

[0157] Each such computer program is preferably stored upon a storage media or device (e.g., solid state memory or media, or magnetic or optical media) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer system, to perform the procedures described herein. The inventive system can also be implemented as a computer-readable storage medium, configured with (e.g., stored with) a computer program, where the storage medium so configured causes a computer system to operate in a specific and predefined manner to perform the functions described herein.

[0158] Several embodiments of the application have been described. However, it will be understood that various modifications can be made without departing from the spirit and scope of the application. Many modifications and variations of this application can be made in light of the above teachings. For example, to facilitate efficient implementation, phase shifting can be used in combination with complex QMF analysis and synthesis filter banks. The analysis filter bank is responsible for filtering the time-domain low-band signal produced by the core decoder into a plurality of sub-bands (e.g., QMF sub-bands). The synthesis filter bank is responsible for combining the regenerated high-band produced by the selected HFR technique (as dictated by the received sbrPatchingMode parameter) with the decoded low-band to produce a wide-band output audio signal. However, a given filter bank implementation operating in a particular sampling rate mode (e.g., normal double-rate operation or reduced sampling rate SBR mode) should not have bitstream dependent phase shifting. The QMF bank used in SBR is a complex indication extension of the theory of cosine modulation filter banks. It can be shown that when a complex exponential modulation is used to extend a cosine modulation filter bank, the aliasing cancellation constraints become obsolete. Thus, for the SBR QMF bank, both the analysis filter h k (n) and the synthesis filter f k (n) can be defined by:

[0159]

[0160] where p0(n) is a real-valued symmetric or asymmetric prototype filter (typically a low-pass prototype filter), M denotes the number of channels and N is the prototype filter order. The number of channels used in the analysis filter bank can be different from the number of channels used in the synthesis filter bank. For example, the analysis filter bank can have 32 channels and the synthesis filter bank can have 64 channels. When the synthesis filter bank is operated in a downsampled mode, the synthesis filter bank can only have 32 channels. Since the subband samples from the filter bank are complex-valued, an additional possible channel dependent phase shift step can be attached to the analysis filter bank. These additional phase shifts need to be compensated before the synthesis filter bank. While in principle the phase shift terms can have arbitrary values without impairing the operation of the QMF analysis / synthesis chain, they can also be constrained to certain values for consistency verification. The SBR signal will be affected by the selection of the phase factors, the low pass signal from the core decoder will not. The audio quality of the output signal will not be affected.

[0161] The coefficients of the prototype filter p0(n) can be defined to have a length L of 640 as shown in Table 4 below.

[0162] Table 4

[0163]

[0164]

[0165]

[0166]

[0167]

[0168]

[0169] The prototype filter p0(n) can also be derived from Table 4 by one or more mathematical operations (e.g. truncation, subsampling, interpolation and decimation).

[0170] While the tuning of SBR related control information generally does not depend on the details of the transposition (as discussed previously), in some embodiments certain elements of the eSBR extension container (bs_extension_id == EXTENSION_ID_ESBR) can be side- chained control data to improve the quality of the reproduced signal. Some side-chained elements can include noise floor data (e.g., a noise floor scale factor and a parameter indicating the direction of the delta coding for each noise floor (in the frequency direction or the time direction)), inverse filtering data (e.g., a parameter indicating the inverse filtering mode selected from no inverse filtering, low level inverse filtering, medium level inverse filtering, and strong level inverse filtering), and missing harmonics data (e.g., a parameter indicating whether a sinusoid should be added to a particular band of the reproduced high frequency band). All of these elements rely on the synthesis model of the transposer of the decoder performed in the encoder and thus can increase the quality of the reproduced signal if properly tuned for the selected transposer.

[0171] In particular, in some embodiments, missing harmonics and inverse filtering control data are transmitted in the eSBR extension container (along with the other bitstream parameters of Table 3) and tuned for the harmonic transposer of eSBR. The additional bitrate required to transmit this two categories of metadata for the harmonic transposer of eSBR is relatively low. Thus, sending tuned missing harmonics and / or inverse filtering control data in the eSBR extension container will increase the quality of the audio produced by the transposer with only a minimal impact on the bitrate. To ensure backwards compatibility with legacy decoders, the parameters tuning the spectral translation operation for SBR can also be sent in the bitstream as part of the SBR control data using implicit or explicit signaling.

[0172] It is understood that the present application can be practiced without the specific details set forth herein. Any element contained herein, in the following claims, is intended to be illustrative only and not restrictive on the claims. Various aspects of the present application will become apparent to those of ordinary skill in the art upon reading the following detailed description, taken in conjunction with the accompanying drawings, in which:

[0173] EEE 1. A method for performing high frequency reconstruction of an audio signal, the method comprising:

[0174] receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low frequency band portion of the audio signal and high frequency reconstruction metadata;

[0175] decoding the audio data to produce a decoded low frequency band audio signal;

[0176] extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata including operational parameters for a high frequency reconstruction process, the operational parameters including a patching mode parameter positioned in an extension container of the encoded audio bitstream, wherein a first value of the patching mode parameter indicates spectral warping and a second value of the patching mode parameter indicates harmonic transposition by phase vocoder frequency spreading;

[0177] filtering the decoded low frequency band audio signal to produce a filtered low frequency band audio signal;

[0178] reproducing a high frequency band portion of the audio signal using the filtered low frequency band audio signal and the high frequency reconstruction metadata, wherein if the patching mode parameter is the first value then the reproducing includes spectral warping and if the patching mode parameter is the second value then the reproducing includes harmonic transposition by phase vocoder frequency spreading; and

[0179] combining the filtered low frequency band audio signal with the reproduced high frequency band portion to form a wide frequency band audio signal.

[0180] EEE2. The method of EEE 1, wherein the extension container includes inverse filtering control data used when the patching mode parameter equals the second value.

[0181] EEE3. The method of any of EEEs 1-2, wherein the extension container further includes missing harmonic control data used when the patching mode parameter equals the second value.

[0182] EEE4. The method of any preceding EEE, wherein the encoded audio bitstream further includes a padding element having an identifier indicating a start of the padding element and padding data following the identifier, wherein the padding data includes the extension container.

[0183] EEE5. The method of EEE 4, wherein the identifier is a three-bit most significant bit first unsigned integer and has a value of 0x6.

[0184] EEE6. The method of EEE 4 or EEE 5, wherein the padding data includes an extension payload, the extension payload including spectral band replication extension data, and the extension payload is identified as having a four-bit most significant bit first unsigned integer and having a value of '1101' or '1110', and optionally,

[0185] wherein the spectral band replication extension data includes:

[0186] a spectral band replication header is selected,

[0187] spectral band replication data, following the header, and

[0188] a spectral band replication extension element, following the spectral band replication data, and wherein a flag is included in the spectral band replication extension element.

[0189] EEE7. The method according to any of EEEs 1 to 6, wherein the high frequency reconstruction metadata comprises envelope scale factors, noise floor scale factors, time / frequency grid information, or parameters indicating a crossover frequency.

[0190] EEE8. The method according to any of EEEs 1 to 7, wherein the filtering is performed by an analysis filter bank comprising analysis filters h k (n) that are modulated versions of a prototype filter p0(n).

[0191]

[0192] where p0(n) is a real-valued symmetric or asymmetric prototype filter, M is the number of channels in the analysis filter bank and N is the order of the prototype filter.

[0193] EEE9. The method according to EEE 8, wherein the prototype filter p0(n) is derived from the coefficients of Table 4 herein.

[0194] EEE10. The method according to EEE 8, wherein the prototype filter p0(n) is derived from the coefficients of Table 4 herein by one or more mathematical operations selected from the group consisting of truncation, sub-sampling, interpolation or decimation.

[0195] EEE11. The method according to any of EEEs 1 to 10, wherein a phase shift is added to the filtered low-band audio signal after the filtering, and the phase shift is compensated for before the combining to reduce the complexity of the method.

[0196] EEE12. The method according to any preceding EEE, wherein the extension container further comprises a flag indicating whether an additional pre-processing is used to avoid discontinuities in the shape of the spectral envelope of the high-band portion when the patching mode parameter is equal to the first value, wherein a first value of the flag enables the additional pre-processing and a second value of the flag disables the additional pre-processing.

[0197] EEE13. The method according to EEE 12, wherein the additional pre-processing comprises calculating a pre-gain curve using linear prediction filter coefficients.

[0198] EEE 14. The method according to any one of EEEs 1 to 13, wherein the extended container is a back-compatible extended container.

[0199] EEE 15. The method according to any one of EEEs 1 to 14, wherein the encoded audio stream is encoded according to a format, and wherein the extended container is an extended container defined in at least one older version of the format.

[0200] EEE 16. A non-transitory computer-readable medium containing instructions that, when executed by a processor, perform the method according to any one of EEEs 1 to 15.

[0201] EEE 17. An audio processing unit for performing high frequency reconstruction of an audio signal, the audio processing unit being configured to perform the method according to any one of EEEs 1 to 15.

Claims

1. A method for performing high frequency reconstruction of an audio signal, the method comprising: receiving an encoded audio bitstream comprising audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata; decoding the audio data to generate a decoded low-band audio signal; extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata comprising operating parameters for a high frequency reconstruction process, the operating parameters comprising a patch mode parameter located in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates spectral translation and a second value of the patch mode parameter indicates harmonic transposition by phase vocoder frequency extension, wherein the encoded audio bitstream further comprises a padding element having an identifier indicating a start of the padding element and padding data following the identifier, wherein the padding data comprises the backward compatible extension container, and wherein the identifier is an unsigned integer with three bits transmitted most significant bit first and having a value of 0x6, and wherein the padding data comprises an extension payload, the extension payload comprising spectral band replication extension data, and wherein the extension payload is identified with a four-bit unsigned integer with most significant bit first and having a value of '1101' or '1110'; filtering the decoded low-band audio signal to produce a filtered low-band audio signal; regenerating a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein the regeneration comprises spectral translation if the patch mode parameter is the first value, and comprises harmonic transposition by phase vocoder frequency extension if the patch mode parameter is the second value; and The filtered low-band audio signal and the regenerated high-band portion are combined to form a wide-band audio signal. 2 . The method of claim 1 , wherein the backward compatible extension container contains inverse filtering control data used when the patch mode parameter is equal to the second value. 3 . The method of claim 1 , wherein the backward compatible extension container further comprises missing harmonic control data for use when the repair mode parameter is equal to the second value.

4. The method of claim 1, wherein after the filtering, a phase shift is added to the filtered low-band audio signal, and before the combining, the phase shift is compensated to reduce the complexity of the method.

5. The method of claim 1 , wherein the retroactively compatible extension container further comprises a flag indicating whether additional preprocessing is used to avoid discontinuities in the shape of the spectrum envelope of the high-frequency band portion when the patch mode parameter is equal to the first value, wherein a first value of the flag enables the additional preprocessing and a second value of the flag disables the additional preprocessing. The method of claim 5 , wherein the additional pre-processing comprises computing a pre-gain curve using linear prediction filter coefficients.

7. A non-transitory computer readable medium containing instructions that, when executed by a processor, perform the method of claim 1.

8. An audio processing unit for performing high frequency reconstruction of an audio signal, the audio processing unit comprising: an input interface for receiving an encoded audio bitstream comprising audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata; a core audio decoder for decoding the audio data to generate a decoded low-band audio signal; a deformatter for extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata comprising operating parameters for a high frequency reconstruction process, the operating parameters comprising a padding element having an identifier indicating a start of the padding element and padding data following the identifier, wherein the padding data comprises a backward compatible extension container, the backward compatible extension container comprising a patch mode parameter, wherein a first value of the patch mode parameter indicates spectral translation and a second value of the patch mode parameter indicates harmonic transposition by phase vocoder frequency extension, and wherein the identifier is an unsigned integer with three bits transmitted most significant bit first and having a value of 0x6, and wherein the padding data comprises an extension payload, the extension payload comprising spectral band replication extension data, and the extension payload is identified as having a four-bit unsigned integer with the most significant bit transmitted first and having a value of '1101' or '1110'; an analysis filter bank for filtering the decoded low-band audio signal to generate a filtered low-band audio signal; a high frequency regenerator for reconstructing a high frequency band portion of the audio signal using the filtered low frequency band audio signal and the high frequency reconstruction metadata, wherein the reconstruction comprises spectral translation if the patch mode parameter is the first value, and comprises harmonic transposition by phase vocoder frequency extension if the patch mode parameter is the second value; and An analysis filter bank is used to combine the filtered low-band audio signal with the regenerated high-band portion to form a wide-band audio signal.

Citation Information

Patent Citations

  • Backward-compatible integration of high frequency reconstruction techniques for audio signals

    CN113990332A

  • Enhanced audio decoder

    CN102598121A

  • Methods for improving high frequency reconstruction

    US20050096917A1

  • Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element

    WO2016146492A1