Integration of high-frequency reconstruction techniques with reduced post-processing delays

By using analytical filter banks and high-frequency reconstruction metadata to process the audio signal, the shortcomings of spectrum band replication technology in low cross-frequency music content are solved, and the reconstruction quality and encoding efficiency of the audio signal are improved.

CN114242087BActive Publication Date: 2025-08-12DOLBY INTERNATIONAL AB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111585681.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-04-25
Filing Date
2019-04-25
Publication Date
2025-08-12
Estimated Expiration
2039-04-25

AI Technical Summary

Technical Problem

Existing spectrum band replication techniques are not effective when processing certain audio types, especially music content with relatively low crossover frequencies, resulting in limited audio quality.

Method used

The decoded low-band audio signals are filtered using an analysis filter bank, and spectrum translation or harmonic transposition is selectively performed based on high-frequency reconstruction metadata and markers, combining the low-band and regenerating the high-band to form a broadband audio signal.

Benefits of technology

The reconstruction quality of audio signals is improved, especially the processing effect of low cross-frequency music content, and the audio encoding efficiency and quality are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114242087B_ABST
    Figure CN114242087B_ABST
Patent Text Reader

Abstract

The present application relates to the integration of high-frequency reconstruction techniques with reduced post-processing delay, and specifically discloses a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream and decoding the audio data to produce a decoded low-band audio signal. The method further includes extracting high-frequency reconstruction metadata and filtering the decoded low-band audio signal using an analysis filter bank to produce a filtered low-band audio signal. The method also includes extracting a flag indicating whether spectral translation or harmonic transposition is performed on the audio data and regenerating a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata based on the flag. The high-frequency regeneration is performed as a post-processing operation with a delay of 3010 samples per audio channel.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Information about divisional applications

[0002] This application is a divisional application. The parent application is an invention patent application filed on April 25, 2019, with application number 201980034811.4 and the title of invention being “Integration of High-Frequency Reconstruction Technology with Reduced Post-Processing Delay.”

[0003] Cross-reference to related applications

[0004] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 662,296, filed April 25, 2018, the entire contents of which are incorporated herein by reference. Technical Field

[0005] Embodiments relate to audio signal processing, and more particularly, to encoding, decoding, or transcoding an audio bitstream using control data specifying a base form of high frequency reconstruction ("HFR") or an enhanced form of HFR to be performed on the audio data. Background Art

[0006] A typical audio bitstream includes both audio data (e.g., encoded audio data) indicating one or more channels of audio content and metadata indicating at least one characteristic of the audio data or audio content. One well-known format for generating encoded audio bitstreams is the MPEG-4 Advanced Audio Coding (AAC) format described in the MPEG standard ISO / IEC 14496-3:2009. In the MPEG-4 standard, AAC stands for "Advanced Audio Coding" and HE-AAC stands for "High Efficiency Advanced Audio Coding."

[0007] The MPEG-4 AAC standard defines several audio profiles that determine which objects and coding tools are present in a compatible encoder or decoder. Three of these audio profiles are (1) the AAC profile, (2) the HE-AAC profile, and (3) the HE-AAC v2 profile. The AAC profile includes the AAC Low Complexity (or "AAC-LC") object type. The AAC-LC object is the counterpart of the MPEG-2 AAC Low Complexity profile, with some adjustments, and does not include both the Spectral Band Replication ("SBR") object type and the Parametric Stereo ("PS") object type. The HE-AAC profile is a superset of the AAC profile and additionally includes the SBR object type. The HE-AAC v2 profile is a superset of the HE-AAC profile and additionally includes the PS object type.

[0008] The SBR object type contains a spectral band replication tool, which is an important high frequency reconstruction ("HFR") coding tool that can significantly improve the compression efficiency of perceptual audio codecs. SBR reconstructs the high-frequency components of the audio signal on the receiver side (e.g., in the decoder). As a result, the encoder only needs to encode and transmit the low-frequency components to allow for much higher audio quality at low data rates. SBR replicates a harmonic sequence that was previously truncated to reduce the data rate based on the available bandwidth-limited signal and control data obtained from the encoder. The ratio between tonal components and noise-like components is maintained by adaptive inverse filtering and, optionally, the addition of noise and sinusoids. In the MPEG-4 AAC standard, the SBR tool performs spectral patching (also known as linear translation or spectral panning), in which several consecutive quadrature mirror filter (QMF) subbands are copied (or "patched") from the transmitted low-band portion of the audio signal to the high-band portion of the audio signal (which is generated in the decoder).

[0009] Spectral patching or linear panning may not be suitable for certain audio types (eg, music content with relatively low crossover frequencies).Thus, techniques for improving spectral band replication are needed. Summary of the Invention

[0010] A first class of embodiments relates to a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream and decoding the audio data to produce a decoded low-band audio signal. The method further includes extracting high-frequency reconstruction metadata and filtering the decoded low-band audio signal using an analysis filter bank to produce a filtered low-band audio signal. The method further includes extracting a flag indicating whether spectral translation or harmonic transposition is performed on the audio data and regenerating a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata based on the flag. Finally, the method includes combining the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal.

[0011] A second class of embodiments relates to an audio decoder for decoding an encoded audio bitstream. The decoder comprises an input interface for receiving the encoded audio bitstream, wherein the encoded audio bitstream comprises audio data representing a low-band portion of an audio signal; and a core decoder for decoding the audio data to produce a decoded low-band audio signal. The decoder also comprises a demultiplexer for extracting high-frequency reconstruction metadata from the encoded audio bitstream, wherein the high-frequency reconstruction metadata comprises operating parameters for a high-frequency reconstruction process that linearly shifts a number of consecutive frequency subbands from the low-band portion of the audio signal to the high-band portion of the audio signal; and an analysis filter bank for filtering the decoded low-band audio signal to produce a filtered low-band audio signal. The decoder further comprises a demultiplexer for extracting a flag from the encoded audio bitstream indicating whether linear shifting or harmonic transposition is performed on the audio data; and a high-frequency regenerator for regenerating the high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata according to the flag. Finally, the decoder includes a synthesis filterbank for combining the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal.

[0012] Other classes of embodiments relate to encoding and transcoding audio bitstreams that contain metadata identifying whether enhanced spectral band replication (eSBR) processing is performed. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a block diagram of an embodiment of a system that may be configured to perform embodiments of the inventive method.

[0014] Figure 2 is a block diagram of an encoder that is an embodiment of the inventive audio processing unit.

[0015] Figure 3 is a block diagram of a system including a decoder, which is an embodiment of the inventive audio processing unit, and optionally also including a post-processor coupled to the decoder.

[0016] Figure 4 is a block diagram of a decoder, which is an embodiment of the inventive audio processing unit.

[0017] Figure 5 is a block diagram of a decoder, which is another embodiment of the inventive audio processing unit.

[0018] Figure 6 is a block diagram of another embodiment of the inventive audio processing unit.

[0019] Figure 7A block diagram of an MPEG-4 AAC bitstream, showing its division into several segments.

[0020] Symbols and terminology

[0021] In this disclosure (including in the claims), the expression "performing an operation on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to a signal or data) is used broadly to mean performing the operation directly on the signal or data or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or preprocessing before the operation is performed on it).

[0022] In this disclosure (including in the claims), the expression "audio processing unit" or "audio processor" is used in a broad sense to refer to a system, device, or apparatus configured to process audio data. Examples of audio processing units include, but are not limited to, encoders, transcoders, decoders, codecs, pre-processing systems, post-processing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools). Almost all consumer electronic products, such as mobile phones, televisions, laptops, and tablet computers, contain audio processing units or audio processors.

[0023] Throughout this disclosure (including in the claims), the terms "coupled" or "coupled" are used broadly to refer to both direct and indirect connections. Thus, if a first device is coupled to a second device, that connection may be through a direct connection or through an indirect connection via other devices and connections. Additionally, components that are integrated into or with other components are also coupled to each other. DETAILED DESCRIPTION

[0024] The MPEG-4 AAC standard contemplates that an encoded MPEG-4 AAC bitstream include metadata that indicates each type of high frequency reconstruction ("HFR") processing applied (if any) by a decoder to decode the audio content of the bitstream, and / or controls such HFR processing, and / or indicates at least one characteristic or parameter of at least one HFR tool used to decode the audio content of the bitstream. Herein, we use the expression "SBR metadata" to refer to this type of metadata for use with spectral band replication ("SBR"), as described or referred to in the MPEG-4 AAC standard. Those skilled in the art will appreciate that SBR is a form of HFR.

[0025] SBR is preferably used as a dual-rate system, in which the basic codec operates at half the original sampling rate, while SBR operates at the original sampling rate. Despite having a higher sampling rate, the SBR encoder works in parallel with the basic core codec. Although SBR is primarily a post-processing in the decoder, important parameters are extracted in the encoder to ensure the most accurate high-frequency reconstruction in the decoder. The encoder estimates the spectral envelope of the SBR range that is suitable for the time and frequency range / resolution of the current input signal segment characteristics. The spectral envelope is estimated by complex QMF analysis and subsequent energy calculation. The time and frequency resolution of the spectral envelope can be selected with a high degree of freedom to ensure the most suitable time-frequency resolution for a given input region segment. Envelope estimation needs to take into account that, before envelope adjustment, transients of the original source, which are mainly located in the high-frequency region (e.g., high-hat), will appear to a lesser extent in the high-frequency band generated by SBR, because the high-frequency band in the decoder is based on a low-frequency band where transients are much less obvious than in the high-frequency band. Compared with the general spectral envelope estimation used in other audio coding algorithms, this aspect places different requirements on the time-frequency resolution of the spectral envelope data.

[0026] In addition to the spectrum envelope, several additional parameters representing the spectrum characteristics of the input signal in different time and frequency regions are also extracted. Since the encoder naturally has the right to access the original signal and the information about how the SBR unit in the decoder will produce the high-frequency band, in view of a specific set of control parameters, the system can handle situations in which the low-frequency band constitutes a strong harmonic series and the high-frequency band that will be regenerated mainly constitutes a random signal component, and situations in which strong tonal components are present in the original high-frequency band and do not have a counterpart in the low-frequency band (the high-frequency band region is based on this). In addition, the SBR encoder works closely with the basic core codec to assess which frequency range should be covered by SBR at a given time. With regard to stereo signals, the SBR data are efficiently encoded before transmission by utilizing entropy coding and the channel dependency of control data.

[0027] The control parameter extraction algorithm usually needs to be carefully tuned at a given bitrate and a given sampling rate according to the base codec. This is due to the fact that lower bitrates usually mean a larger SBR range than higher bitrates and different sampling rates correspond to different temporal resolutions of the SBR frame.

[0028] An SBR decoder typically includes several different parts. It includes a bitstream decoding module, a high-frequency reconstruction (HFR) module, an additional high-frequency component module, and an envelope adjuster module. The system is based on a complex-valued QMF filter bank (for high-quality SBR) or a real-valued QMF filter bank (for low-power SBR). Embodiments of the present invention are applicable to both high-quality SBR and low-power SBR. In the bitstream extraction module, control data is read and decoded from the bitstream. Before reading envelope data from the bitstream, the time-frequency grid of the current frame is obtained. The basic core decoder decodes the audio signal of the current frame (although at a lower sampling rate) to generate time-domain audio samples. The resulting frame of audio data is used by the HFR module to perform high-frequency reconstruction. Next, a QMF filter bank is used to analyze the decoded low-band signal. Subsequently, high-frequency reconstruction and envelope adjustment are performed on the sub-band samples of the QMF filter bank. Based on given control parameters, the high frequency is reconstructed from the low-band in a flexible manner. In addition, based on the control data, the reconstructed high-frequency band is adaptively filtered based on the sub-band channels to ensure appropriate spectral characteristics for a given time / frequency region.

[0029] The top layer of an MPEG-4 AAC bitstream is a sequence of data blocks ("raw_data_block" elements), each of which is a data segment (referred to herein as a "block") containing audio data (typically within a period of 1024 or 960 samples) and related information and / or other data. Herein, we use the term "block" to refer to a segment of an MPEG-4 AAC bitstream that includes audio data (and corresponding metadata and optionally other related data) that identifies or indicates one (but not more than one) "raw_data_block" element.

[0030] Each block of an MPEG-4 AAC bitstream may include several syntax elements (each of which is also embodied as a data segment in the bitstream). Seven types of these syntax elements are defined in the MPEG-4 AAC standard. Each syntax element is identified by a different value of the data element "id_syn_ele". Examples of syntax elements include "single_channel_element()", "channel_pair_element()", and "fill_element()". A single channel element is a container that contains audio data for a single audio channel (a mono audio signal). A channel pair element contains audio data for two audio channels (i.e., a stereo audio signal).

[0031] A padding element is an information container consisting of an identifier (e.g., the value of the "id_syn_ele" element described above) followed by data (referred to as "padding data"). Padding elements are traditionally used to adjust the instantaneous bit rate of a bitstream transmitted over a constant-rate channel. A constant data rate can be achieved by adding an appropriate amount of padding data to each block.

[0032] According to embodiments of the present invention, padding data may include one or more extension payloads that expand the types of data (e.g., metadata) that can be transmitted in a bitstream. A decoder that receives a bitstream with padding data containing new data types may optionally be used by a device receiving the bitstream (e.g., a decoder) to expand the functionality of the device. Therefore, those skilled in the art will appreciate that padding elements are a special type of data structure and differ from data structures typically used to transmit audio data (e.g., audio payloads containing channel data).

[0033] In some embodiments of the present invention, the identifier used to identify the padding element may consist of a 3-bit unsigned integer ("uimsbf") with the most significant bit transmitted first and a value of 0x6. In a block, several instances of the same type of syntax element (e.g., several padding elements) may appear.

[0034] Another standard for encoding audio bitstreams is the MPEG Unified Speech and Audio Coding (USAC) standard (ISO / IEC 23003-3:2012). The MPEG USAC standard describes the use of a spectral band replication process (including the SBR process described in the MPEG-4 AAC standard and also other enhancements to the spectral band replication process) to encode and decode audio content. This process applies the spectral band replication tools (sometimes referred to herein as "enhanced SBR tools" or "eSBR tools"), which are an extension and enhancement of the SBR toolset described in the MPEG-4 AAC standard. Thus, eSBR (as defined in the USAC standard) is an improvement over SBR (as defined in the MPEG-4 AAC standard).

[0035] In this document, we use the expression "enhanced SBR processing" (or "eSBR processing") to refer to spectral band replication processing that uses at least one eSBR tool not described or mentioned in the MPEG-4 AAC standard, such as at least one eSBR tool described or mentioned in the MPEG USAC standard. Examples of these eSBR tools are harmonic transposition and QMF patching additional pre-processing or "pre-flattening."

[0036] A harmonic transposer of integer order T maps a sinusoid with frequency ω into a sinusoid with frequency Tω while preserving signal duration. Three orders, T = 2, 3, and 4, are typically used in sequence to generate each portion of the desired output frequency range using the smallest possible transposition order. If an output higher than the 4th order transposition range is required, it can be generated by frequency shifting. Processing is performed in the fundamental frequency time domain as close to critical sampling as possible to minimize computational complexity.

[0037] The harmonic transposer can be based on QMF or DFT. When using a QMF-based harmonic transposer, bandwidth extension of the core encoder time-domain signal is fully implemented in the QMF domain using a modified phase vocoder structure to perform decimation and subsequent time stretching for each QMF subband. Transposition using several transposition factors (e.g., T = 2, 3, 4) is implemented in a common QMF analysis / synthesis transform stage. Since the QMF-based harmonic transposer does not feature signal-adaptive frequency-domain oversampling, the corresponding flag in the bitstream (sbrOversamplingFlag[ch]) can be ignored.

[0038] When using a DFT-based harmonic transposer, the factor 3 and 4 transposers (3rd and 4th order transposers) are preferably integrated into the factor 2 transposer (2nd order converter) by interpolation to reduce complexity. For each frame (corresponding to coreCoderFrameLength core encoder samples), the nominal "full scale" transform size of the transposer is first determined by the signal adaptive frequency domain oversampling flag (sbrOversamplingFlag[ch]) in the bitstream.

[0039] When sbrPatchingMode==1 to indicate that linear transposition will be used to generate the high frequency band, an additional step can be introduced to avoid the shape discontinuity of the spectrum envelope of the high frequency signal being input to the subsequent envelope adjuster. This improves the operation of the subsequent envelope adjustment stage to result in a high frequency band signal that is perceived as more stable. The operation of the additional preprocessing is beneficial to the signal type in which the rough spectrum envelope of the low frequency band signal used for high frequency reconstruction shows a large level of variation. However, the value of the bit stream element can be determined in the encoder by applying any type of signal-dependent classification. Preferably, additional preprocessing is started by a 1-bit bit stream element bs_sbr_preprocessing. When bs_sbr_preprocessing is set to 1, additional processing is enabled. When bs_sbr_preprocessing is set to 0, additional preprocessing is disabled. The additional processing preferably utilizes the pre-gain curve used by the high frequency generator to proportionally adjust the low frequency band X of each patching. Low For example, the pre-gain curve can be calculated according to the following equation:

[0040] preGain(k)=10 (meanNrg-lowEnvSlope(k)) / 20 , 0≤k<k0

[0041] where k0 is the first QMF subband in the mainband table and lowEnvSlope is calculated using a function that calculates the coefficients of a best-fit polynomial (in the least squares sense), such as polyfit(). For example, (using a cubic polynomial)

[0042] polyfit(3,k0,x_lowband,lowEnv,lowEnvSlope);

[0043] And among them

[0044]

[0045] where x_lowband(k)=[0...k0-1], numTimeSlot is the number of SBR envelope time slots present in the frame, RATE is a constant indicating the number of QMF subband samples per time slot (e.g., 2), are the linear prediction filter coefficients (obtained from the covariance method) and where

[0046]

[0047] A bitstream produced in accordance with the MPEG USAC standard (sometimes referred to herein as a "USAC bitstream") includes encoded audio content and typically includes metadata indicating each type of spectral band replication processing applied by a decoder to decode the audio content of the USAC bitstream and / or metadata that controls such spectral band replication processing and / or indicates at least one characteristic or parameter of at least one SBR tool and / or eSBR tool used to decode the audio content of the USAC bitstream.

[0048] In this document, we use the expression "enhanced SBR metadata" (or "eSBR metadata") to refer to metadata that indicates each type of spectral band replication process applied by a decoder to decode the audio content of an encoded audio bitstream (e.g., a USAC bitstream) and / or controls such spectral band replication process and / or indicates at least one characteristic or parameter of at least one SBR tool and / or eSBR tool used for decoding such audio content but not described or mentioned in the MPEG-4 AAC standard. An example of eSBR metadata is metadata (indicating or used for controlling spectral band replication processes) that is described or mentioned in the MPEG USAC standard but not described or mentioned in the MPEG-4 AAC standard. Thus, eSBR metadata in this document refers to metadata that is not SBR metadata, and SBR metadata in this document refers to metadata that is not eSBR metadata.

[0049] A USAC bitstream may include both SBR metadata and eSBR metadata. More specifically, a USAC bitstream may include both eSBR metadata that controls eSBR processing performed by a decoder and SBR metadata that controls SBR processing performed by a decoder. According to an exemplary embodiment of the present invention, eSBR metadata (e.g., eSBR-specific configuration data) is included (according to the present invention) in an MPEG-4 AAC bitstream (e.g., in an sbr_extension() container at the end of the SBR payload).

[0050] During decoding of an encoded bitstream using an eSBR toolset (including at least one eSBR tool), an eSBR process is performed by the decoder to regenerate the high-frequency band of the audio signal based on a replica of the harmonic sequence that was truncated during encoding. This eSBR process typically adjusts the spectral envelope of the generated high-frequency band and applies inverse filtering, and adds noise and sinusoidal components to recreate the spectral characteristics of the original audio signal.

[0051] According to an exemplary embodiment of the present invention, eSBR metadata (e.g., a small amount of control bits of eSBR metadata) is included in one or more metadata sections of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that also includes encoded audio data in other sections (audio data sections). Typically, at least one such metadata section for each block of the bitstream is (or includes) a filler element (including an identifier indicating the start of the filler element), and the eSBR metadata is included in the filler element following the identifier.

[0052] Figure 1 FIG1 is a block diagram of an exemplary audio processing chain (audio data processing system) in which one or more elements of the system may be configured according to embodiments of the present invention. The system includes the following elements coupled together as shown: an encoder 1, a transport subsystem 2, a decoder 3, and a post-processing unit 4. In variations of the system shown, one or more elements may be omitted, or additional audio data processing units may be included.

[0053] In some implementations, the encoder 1 (which optionally includes a pre-processing unit) is configured to accept PCM (time domain) samples including audio content as input and output an encoded audio bitstream (having a format that conforms to the MPEG-4 AAC standard) indicating the audio content. The data indicating the bitstream of audio content is sometimes referred to herein as "audio data" or "encoded audio data." If the encoder is configured according to an exemplary embodiment of the present invention, the audio bitstream output from the encoder includes eSBR metadata (and typically other metadata as well) as the audio data.

[0054] One or more encoded audio bitstreams output from encoder 1 may be asserted to an encoded audio transmission subsystem 2. Subsystem 2 is configured to store and / or transmit each encoded bitstream output from encoder 1. The encoded audio bitstreams output from encoder 1 may be stored by subsystem 2 (e.g., in the form of a DVD or Blu-ray disc), or transmitted by subsystem 2 (which may implement a transmission link or network), or may be stored and transmitted by subsystem 2.

[0055] Decoder 3 is configured to decode an encoded MPEG-4 AAC audio bitstream (generated by encoder 1) that it receives via subsystem 2. In some embodiments, decoder 3 is configured to extract eSBR metadata from each block of the bitstream and decode the bitstream (including by performing eSBR processing using the extracted eSBR metadata) to produce decoded audio data (e.g., a stream of decoded PCM audio samples). In some embodiments, decoder 3 is configured to extract SBR metadata from the bitstream (but ignore the eSBR metadata included in the bitstream) and decode the bitstream (including by performing SBR processing using the extracted SBR metadata) to produce decoded audio data (e.g., a stream of decoded PCM audio samples). Typically, decoder 3 includes a buffer that stores (e.g., in a non-transitory manner) segments of the encoded audio bitstream received from subsystem 2.

[0056] Figure 1 The post-processing unit 4 is configured to accept the decoded audio data stream (e.g., decoded PCM audio samples) from the decoder 3 and perform post-processing on it. The post-processing unit may also be configured to render the post-processed audio content (or the decoded audio received from the decoder 3) for playback on one or more speakers.

[0057] Figure 2 is a block diagram of encoder 100, which is an embodiment of the inventive audio processing unit. Any component or element of encoder 100 can be implemented in hardware, software, or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., an ASIC, FPGA, or other integrated circuit). Encoder 100 includes encoder 105, stuffer / formatter stage 107, metadata generation stage 106, and buffer memory 109, connected as shown. Typically, encoder 100 also includes other processing elements (not shown). Encoder 100 is configured to convert an input audio bitstream into an encoded output MPEG-4 AAC bitstream.

[0058] Metadata generator 106 is coupled and configured to generate metadata (including eSBR metadata and SBR metadata) (and / or pass to stage 107 ) for inclusion by stage 107 in the encoded bitstream output from encoder 100 .

[0059] Encoder 105 is coupled and configured to encode input audio data (eg, by performing compression thereon) and assert the resulting encoded audio to stage 107 for inclusion in the encoded bitstream output from stage 107 .

[0060] Stage 107 is configured to multiplex the encoded audio from encoder 105 and metadata (including eSBR metadata and SBR metadata) from generator 106 to produce an encoded bitstream output from stage 107, preferably such that the encoded bitstream has a format specified by one embodiment of the present invention.

[0061] Buffer memory 109 is configured to store (e.g., in a non-transitory manner) at least one block of the encoded audio bitstream output from stage 107 and then assert a sequence of blocks of the encoded audio bitstream from buffer memory 109 as output from encoder 100 to a transmission system.

[0062] Figure 3 is a block diagram of a system including a decoder 200 (which is an embodiment of the inventive audio processing unit) and optionally also including a post-processor 300 coupled to the decoder 200. Any components or elements of the decoder 200 and the post-processor 300 can be implemented as one or more processes and / or one or more circuits (e.g., ASICs, FPGAs, or other integrated circuits) in hardware, software, or a combination of hardware and software. The decoder 200 includes a buffer memory 201, a bitstream payload deformatter (parser) 205, an audio decoding subsystem 202 (sometimes referred to as a "core" decoding stage or "core" decoding subsystem), an eSBR processing stage 203, and a control bit generation stage 204, connected as shown. Typically, the decoder 200 also includes other processing elements (not shown).

[0063] Buffer memory (buffer) 201 stores (eg, in a non-transitory manner) at least one block of an encoded MPEG-4 AAC audio bitstream received by decoder 200. In operation of decoder 200, a sequence of blocks of the bitstream is asserted from buffer 201 to deformatter 205.

[0064] exist Figure 3 Example (or to be described Figure 4 In a variation of the embodiment), the APU (which is not a decoder) (e.g. Figure 6 APU 500 includes a buffer memory (e.g., the same as buffer 201) that stores (e.g., in a non-transitory manner) Figure 3 or Figure 4 The buffer 201 receives at least one block of the same type of encoded audio bitstream (eg, an MPEG-4 AAC audio bitstream) (ie, an encoded audio bitstream containing eSBR metadata).

[0065] Reference again Figure 3 , the deformatter 205 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (including quantization envelope data) and eSBR metadata (and typically other metadata as well) therefrom to assert at least the eSBR metadata and SBR metadata to the eSBR processing stage 203 and typically also assert the other extracted metadata to the decoding subsystem 202 (and optionally also to the control bit generator 204). The deformatter 205 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.

[0066] Figure 3 The system also optionally includes a post-processor 300. Post-processor 300 includes a buffer memory (buffer) 301 and other processing elements (not shown), including at least one processing element coupled to buffer 301. Buffer 301 stores (e.g., in a non-transitory manner) at least one block (or frame) of decoded audio data received by post-processor 300 from decoder 200. The processing elements of post-processor 300 are coupled and configured to receive and use metadata output from decoding subsystem 202 (and / or deformatter 205) and / or control bits output from stage 204 of decoder 200 to adaptively process the sequence of blocks (or frames) of decoded audio output from buffer 301.

[0067] The audio decoding subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the parser 205 (this decoding may be referred to as the "core" decoding operation) to produce decoded audio data and to pass the decoded audio data to the eSBR processing stage 203. Decoding is performed in the frequency domain and typically includes inverse quantization followed by spectral processing. Typically, the final processing stage in the subsystem 202 applies a frequency-domain to time-domain transform to the decoded frequency-domain audio data, resulting in the subsystem output being time-domain decoded audio data. Stage 203 is configured to apply the SBR tools and eSBR tools indicated by the eSBR metadata and eSBR (extracted by the parser 205) to the decoded audio data (i.e., perform SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata) to produce fully decoded audio data that is output from the decoder 200 (e.g., to the post-processor 300). Typically, decoder 200 includes memory (accessible by subsystem 202 and stage 203) that stores deformatted audio data and metadata output from deformatter 205, and stage 203 is configured to access audio data and metadata (including SBR metadata and eSBR metadata) as needed during SBR and eSBR processing. The SBR and eSBR processing in stage 203 can be considered post-processing of the output of core decoding subsystem 202. Decoder 200 also optionally includes a final upmix subsystem (which can use PS metadata extracted by deformatter 205 and / or control bits generated in subsystem 204 to apply parametric stereo ("PS") tools defined in the MPEG-4 AAC standard) coupled and configured to perform upmixing on the output of stage 203 to produce fully decoded upmixed audio output from decoder 200. Alternatively, post-processor 300 is configured to perform upmixing on the output of decoder 200 (eg, using PS metadata extracted by deformatter 205 and / or control bits generated in subsystem 204).

[0068] In response to metadata extracted by the deformatter 205, the control bit generator 204 may generate control data, and the control data may be used within the decoder 200 (e.g., for use in the final upmix subsystem) and / or asserted as an output of the decoder 200 (e.g., to the post-processor 300 for post-processing). In response to metadata extracted from the input bitstream (and optionally also in response to the control data), the stage 204 may generate a control bit (and assert the control bit to the post-processor 300) to indicate that the decoded audio data output from the eSBR processing stage 203 should undergo a particular type of post-processing. In some implementations, the decoder 200 is configured to assert the metadata extracted from the input bitstream by the deformatter 205 to the post-processor 300, and the post-processor 300 is configured to perform post-processing on the decoded audio data output from the decoder 200 using the metadata.

[0069] Figure 4 FIG2 is a block diagram of an audio processing unit ("APU") 210, which is another embodiment of the inventive audio processing unit. APU 210 is a conventional decoder that is not configured to perform eSBR processing. Any component or element of APU 210 may be implemented as one or more processes and / or one or more circuits (e.g., an ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware and software. APU 210 includes a buffer memory 201, a bitstream payload deformatter (parser) 215, an audio decoding subsystem 202 (sometimes referred to as a "core" decoding stage or "core" decoding subsystem), and an SBR processing stage 213, connected as shown. Typically, APU 210 also includes other processing elements (not shown). APU 210 may represent, for example, an audio encoder, decoder, or transcoder.

[0070] The components 201 and 202 of the APU 210 are identical to ( Figure 3 ) decoder 200, and their above description will not be repeated. In operation of the APU 210, a block sequence of an encoded audio bitstream (MPEG-4 AAC bitstream) received by the APU 210 is asserted from the buffer 201 to the deformatter 215.

[0071] The deformatter 215 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (including quantization envelope data) therefrom and typically also extract other metadata therefrom, but ignoring eSBR metadata that may be included in the bitstream according to any embodiment of the present invention. The deformatter 215 is configured to assert at least the SBR metadata to the SBR processing stage 213. The deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.

[0072] The audio decoding subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the deformatter 215 (this decoding may be referred to as a "core" decoding operation) to produce decoded audio data and to pass the decoded audio data to the SBR processing stage 213. Decoding is performed in the frequency domain. Typically, the final processing stage in the subsystem 202 applies a frequency-domain to time-domain transform to the decoded frequency-domain audio data, so that the output of the subsystem is time-domain decoded audio data. Stage 213 is configured to apply SBR tools (but not eSBR tools) indicated by the SBR metadata (extracted by the deformatter 215) to the decoded audio data (i.e., use the SBR metadata to perform SBR processing on the output of the decoding subsystem 202) to produce fully decoded audio data that is output from the APU 210 (e.g., to the post-processor 300). Typically, APU 210 includes memory (accessible by subsystem 202 and stage 213) to store deformatted audio data and metadata output from deformatter 215, and stage 213 is configured to access audio data and metadata (including SBR metadata) as needed during SBR processing. The SBR processing in stage 213 can be considered post-processing of the output of core decoding subsystem 202. APU 210 also optionally includes a final upmix subsystem (which can use the PS metadata extracted by deformatter 215 to apply the parametric stereo "PS" tool defined in the MPEG-4 AAC standard) coupled and configured to perform upmixing on the output of stage 213 to produce fully decoded upmixed audio output from APU 210. Alternatively, a post-processor is configured to perform upmixing on the output of APU 210 (e.g., using PS metadata extracted by deformatter 215 and / or control bits generated in APU 210).

[0073] Various implementations of encoder 100, decoder 200, and APU 210 are configured to perform different embodiments of the inventive method.

[0074] According to some embodiments, eSBR metadata (e.g., a small number of control bits that are eSBR metadata) is included in an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) so that legacy decoders (which are not configured to parse the eSBR metadata or use any eSBR tools related to the eSBR metadata) can ignore the eSBR metadata but still decode the bitstream as best they can without using the eSBR metadata or any eSBR tools related to the eSBR metadata, typically without a significant loss in decoded audio quality. However, eSBR decoders (which are configured to parse the bitstream to identify the eSBR metadata and use at least one eSBR tool in response to the eSBR metadata) will benefit from using at least one such eSBR tool. Thus, embodiments of the present invention provide methods for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner.

[0075] Typically, the eSBR metadata in a bitstream indicates (e.g., indicates at least one characteristic or parameter thereof) one or more of the following eSBR tools (which are described in the MPEG USAC standard and may or may not be applied by the encoder during generation of the bitstream):

[0076] Harmonic transpose; and

[0077] QMF patching with additional pre-processing (pre-flattening).

[0078] For example, the eSBR metadata included in the bitstream may indicate the values of the parameters (as described in the MPEG USAC standard and this disclosure): sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch], and bs_sbr_preprocessing.

[0079] In this document, the notation X[ch] (where X is a parameter) indicates that the parameter is related to a channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, we sometimes omit the expression [ch] and assume that the relevant parameter is related to the channel of the audio content.

[0080] Herein, the notation X[ch][env] (where X is a parameter) indicates that the parameter is related to the SBR envelope ("env") of a channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, we sometimes omit the notation [env] and [ch] and assume that the relevant parameter is related to the SBR envelope of the channel of the audio content.

[0081] During decoding of the encoded bitstream, harmonic transposition is performed during the eSBR processing stage of the decoding (for each channel "ch" of the audio content indicated by the bitstream) controlled by the following eSBR metadata parameters: sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch].

[0082] The value "sbrPatchingMode[ch]" indicates the transpose type used in eSBR: sbrPatchingMode[ch]=1 indicates linear transposition patching described in section 4.6.18 of the MPEG-4 AAC standard (used with high-quality SBR or low-power SBR); sbrPatchingMode[ch]=0 indicates harmonic SBR patching described in sections 7.5.3 or 7.5.4 of the MPEG USAC standard.

[0083] The value "sbrOversamplingFlag[ch]" indicates that signal-adaptive frequency-domain oversampling in eSBR is used in combination with the DFT-based harmonic SBR patching described in section 7.5.3 of the MPEG USAC standard. This flag controls the size of the DFT used in the transposer: 1 indicates that signal-adaptive frequency-domain oversampling is enabled as described in section 7.5.3.1 of the MPEG USAC standard; 0 indicates that signal-adaptive frequency-domain oversampling is disabled as described in section 7.5.3.1 of the MPEG USAC standard.

[0084] The value "sbrPitchInBinsFlag[ch]" controls the interpretation of the sbrPitchInBins[ch] parameter: 1 indicates that the value of sbrPitchInBins[ch] is valid and greater than 0; 0 indicates that the value of sbrPitchInBins[ch] is set to 0.

[0085] The value "sbrPitchInBins[ch]" controls the addition of cross-product terms in the SBR harmonic transposer. The value sbrPitchinBins[ch] is an integer value in the range [0,127] and represents the distance measured in the frequency bins of the 1536-line DFT applied to the core encoder's sampling frequency.

[0086] If the MPEG-4 AAC bitstream indicates SBR channel pairs whose channels are uncoupled (rather than a single SBR channel), then the bitstream indicates two instances of the above syntax (for harmonic or non-harmonic transposition), one instance of sbr_channel_pair_element() for each channel.

[0087] The harmonic transposition of the eSBR tool generally improves the quality of the decoded music signal at relatively low crossover frequencies. Non-harmonic transposition (i.e., traditional spectral patching) generally improves speech signals. Therefore, the starting point for deciding which type of transposition is preferred for encoding specific audio content is to select the transposition method based on speech / music detection, where harmonic transposition is used for music content and spectral patching is used for tempo content.

[0088] The performance of pre-flattening during eSBR processing is controlled by the value of a 1-bit eSBR metadata parameter called "bs_sbr_preprocessing", in the sense that pre-flattening is performed or not performed depending on the value of this single bit. When using the SBR QMF patching algorithm described in section 4.6.18.6.3 of the MPEG-4 AAC standard, a pre-flattening step may be performed (when indicated by the "bs_sbr_preprocessing" parameter) to attempt to avoid discontinuities in the shape of the spectral envelope of the high-frequency signal input to the subsequent envelope adjuster (the envelope adjuster performs another stage of eSBR processing). Pre-flattening generally improves the operation of the subsequent envelope adjustment stage, resulting in a high-band signal that is perceived as more stable.

[0089] According to some embodiments of the present invention, the total bit rate requirement included in the MPEG-4 AAC bitstream eSBR metadata indicating the above-mentioned eSBR tools (harmonic transposition and pre-flattening) is expected to be on the order of hundreds of bits per second, since only the differential control data required to perform the eSBR process is transmitted. Legacy decoders can ignore this information since it is included in a backwards-compatible manner (as will be explained later). Therefore, the adverse impact on bit rate associated with including the eSBR metadata is negligible for several reasons, including the following:

[0090] The bitrate loss (attributable to the inclusion of the eSBR metadata) is a very small contribution to the total bitrate, as only the differential control data required to perform the eSBR processing is transmitted (and not the simulcast of the SBR control data); and

[0091] • The tuning of the SBR related control information does not generally depend on the details of the transposition. Examples where the control data depends on the operation of the transposer will be discussed later in this application.

[0092] Thus, embodiments of the present invention provide methods for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner. This efficient transmission of eSBR control data reduces memory requirements in decoders, encoders, and transcoders employing aspects of the present invention, while having no significant adverse impact on bit rate. Furthermore, the complexity and processing requirements associated with performing eSBR according to embodiments of the present invention are also reduced because the SBR data only needs to be processed once and is not simulcast, as would be the case when eSBR is treated as a completely separate object type in MPEG-4 AAC rather than being integrated into the MPEG-4 AAC codec in a backward-compatible manner.

[0093] Next, refer to Figure 7 , we describe elements of a block ("raw_data_block") of an MPEG-4 AAC bitstream (which contains eSBR metadata) according to some embodiments of the present invention. Figure 7 is a diagram of a block of an MPEG-4 AAC bitstream ("raw_data_block"), showing some segments of the MPEG-4 AAC bitstream.

[0094] A block of an MPEG-4 AAC bitstream may contain at least one "single_channel_element()" (e.g. Figure 7 ) and / or at least one "channel_pair_element()" ( Figure 7 The block may also contain several "fill_element" (e.g. Figure 7 Each "single_channel_element()" contains an identifier (e.g., a character string) indicating the start of a single channel element. Figure 7 ) and may contain audio data indicating different channels of a multi-channel audio program. Each "channel_pair_element()" contains an identifier ( Figure 7 ), and may include audio data indicating two channels of programming.

[0095] The fill_element (herein referred to as fill element) of an MPEG-4 AAC bitstream contains an identifier ( Figure 7 The identifier ID2 may be composed of a 3-bit unsigned integer ("uimsbf") with a value of 0x6, with the most significant bit transmitted first. The padding data may include an extension_payload() element (sometimes referred to herein as an extension payload) whose syntax is shown in Table 4.57 of the MPEG-4 AAC standard. Several types of extension payloads exist and are identified by an "extension_type" parameter, which is a 4-bit unsigned integer ("uimsbf") with the most significant bit transmitted first.

[0096] The padding data (eg, its extended payload) may include a header or an identifier (eg, Figure 7The header initializes the "SBR object" type called sbr_extension_data() in the MPEG-4 AAC standard. For example, a spectral band replication (SBR) extension payload is identified using a value of "1101" or "1110" for the extension_type field in the header, where the identifier "1101" identifies an extension payload with SBR data and "1110" identifies an extension payload containing SBR data with a cyclic redundancy check (CRC) to verify the correctness of the SBR data.

[0097] When the header (e.g., extension_type field) initializes the SBR object type, SBR metadata (sometimes referred to herein as "spectral band replication data," and referred to as sbr_data() in the MPEG-4 AAC standard) follows the header, and at least one spectral band replication extension element (e.g., Figure 7 The SBR metadata may be followed by an "SBR extension element" (e.g., a filler element 1 of the spectral band replication extension element). This spectral band replication extension element (segment of the bitstream) is called an "sbr_extension()" container in the MPEG-4 AAC standard. The spectral band replication extension element optionally contains a header (e.g., Figure 7 The "SBR extension header" of the padding element 1).

[0098] The MPEG-4 AAC standard anticipates that the Spectral Band Replication extension element may contain PS (Parametric Stereo) data for program audio data. The MPEG-4 AAC standard anticipates that when the header of a filler element (e.g., its extension payload) initializes an SBR object type (e.g., Figure 7 "Header 1" of the padding element) and the spectrum band replication extension element of the padding element contains PS data, the padding element (e.g., its extended payload) includes the spectrum band replication data and the "bs_extension_id" parameter, whose value (i.e., bs_extension_id = 2) indicates that the PS data is included in the spectrum band replication extension element of the padding element.

[0099] According to some embodiments of the present invention, eSBR metadata (e.g., a flag indicating whether enhanced spectral band replication (eSBR) processing is performed on the audio content of a block) is included in a spectral band replication extension element of a filler element. For example, this flag is in Figure 7 The flag is indicated in the filler element 1 of the Spectrum Band Duplication Extension Element, where the flag appears after the header of the "SBR Extension Element" of the filler element 1 (the "SBR Extension Header" of the filler element 1). This flag and additional eSBR metadata are optionally included in the Spectrum Band Duplication Extension Element after the header of the Spectrum Band Duplication Extension Element (e.g., after the SBR Extension Header). Figure 7According to some embodiments of the present invention, the filling element containing eSBR metadata also includes a "bs_extension_id" parameter, whose value (e.g., bs_extension_id=3) indicates that eSBR metadata is included in the filling element and eSBR processing is performed on the audio content of the relevant block.

[0100] According to some embodiments of the present invention, eSBR metadata is included in the filler elements (e.g. Figure 7 This is because a padding element containing extension_payload() (which has SBR data or SBR data with CRC) does not contain any other extension payloads of any other extension type. Therefore, in embodiments where the eSBR metadata stores its own extension payload, a separate padding element is used to store the eSBR metadata. This padding element contains an identifier (e.g., Figure 7 The padding data may include an extension_payload() element (sometimes referred to herein as an extension payload) whose syntax is shown in Table 4.57 of the MPEG-4 AAC standard. The padding data (e.g., its extension payload) includes a header indicating an eSBR object (e.g., Figure 7 The "header 2" of the filler element 2 of the header (i.e., the header initializes the enhanced spectral band replication (eSBR) object type), and the filler data (e.g., its extended payload) includes the eSBR metadata after the header. For example, Figure 7 The filler element 2 of the block contains this header ("header 2") and also contains eSBR metadata after the header (i.e., the "flag" in filler element 2 that indicates whether enhanced spectral band replication (eSBR) processing was performed on the audio content of the block). Additional eSBR metadata is also optionally included in the following block after header 2. Figure 7 In the embodiment described in this paragraph, the header (e.g. Figure 7 The header 2) has an identification value that is not the conventional value specified in Table 4.57 of the MPEG-4 AAC standard, but instead indicates an eSBR extended payload (so that the extension_type field of the header indicates that the padding data contains eSBR metadata).

[0101] In a first class of embodiments, the present invention is an audio processing unit (e.g., a decoder) comprising:

[0102] Memory (e.g. Figure 34 ), which is configured to store at least one block of an encoded audio bitstream (e.g., at least one block of an MPEG-4 AAC bitstream);

[0103] Bitstream payload deformatter (e.g. Figure 3 Component 205 or Figure 4 215) coupled to the memory and configured to demultiplex at least a portion of the block of the bitstream; and

[0104] Decoding subsystem (e.g. Figure 3 Elements 202 and 203 or Figure 4 202 and 213 ) coupled and configured to decode at least a portion of the audio content of the block of the bitstream, wherein the block comprises:

[0105] A filler element comprising an identifier indicating the start of the filler element (e.g., an "id_syn_ele" identifier having a value of 0x6 of Table 4.85 of the MPEG-4 AAC standard) and filler data following the identifier, wherein the filler data comprises:

[0106] At least one flag that identifies whether to perform enhanced spectral band replication (eSBR) processing on the audio content of the block (e.g., using spectral band replication data and eSBR metadata contained in the block).

[0107] The flag is eSBR metadata, and an example of the flag is the sbrPatchingMode flag. Another example of the flag is the harmonicSBR flag. Both of these flags indicate whether a basic form of spectral band replication or an enhanced form of spectral replication is performed on the audio data of the block. The basic form of spectral replication is spectral patching, and the enhanced form of spectral band replication is harmonic transposition.

[0108] In some embodiments, the padding data also includes additional eSBR metadata (ie, eSBR metadata other than the markers).

[0109] The memory may be a buffer memory (eg Figure 4 Implementation of a buffer 201) that stores (e.g., in a non-temporal manner) the at least one block of the encoded audio bitstream.

[0110] It is estimated that the complexity of performing eSBR processing (using eSBR harmonic transposition and pre-flattening) by an eSBR decoder during decoding of an MPEG-4 AAC bitstream containing eSBR metadata (indicating these eSBR tools) will be as follows (for a typical decoding with the indicated parameters):

[0111] Harmonic transposition (16kbps, 14400 / 28800Hz)

[0112] DFT-based: 3.68WMOPS (weighted million operations per second);

[0113] Based on QMF: 0.98WMOPS;

[0114] QMF patch pre-processing (pre-flattening): 0.1WMOPS.

[0115] It is well known that for transients, DFT-based transpose usually performs better than QMF-based transpose.

[0116] According to some embodiments of the present invention, a filler element (of an encoded audio bitstream) that contains eSBR metadata also contains a parameter (e.g., a "bs_extension_id" parameter) whose value (e.g., bs_extension_id=3) indicates that eSBR metadata is contained in the filler element and that eSBR processing is performed on the audio content of the associated block, and / or a parameter (e.g., the same "bs_extension_id" parameter) whose value (e.g., bs_extension_id=2) indicates that the sbr_extension() container of the filler element contains PS data. For example, as indicated in Table 1 below, this parameter with a value of bs_extension_id=2 may indicate that the sbr_extension() container of the filler element contains PS data, and this parameter with a value of bs_extension_id=3 may indicate that the sbr_extension() container of the filler element contains eSBR metadata:

[0117] Table 1

[0118] bs_extension_id meaning 0 reserve 1 reserve 2 EXTENSION_ID_PS 3 EXTENSION_ID_ESBR

[0119] According to some embodiments of the present invention, the syntax of each spectrum band duplication extension element containing eSBR metadata and / or PS data is as indicated in Table 2 below (wherein "sbr_extension()" indicates that it is a container of the spectrum band duplication extension element, "bs_extension_id" is as described in Table 1 above, "ps_data" indicates PS data, and "esbr_data" indicates eSBR metadata):

[0120] Table 2

[0121]

[0122]

[0123] In an exemplary embodiment, esbr_data() mentioned in Table 2 above indicates the values of the following metadata parameters:

[0124] 1. 1-bit metadata parameter “bs_sbr_preprocessing”; and

[0125] 2. For each channel ("ch") of the audio content of the encoded bitstream to be decoded, each of the above parameters is "sbrPatchingMode[ch]", "SbrOversamplingFlag[ch]", "SbrPitchInBinsFlag[ch]", and "sbrPitchInBins[ch]".

[0126] For example, in some embodiments, esbr_data() may have the syntax indicated in Table 3 to indicate these metadata parameters:

[0127] Table 3

[0128]

[0129]

[0130]

[0131] The above syntax enables efficient implementation of enhanced forms of spectral band replication (e.g., harmonic transposition) as extensions to legacy decoders. Specifically, the eSBR data of Table 3 contains only the parameters required to perform the enhanced form of spectral band replication, which are not already supported in the bitstream and cannot be directly derived from the parameters already supported in the bitstream. All other parameters and processing data required to perform the enhanced form of spectral band replication are extracted from existing parameters in defined locations in the bitstream.

[0132] For example, an MPEG-4 HE-AAC or HE-AAC v2 compatible decoder can be extended to include an enhanced form of spectral band replication, such as harmonic transposition. This enhanced form of spectral band replication is in addition to the basic form of spectral band replication already supported by the decoder. In the context of an MPEG-4 HE-AAC or HE-AAC v2 compatible decoder, this basic form of spectral band replication is the QMF spectral patching (SBR) tool, as defined in section 4.6.18 of the MPEG-4 AAC standard.

[0133] When performing an enhanced form of spectral band replication, the extended HE-AAC decoder can reuse many of the bitstream parameters that are already included in the SBR extension payload of the bitstream. The specific parameters that can be reused include, for example, various parameters that determine the main frequency band table. These parameters include bs_start_freq (a parameter that determines the start of the main frequency table parameters), bs_stop_freq (a parameter that determines the stop of the main frequency table), bs_freq_scale (a parameter that determines the number of frequency bands per octave), and bs_alter_scale (a parameter that alters the scale of the frequency bands). The reusable parameters also include parameters that determine the noise band table (bs_noise_bands) and the limiter band table parameters (bs_limiter_bands). Therefore, in various embodiments, at least some of the equivalent parameters specified in the USAC standard are omitted from the bitstream to thereby reduce the control burden of the bitstream. Typically, when a parameter specified in the AAC standard has an equivalent parameter specified in the USAC standard, the equivalent parameter specified in the USAC standard has the same name as the parameter specified in the AAC standard, such as the envelope scale factor E OrigMapped However, the equivalent parameters specified in the USAC standard generally have different values, being "tuned" according to the enhanced SBR process defined in the USAC standard rather than the SBR process defined in the AAC standard.

[0134] It is recommended to enable enhanced SBR to improve the subjective quality of audio content with harmonic frequency structure and strong tonal characteristics, especially at low bit rates. The values of the corresponding bitstream elements (i.e., esbr_data()) that control these tools can be determined in the encoder by applying a signal-dependent classification mechanism. In general, the use of the harmonic patching method (sbrPatchingMode == 1) is preferred for encoding music signals at very low bit rates, where the audio bandwidth of the core codec is significantly limited. This is particularly true when these signals contain a significant harmonic structure. In contrast, the use of the conventional SBR patching method is preferred for speech and mixed signals because it provides better preservation of the temporal structure of speech.

[0135] To improve the performance of the harmonic transposer, a preprocessing step can be enabled (bs_sbr_preprocessing == 1) which attempts to avoid introducing spectral discontinuities in the signal to the subsequent envelope adjuster. The operation of the tool is beneficial for signal types where the coarse spectral envelope of the low-band signal used for high-frequency reconstruction shows large fluctuation levels.

[0136] To improve the transient response of harmonic SBR patching, signal-adaptive frequency-domain oversampling can be applied (sbrOversamplingFlag == 1). Since signal-adaptive frequency-domain oversampling increases the computational complexity of the transposer and only benefits frames containing transients, its use is controlled by bitstream elements, which are transmitted once per frame and per independent SBR channel.

[0137] Decoders operating in the proposed enhanced SBR mode typically need to be able to switch between traditional SBR patching and enhanced SBR patching. Therefore, a delay that can be as long as the duration of one core audio frame may be introduced, depending on the decoder settings. Typically, the delays for both traditional and enhanced SBR patching will be similar.

[0138] In addition to many parameters, other data elements may also be reused by the extended HE-AAC decoder when performing an enhanced form of spectral band replication according to an embodiment of the present invention. For example, envelope data and noise floor data may also be extracted from the bs_data_env (envelope scale factor) and bs_noise_env (noise floor scale factor) data and used during the enhanced form of spectral band replication.

[0139] Essentially, these embodiments leverage configuration parameters and envelope data in the SBR extension payload that are already supported by legacy HE-AAC or HE-AAC v2 decoders to enable an enhanced form of spectral band replication that requires as little additional transmission data as possible. The metadata is initially tuned to a basic form of HFR (e.g., spectral shifting operations for SBR), but according to embodiments, is used for an enhanced form of HFR (e.g., harmonic transposition for eSBR). As previously discussed, the metadata generally represents operational parameters (e.g., envelope scaling factors, noise floor scaling factors, time / frequency grid parameters, sine wave addition information, variable crossover frequencies / bands, inverse filtering modes, envelope resolution, smoothing modes, frequency interpolation modes) that are tuned and designed for use with a basic form of HFR (e.g., linear spectral shifting). However, this metadata can be used in combination with additional metadata parameters specific to an enhanced form of HFR (e.g., harmonic transposition) to efficiently and effectively process audio data using the enhanced form of HFR.

[0140] Thus, an extension decoder that supports the enhanced form of spectral band replication can be generated in a very efficient manner by relying on already defined bitstream elements (e.g., bitstream elements in the SBR extension payload) and adding only the parameters (in the filler element extension payload) required to support the enhanced form of spectral band replication. This data reduction feature, combined with placing the newly added parameters in reserved data fields (e.g., the extension container), substantially reduces the barrier to production of decoders that support the enhanced form of spectral band replication by ensuring bitstream backward compatibility with legacy decoders that do not support the enhanced form of spectral band replication.

[0141] In Table 3, the numbers in the right row indicate the number of bits of the corresponding parameters in the left row.

[0142] In some embodiments, the SBR object type defined in MPEG-4 AAC is updated to include aspects of SBR tools and enhanced SBR (eSBR) tools, as predicted by the SBR extension element (bs_extension_id == EXTENSION_ID_ESBR). If a decoder detects and supports this SBR extension element, it adopts the predicted aspects of the enhanced SBR tools. An SBR object type updated in this manner is referred to as SBR enhancement.

[0143] In some embodiments, the present invention is a method comprising the steps of encoding audio data to produce an encoded bitstream (e.g., an MPEG-4 AAC bitstream), including by including eSBR metadata in at least one segment of at least one block of the encoded bitstream and including audio data in at least another segment of the block. In a typical embodiment, the method comprises the steps of multiplexing the audio data in each block of the encoded bitstream with the eSBR metadata. In typical decoding of the encoded bitstream in an eSBR decoder, the decoder extracts the eSBR metadata from the bitstream (including by parsing and demultiplexing the eSBR metadata and audio data) and uses the eSBR metadata to process the audio data to produce a decoded audio data stream.

[0144] Another aspect of the present invention is an eSBR decoder configured to perform eSBR processing (e.g., using at least one of the eSBR tools known as harmonic transposition or pre-flattening) during decoding of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that does not include eSBR metadata. Figure 5 To describe an example of this decoder.

[0145] Figure 5 The eSBR decoder 400 includes a buffer memory 201 connected as shown (which is the same as Figure 3 and 4 Memory 201), bit stream payload deformatter 215 (which is the same as Figure 4 deformatter 215), audio decoding subsystem 202 (sometimes referred to as the "core" decoding stage or "core" decoding subsystem, and is identical to Figure 3 core decoding subsystem 202), eSBR control data generation subsystem 401 and eSBR processing stage 203 (which is the same as Figure 3 Typically, decoder 400 also includes other processing elements (not shown).

[0146] In operation of the decoder 400 , a sequence of blocks of an encoded audio bitstream (MPEG-4 AAC bitstream) received by the decoder 400 is asserted from the buffer 201 to the deformatter 215 .

[0147] The deformatter 215 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (including quantization envelope data) and typically other metadata therefrom. The deformatter 215 is configured to assert at least the SBR metadata to the eSBR processing stage 203. The deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.

[0148] The audio decoding subsystem 202 of the decoder 400 is configured to decode the audio data extracted by the deformatter 215 (this decoding may be referred to as the "core" decoding operation) to produce decoded audio data and to pass the decoded audio data to the eSBR processing stage 203. Decoding is performed in the frequency domain. Typically, the final processing stage in the subsystem 202 applies a frequency-domain to time-domain transform to the decoded frequency-domain audio data, so that the output of the subsystem is time-domain decoded audio data. Stage 203 is configured to apply the SBR tools (and eSBR tools) indicated by the SBR metadata (extracted by the deformatter 215) and the eSBR metadata generated in the subsystem 401 to the decoded audio data (i.e., using the SBR and eSBR metadata to perform SBR and eSBR processing on the output of the decoding subsystem 202) to produce fully decoded audio data output by the decoder 400. Typically, decoder 400 includes memory (accessible by subsystem 202 and stage 203) that stores deformatted audio data and metadata output from deformatter 215 (and optionally subsystem 401), and stage 203 is configured to access the audio data and metadata as needed during SBR and eSBR processing. The SBR processing in stage 203 can be considered post-processing of the output of core decoding subsystem 202. Decoder 400 also optionally includes a final upmix subsystem (which can use the PS metadata extracted by deformatter 215 to apply the parametric stereo "PS" tool defined in the MPEG-4 AAC standard) coupled and configured to perform upmixing on the output of stage 203 to produce fully decoded upmixed audio output from APU 210.

[0149] Parametric stereo is a coding tool that uses a linear downmix of the left and right channels of a stereo signal and a set of spatial parameters that describe the stereo image to represent a stereo signal. Parametric stereo typically employs three types of spatial parameters: (1) inter-channel intensity difference (IID), which describes the intensity difference between channels; (2) inter-channel phase difference (IPD), which describes the phase difference between channels; and (3) inter-channel coherence (ICC), which describes the coherence (or similarity) between channels. Coherence can be measured as the maximum value of the cross-correlation that varies as a function of time or phase. These three parameters typically achieve high-quality reconstruction of the stereo image. However, the IPD parameter only specifies the relative phase differences between the channels of the stereo input signal and does not indicate the distribution of these phase differences over the left and right channels. Therefore, a fourth type of parameter that describes the total phase offset or total phase difference (OPD) can be used in addition. In the stereo reconstruction process, consecutive window segments of both the received downmix signal s[n] and the decorrelated version of the received downmix d[n] are processed together with the spatial parameters to produce the left (l k (n)) and right (r k (n))Reconstruct the signal:

[0150] l k (n) = H 11 (k,n)s k (n)+H 21 (k,n)d k (n)

[0151] r k (n) = H 12 (k,n)s k (n)+H 22 (k,n)d k (n)

[0152] Among them H 11 、H 12 、H 21 and H 22 Finally, the signal l is transformed by frequency to time transformation. k (n) and r k (n) Transform back to the time domain.

[0153] Figure 5 The control data generation subsystem 401 is coupled to and configured to detect at least one property of the encoded audio bitstream to be decoded and, in response to at least one result of the detection step, generate eSBR control data (which may be or include any type of eSBR metadata included in the encoded audio bitstream according to other embodiments of the present invention). The eSBR control data is asserted to stage 203 to trigger the application of individual eSBR tools or combinations of eSBR tools and / or to control the application of these eSBR tools after detecting a particular property (or combination of properties) of the bitstream. For example, to control the performance of eSBR processing using harmonic transposition, some embodiments of the control data generation subsystem 401 will include: a music detector (e.g., a simplified version of a conventional music detector) for setting the sbrPatchingMode[ch] parameter in response to detecting whether the bitstream indicates music (and asserting the set parameter to stage 203); a transient detector for setting the sbrOversamplingFlag[ch] parameter in response to detecting the presence or absence of transients in the audio content indicated by the bitstream (and asserting the set parameter to stage 203); and / or a pitch detector for setting the sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch] parameters in response to detecting pitches in the audio content indicated by the bitstream (and asserting the set parameters to stage 203). Other aspects of the invention are methods of audio bitstream decoding performed by any embodiment of the inventive decoder described in this and the previous paragraphs.

[0154] Aspects of the present invention include the types of encoding or decoding methods that any embodiment of the inventive APU, system, or device is configured (e.g., programmed) to perform. Other aspects of the present invention include systems or devices configured (e.g., programmed) to perform any embodiment of the inventive method and computer-readable media (e.g., optical disks) storing (e.g., in a non-transitory manner) code for implementing any embodiment of the inventive method or its steps. For example, the inventive system may be or include a programmable general-purpose processor, a digital signal processor, or a microprocessor, any of which is programmed and / or otherwise configured using software or firmware to perform various operations on data (including embodiments of the inventive method or its steps). Such a general-purpose processor may be or include a computer system that includes an input device, memory, and processing circuitry that is programmed (and / or otherwise configured) to execute an embodiment of the inventive method (or its steps) in response to data asserted thereto.

[0155] Embodiments of the present invention may be implemented in hardware, firmware, or software, or a combination of both (e.g., as a programmable logic array). Unless otherwise stated, the algorithms or processes included as part of the present invention are not inherently related to any particular computer or other apparatus. In particular, various general-purpose machines may be used with programs written according to the teachings herein, or it may be more convenient to construct more specialized apparatus (e.g., integrated circuits) to perform the required method steps. Thus, the present invention may be implemented in one or more computer programs executing on one or more programmable computer systems (e.g., Figure 1 components, or Figure 2 The encoder 100 (or its components), or Figure 3 The decoder 200 (or its components), or Figure 4 The decoder 210 (or its components) or Figure 5 The present invention relates to any embodiment of the decoder 400 (or elements thereof) of the present invention, wherein the one or more programmable computer systems each include at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known manner.

[0156] Each such program can be implemented in any desired computer language (including machine, assembly or high-level procedural, logical or object-oriented programming languages) to communicate with a computer system. In either case, the language can be a compiled or interpreted language.

[0157] For example, when implemented by a sequence of computer software instructions, the various functions and steps of the embodiments of the present invention may be implemented by a multi-threaded sequence of software instructions running in hardware suitable for digital signal processing, in which case the various devices, steps, and functions of the embodiments may correspond to portions of the software instructions.

[0158] Each such computer program is preferably stored or downloaded onto a storage medium or device (e.g., solid-state memory or media or magnetic or optical media) readable by a general or special purpose programmable computer to configure and operate the computer when the storage medium or device is read by a computer system to execute the program described herein. The present invention system may also be implemented as a computer-readable storage medium configured with (i.e., storing) a computer program, wherein the storage medium so configured causes the computer system to operate in a specific and predefined manner to perform the functions described herein.

[0159] Many embodiments of the present invention have been described. However, it will be appreciated that various modifications may be made without departing from the spirit and scope of the present invention. Many modifications and variations of the present invention may be made in light of the above teachings. For example, to facilitate efficient implementation, phase shifting may be combined with complex QMF analysis and synthesis filter banks. The analysis filter bank is responsible for filtering the time domain low-band signal generated by the core decoder into a plurality of sub-bands (e.g., QMF sub-bands). The synthesis filter bank is responsible for combining the regenerated high-band (as indicated by the received sbrPatchingMode parameter) generated by the selected HFR technique with the decoded low-band to produce a broadband output audio signal. However, a given filter bank implementation operating in a certain sampling rate mode (e.g., normal double-rate operation or downsampled SBR mode) should not have a phase shift that is dependent on the bitstream. The QMF bank used in SBR is a complex exponential extension of the theory of a cosine modulated filter bank. It can be shown that when complex exponential modulation is used to extend the cosine modulated filter bank, the aliasing elimination constraint becomes obsolete. Therefore, for the SBR QMF bank, the analysis filter h k (n) and synthesis filter f k (n) Both can be defined by the following equations:

[0160]

[0161] Where p0(n) is a real-valued symmetric or asymmetric prototype filter (typically a low-pass prototype filter), M represents the number of channels, and N is the prototype filter order. The number of channels used in the analysis filterbank can be different from the number of channels used in the synthesis filterbank. For example, the analysis filterbank may have 32 channels and the synthesis filterbank may have 64 channels. When operating the synthesis filterbank in downsampling mode, the synthesis filterbank may only have 32 channels. Since the subband samples from the filterbank are complex-valued, an additive, channel-dependent phase shift step can be added to the analysis filterbank. These additional phase shifts need to be compensated for before the synthesis filterbank. Although the phase shift term can, in principle, have arbitrary values without disrupting the operation of the QMF analysis / synthesis chain, it can also be constrained to certain values for consistency verification. The SBR signal is affected by the choice of phase factor, while the low-pass signal from the core decoder is not. The audio quality of the output signal is unaffected.

[0162] The coefficients p0(n) of the prototype filter may be defined as a length L of 640, as shown in Table 4 below.

[0163] Table 4

[0164]

[0165]

[0166]

[0167]

[0168]

[0169]

[0170]

[0171]

[0172] The prototype filter p0(n) may also be derived from Table 4 by one or more mathematical operations such as rounding, subsampling, interpolation, and decimation.

[0173] Although the tuning of SBR-related control information generally does not depend on the details of transposition (as previously discussed), in some embodiments, certain elements of the control data may be simulcast in the eSBR extension container (bs_extension_id == EXTENSION_ID_ESBR) to improve the quality of the reproduced signal. Some of the simulcast elements may include noise floor data (e.g., a noise floor scale factor and a parameter indicating the direction (frequency or time) of differential encoding of each noise floor), inverse filtering data (e.g., a parameter indicating an inverse filtering mode selected from no inverse filtering, low inverse filtering level, moderate inverse filtering level, and strong inverse filtering level), and missing harmonics data (e.g., a parameter indicating whether a sine wave should be added to a particular frequency band of the reproduced high frequency band). All of these elements rely on a synthetic simulation of the decoder's transposer implemented in the encoder and can therefore improve the quality of the reproduced signal after being properly tuned according to the selected transposer.

[0174] Specifically, in some embodiments, missing harmonics and inverse filter control data (along with the other bitstream parameters of Table 3) are transmitted in an eSBR extension container and tuned according to the eSBR harmonic transposer. The additional bitrate required to transmit these two types of metadata for the eSBR harmonic transposer is relatively low. Therefore, sending the tuned missing harmonics and / or inverse filter control data in the eSBR extension container will improve the quality of the audio produced by the transposer while only slightly affecting the bitrate. To ensure backward compatibility with legacy decoders, parameters tuned for SBR's spectral shifting operation can also be sent as part of the SBR control data in the bitstream using implicit or explicit signaling.

[0175] The complexity of the decoder with SBR enhancement described in this application must be limited so as not to significantly increase the overall computational complexity of the implementation. Preferably, when using the eSBR tool, the PCU (MOP) of the SBR object type is equal to or lower than 4.5, and when using the eSBR tool, the RCU of the SBR object type is equal to or lower than 3. The approximate processing power is given in Processor Complexity Units (PCU) (specified by an integer number of MOPS). The approximate RAM usage is given in RAM Complexity Units (RCU) (specified by an integer number of kWords (1000 words)). The number of RCUs does not include working buffers that can be shared between different objects and / or channels. In addition, the PCU is proportional to the sampling frequency. The PCU value is given in MOPS (millions of operations per second) per channel and the RCU value is given in kilowords per channel.

[0176] Special attention should be paid to compressed data, such as HE-AAC encoded audio, which can be decoded by different decoder configurations. In this case, decoding can be done in a backward-compatible mode (AAC only) as well as in an enhanced mode (AAC + SBR). If the compressed data allows both backward-compatible and enhanced decoding, and if the decoder operates in an enhanced mode such that it uses a post-processor that inserts some additional delay (such as the SBR post-processor in HE-AAC), then it is necessary to ensure that this additional time delay caused by the backward-compatible mode is taken into account when presenting the combined unit, as described by the corresponding value n. To ensure correct handling of the combined timestamp (so that the audio remains synchronized with other media), when the decoder operating mode includes the SBR enhancements described in this application (including eSBR), the additional delay introduced by post-processing, given as the number of samples (per audio channel) at the output sampling rate, is 3010. Therefore, for an audio combined unit, when the decoder operating mode includes the SBR enhancements described in this application, the combined time applies to the 3011th audio sample within the combined unit.

[0177] SBR enhancement should be enabled to improve the subjective quality of audio content with harmonic frequency structure and strong tonal characteristics, especially at low bit rates. The values of the corresponding bitstream elements (i.e., esbr_data()) that control these tools can be determined in the encoder by applying a signal-dependent classification mechanism.

[0178] In general, the use of harmonic patching (sbrPatchingMode == 0) is preferred for music signals encoded at very low bit rates, where the core codec's audio bandwidth is significantly limited. This is particularly true when these signals contain a pronounced harmonic structure. Conversely, the use of the conventional SBR patching method is preferred for speech and mixed signals, as it provides better preservation of speech's temporal structure.

[0179] To improve the performance of the MPEG-4 SBR transposer, a preprocessing step can be enabled (bs_sbr_preprocessing == 1) which avoids introducing spectral discontinuities of the signal into the subsequent envelope adjuster. The operation of the tool is beneficial for signal types where the coarse spectral envelope of the low-band signal used for high-frequency reconstruction shows large fluctuation levels.

[0180] To improve the transient response of harmonic SBR patching (sbrPatchingMode == 0), signal-adaptive frequency-domain oversampling (sbrOversamplingFlag == 1) can be applied. Since signal-adaptive frequency-domain oversampling increases the computational complexity of the transposer but only benefits frames containing transients, its use is controlled by bitstream elements, which are transmitted once per frame and per independent SBR channel.

[0181] Typical bitrate settings for HE-AACv2 with SBR enhancement (i.e., the harmonic transposer with the eSBR tool enabled) suggest a range of 20 kbp to 32 kbp for stereo audio content at a sampling rate of 44.1 kHz or 48 kHz. The relative subjective quality gain of SBR enhancement increases towards the lower bitrate boundaries, and a suitably configured encoder allows this range to be extended to even lower bitrates. The bitrates provided above are only suggestions and may be adapted to specific service requirements.

[0182] Decoders operating in the proposed enhanced SBR mode typically need to be able to switch between traditional SBR patching and enhanced SBR patching. Therefore, a delay that can be as long as the duration of one core audio frame can be introduced, depending on the decoder settings. Typically, the delays for both traditional and enhanced SBR patching will be similar.

[0183] It is to be understood that within the scope of the appended claims, the invention may be practiced otherwise than as specifically described herein.Any element signs contained in the following claims are for illustration only and should in no way be used to interpret or limit the claims.

[0184] Various aspects of the invention may be understood from the following enumerated example embodiments (EEE):

[0185] EEE 1. A method for performing high-frequency reconstruction of an audio signal, the method comprising:

[0186] receiving an encoded audio bitstream comprising audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata;

[0187] decoding the audio data to generate a decoded low-band audio signal;

[0188] extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata comprising operating parameters of a high frequency reconstruction process, the operating parameters comprising a patch mode parameter located in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates spectral translation and a second value of the patch mode parameter indicates harmonic transposition by phase vocoder frequency stretching;

[0189] filtering the decoded low-band audio signal to produce a filtered low-band audio signal;

[0190] reproducing a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein the regeneration comprises spectral translation if the patch mode parameter is the first value, and comprises harmonic transposition by a phase vocoder frequency stretching if the patch mode parameter is the second value; and

[0191] combining the filtered low-band audio signal with the regenerated high-band portion to form a wideband audio signal,

[0192] wherein the filtering, regeneration, and combining are performed as post-processing operations with a delay of 3010 samples or less per audio channel, and wherein the spectral shifting comprises maintaining a ratio between tonal and noise-like components by adaptive inverse filtering.

[0193] EEE 2. The method according to EEE 1, wherein the encoded audio bitstream further includes a filler element having an identifier indicating a start of the filler element and filler data following the identifier, wherein the filler data includes the backward-compatible extension container. EEE 2.

[0194] EEE 3. The method according to EEE 2, wherein the identifier is a 3-bit unsigned integer with the most significant bit transmitted first and having a value of 0x6.

[0195] EEE 4. The method according to EEE 2 or EEE 3, wherein the padding data includes an extended payload, the extended payload includes spectrum band replication extension data, and the extended payload is identified by a 4-bit unsigned integer with the most significant bit transmitted first and having a value of "1101" or "1110", and optionally,

[0196] The spectrum band replication extension data includes:

[0197] Optional spectrum with duplicate header,

[0198] Spectral band copy data, which follows the header, and

[0199] A spectrum band replication extension element is located after the spectrum band replication data, and wherein the flag is included in the spectrum band replication extension element.

[0200] EEE 5. The method according to any one of EEEs 1 to 4, wherein the high-frequency reconstruction metadata comprises an envelope scaling factor, a noise floor scaling factor, time / frequency grid information, or a parameter indicating a crossover frequency. EEE 5.

[0201] EEE 6. A method according to any one of EEEs 1 to 5, wherein the backward compatible extension container further includes a flag indicating whether additional preprocessing is used to avoid shape discontinuities of the spectral envelope of the high-frequency band portion when the patch mode parameter is equal to the first value, wherein the first value of the flag enables the additional preprocessing and the second value of the flag disables the additional preprocessing.

[0202] EEE 7. The method according to EEE 6, wherein the additional preprocessing comprises calculating a pre-gain curve using linear prediction filter coefficients. EEE 7.

[0203] EEE 8. A method according to any one of EEEs 1 to 5, wherein the backward compatible extension container further includes a flag indicating whether signal adaptive frequency domain oversampling is applied when the patch mode parameter is equal to the second value, wherein the first value of the flag enables the signal adaptive frequency domain oversampling and the second value of the flag disables the signal adaptive frequency domain oversampling.

[0204] EEE 9. The method according to EEE 8, wherein the signal adaptive frequency domain oversampling is applied only to frames containing transients. EEE 9.

[0205] EEE 10. The method of any one of the preceding EEEs, wherein the harmonic transposition by phase vocoder frequency stretching is performed at an estimated complexity equal to or lower than 4.5 million operations per second and 3 kilowords of memory. EEE 11.

[0206] EEE 11. A non-transitory computer-readable medium containing instructions that, when executed by a processor, perform the method according to any one of EEEs 1 to 10. EEE 12.

[0207] EEE 12. A computer program product having instructions which, when executed by a computing device or system, cause the computing device or system to perform the method according to any one of EEEs 1 to 10.

[0208] EEE 13. An audio processing unit for performing high frequency reconstruction of an audio signal, the audio processing unit comprising:

[0209] an input interface for receiving an encoded audio bitstream comprising audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata;

[0210] a core audio decoder for decoding the audio data to generate a decoded low-band audio signal;

[0211] a deformatter for extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata comprising operating parameters for a high frequency reconstruction process, the operating parameters comprising a patch mode parameter located in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates spectral translation and a second value of the patch mode parameter indicates harmonic transposition by phase vocoder frequency stretching;

[0212] an analysis filter bank for filtering the decoded low-band audio signal to generate a filtered low-band audio signal;

[0213] a high frequency regenerator for reconstructing a high frequency band portion of the audio signal using the filtered low frequency band audio signal and the high frequency reconstruction metadata, wherein the reconstruction comprises spectral translation if the patch mode parameter is the first value, and comprises harmonic transposition by phase vocoder frequency stretching if the patch mode parameter is the second value; and

[0214] a synthesis filter bank for combining the filtered low-band audio signal with the regenerated high-band portion to form a wideband audio signal,

[0215] wherein the analysis filter bank, high frequency regenerator, and synthesis filter bank are performed in a post-processor having a delay of 3010 samples or less per audio channel, and wherein the spectral shifting comprises maintaining a ratio between tonal components and noise-like components by adaptive inverse filtering.

[0216] EEE 14. The audio processing unit according to EEE 13, wherein the harmonic transposition by phase vocoder frequency stretching is performed with an estimated complexity equal to or lower than 4.5 million operations per second and 3 kilowords of memory. EEE 15.

Claims

1. A method for performing high frequency reconstruction of an audio signal, the method comprising: receiving an encoded audio bitstream comprising audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata; decoding the audio data to generate a decoded low-band audio signal; extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata comprising operating parameters of a high frequency reconstruction process, the operating parameters comprising a patch mode parameter located in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates spectral translation and a second value of the patch mode parameter indicates harmonic transposition by phase vocoder frequency stretching; filtering the decoded low-band audio signal to produce a filtered low-band audio signal; as well as reproducing a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein the regeneration comprises spectral translation if the patch mode parameter is the first value, and comprises harmonic transposition by a phase vocoder frequency stretching if the patch mode parameter is the second value, wherein the filtering and regeneration are performed as a post-processing operation with a delay of 3010 samples per audio channel, and wherein the spectral shifting comprises maintaining the ratio between tonal and noise-like components by adaptive inverse filtering.

2. The method according to claim 1, wherein the backward compatible extension container further includes a flag indicating whether signal adaptive frequency domain oversampling is applied when the patch mode parameter is equal to the second value, wherein a first value of the flag enables the signal adaptive frequency domain oversampling and a second value of the flag disables the signal adaptive frequency domain oversampling. The method according to claim 2 , wherein the signal adaptive frequency domain oversampling is applied only to frames containing transients.

4. The method of claim 1, wherein the harmonic transposition by phase vocoder frequency stretching is performed at an estimated complexity of at or below 4.5 million operations per second and at or below 3 kilowords of memory.

5. A non-transitory computer-readable medium comprising instructions, which, when executed by a processor, perform the method of claim 1.

6. A computer program product stored in a non-transitory computer-readable medium having instructions that, when executed by a computing device or system, cause the computing device or system to perform the method according to claim 1.

7. An audio processing unit for performing high frequency reconstruction of an audio signal, the audio processing unit comprising: an input interface for receiving an encoded audio bitstream comprising audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata; a core audio decoder for decoding the audio data to generate a decoded low-band audio signal; a deformatter for extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata comprising operating parameters for a high frequency reconstruction process, the operating parameters comprising a patch mode parameter located in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates spectral translation and a second value of the patch mode parameter indicates harmonic transposition by phase vocoder frequency stretching; an analysis filter bank for filtering the decoded low-band audio signal to generate a filtered low-band audio signal; as well as a high frequency regenerator for reconstructing a high frequency band portion of the audio signal using the filtered low frequency band audio signal and the high frequency reconstruction metadata, wherein the reconstruction comprises spectral translation if the patch mode parameter is the first value, and comprises harmonic transposition by a phase vocoder frequency stretching if the patch mode parameter is the second value; wherein the analysis filter bank and the high frequency regenerator are performed in a post-processor with a delay of 3010 samples per audio channel, and wherein the spectral shifting comprises maintaining the ratio between tonal and noise-like components by adaptive inverse filtering.

Citation Information

Patent Citations

  • Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element

    CN107430867A

  • Integration of high-frequency reconstruction techniques with reduced post-processing latency

    CN112204659B