Backward-Compatible Integration of High-Frequency Reconstruction Techniques for Audio Signals

By introducing analytical filter banks and high-frequency reconstruction metadata into the audio decoder, using harmonic transposition and QMF-patch preprocessing technology, the problem of low spectrum band replication efficiency in existing audio encoding technologies is solved, and the high-frequency reconstruction quality and encoding efficiency of audio signals are improved.

CN113990332BActive Publication Date: 2025-08-01DOLBY INTERNATIONAL AB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111240006.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-01-26
Filing Date
2019-01-28
Publication Date
2025-08-01
Estimated Expiration
2039-01-28

AI Technical Summary

Technical Problem

Existing audio encoding technologies are inefficient in spectrum band replication, especially for music content with low crossover frequencies, spectrum patching or linear translation is not ideal, resulting in a degradation of audio quality.

Method used

By introducing analytical filter banks and high-frequency reconstruction metadata into the audio decoder, using harmonic transposition and QMF-patch preprocessing technology, combined with the Enhanced Spectral Band Replication Tool (eSBR) for high-frequency reconstruction, improve the spectrum characteristics of the audio signal.

Benefits of technology

Improve the high-frequency reconstruction quality of audio signals, especially for music content, enhance the efficiency and quality of audio encoding, and maintain backtrack compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113990332B_ABST
    Figure CN113990332B_ABST
Patent Text Reader

Abstract

This application relates to a backward-compatible integration of high-frequency reconstruction techniques for audio signals. The present invention discloses a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream and decoding the audio data to produce a decoded low-band audio signal. The method further includes extracting high-frequency reconstruction metadata and filtering the decoded low-band audio signal using an analysis filter bank to produce a filtered low-band audio signal. The method also includes extracting a flag indicating that spectral translation or harmonic transposition has been performed on the audio data and regenerating a high-band portion of the audio signal based on the flag using the filtered low-band audio signal and the high-frequency reconstruction metadata.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Relevant information on divisional applications

[0002] This application is a divisional application. The parent case of this divisional application is the patent application No. 201980010173.2, titled "Backward-Compatible Integration of High-Frequency Reconstruction Techniques for Audio Signals", which is the Chinese national phase entry of the PCT international application with the filing date of January 28, 2019 and the application number of PCT / US2019 / 015442.

[0003] Cross-reference to related applications

[0004] This application claims the benefit of priority from the following prior applications: U.S. Provisional Application No. 62 / 622,205, filed on January 26, 2018, which is hereby incorporated herein by reference. Technical Field

[0005] Embodiments relate to audio signal processing, and more particularly, to the encoding, decoding, or transcoding of audio bitstreams, where control data indicates either a basic form of high-frequency reconstruction ("HFR") or an enhanced form of HFR to be performed on audio data. Background Art

[0006] Typical information bitstreams contain both audio data (e.g., encoded audio data) indicating one or more channels of audio content and metadata indicating at least one characteristic of the audio data or audio content. One well-known format for generating an encoded audio bitstream is the MPEG-4 Advanced Audio Coding (AAC) format described in the MPEG standard ISO / IEC 14496-3:2009. In the MPEG-4 standard, AAC stands for "Advanced Audio Coding" and HE-AAC stands for "High Efficiency Advanced Audio Coding".

[0007] The MPEG-4 AAC standard defines several audio profiles that determine which objects and coding tools are present in a compliant encoder or decoder. Three of these audio profiles are (1) the AAC profile, (2) the HE-AAC profile, and (3) the HE-AAC v2 profile. The AAC profile contains the AAC Low Complexity (or "AAC-LC") object type. The AAC-LC object is a counterpart of the MPEG-2 AAC Low Complexity profile with some adaptations and does not contain the Spectral Band Replication ("SBR") object type or the Parametric Stereo ("PS") object type. The HE-AAC profile is a superset of the AAC profile and additionally contains the SBR object type. The HE-AAC v2 profile is a superset of the HE-AAC profile and additionally contains the PS object type.

[0008] The SBR object type contains a spectral band replication tool, which is an important high-frequency reconstruction ("HFR") coding tool that significantly improves the compression efficiency of perceptual audio codecs. SBR reconstructs the high-frequency components of the audio signal on the receiver side (e.g., in the decoder). Thus, the encoder only needs to encode and transmit the low-frequency components, allowing for higher audio quality at low data rates. SBR is based on the replication of a previously truncated harmonic sequence in order to reduce the data rate from the available bandwidth-limited signal and the control data obtained from the encoder. The ratio between the tonal and noise-like components is maintained through adaptive inverse filtering and the selective addition of noise and sine waves. In the MPEG-4 AAC standard, the SBR tool performs spectral patching (also known as linear translation or spectral translation), where several consecutive quadrature mirror filter (QMF) subbands are replicated (or "patched") from the transmitted low-frequency part of the audio signal to the high-frequency band part of the audio signal (which is generated in the decoder).

[0009] For certain audio types (e.g., music content with a relatively low crossover frequency), spectral patching or linear translation may not be ideal. Thus, techniques for improving spectral band replication are needed. Summary of the Invention

[0010] A first category of embodiments is disclosed, which relates to a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream and decoding the audio data to produce a decoded low-frequency band audio signal. The method further includes extracting high-frequency reconstruction metadata and filtering the decoded low-frequency band audio signal using an analysis filter bank to produce a filtered low-frequency band audio signal. The method further includes extracting a flag indicating that spectral translation or harmonic transposition has been performed on the audio data and regenerating the high-frequency band part of the audio signal based on the flag using the filtered low-frequency band audio signal and the high-frequency reconstruction metadata. Finally, the method includes combining the filtered low-frequency band audio signal with the regenerated high-frequency band part to form a wide-band audio signal.

[0011] A second class of embodiments relates to an audio decoder for decoding an encoded audio bitstream. The decoder comprises: an input interface for receiving the encoded audio bitstream, wherein the encoded audio bitstream contains audio data representing a low-frequency band portion of an audio signal; and a core decoder for decoding the audio data to produce a decoded low-frequency band audio signal. The decoder further comprises: a demultiplexer for extracting high-frequency reconstruction metadata from the encoded audio bitstream, wherein the high-frequency reconstruction metadata contains operation parameters for a high-frequency reconstruction process of linearly translating a consecutive number of sub-bands from the low-frequency band portion of the audio signal to the high-frequency band portion of the audio signal; and an analysis filter bank for filtering the decoded low-frequency band audio signal to produce a filtered low-frequency band audio signal. The decoder further comprises: a demultiplexer for extracting a flag indicating whether linear translation or harmonic transposition is performed on the audio data from the encoded audio bitstream; and a high-frequency regenerator for regenerating the high-frequency band portion of the audio signal using the filtered low-frequency band audio signal and the high-frequency reconstruction metadata according to the flag. Finally, the decoder comprises a synthesis filter bank for combining the filtered low-frequency band audio signal and the regenerated high-frequency band portion to form a wide-band audio signal.

[0012] Other classes of embodiments relate to encoding and transcoding audio bitstreams containing metadata identifying whether enhanced spectral band replication (eSBR) processing is performed. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a block diagram of an embodiment of a system that can be configured to perform an embodiment of the method of the present invention.

[0014] Figure 2 is a block diagram of an encoder that is an embodiment of an audio processing unit of the present invention.

[0015] Figure 3 is a block diagram of a system that includes a decoder that is an embodiment of an audio processing unit of the present invention and optionally also a post-processor coupled to the decoder.

[0016] Figure 4 is a block diagram of a decoder that is an embodiment of an audio processing unit of the present invention.

[0017] Figure 5 is a block diagram of a decoder that is another embodiment of an audio processing unit of the present invention.

[0018] Figure 6 is a block diagram of another embodiment of an audio processing unit of the present invention.

[0019] Figure 7It is a schema of a block of an MPEG-4 AAC bitstream, and the block contains the segments into which it is divided.

[0020] Notation and nomenclature

[0021] Throughout the present invention (including in the claims), the expression "perform an operation on" a signal or data is used in a broad sense to mean performing an operation directly on the signal or data or on a processed version of the signal or data (e.g., a version of a signal that has undergone initial filtering or preprocessing before the operation is performed on it).

[0022] Throughout the present invention (including in the claims), the expression "audio processing unit" or "audio processor" is used in a broad sense to mean a system, apparatus, or device configured to process audio data. Examples of audio processing units include (but are not limited to) encoders, codecs, decoders, codecs, preprocessing systems, postprocessing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools). Substantially all consumer electronic devices (e.g., mobile phones, televisions, laptop computers, and tablet computers) contain an audio processing unit or an audio processor.

[0023] Throughout the present invention (including in the claims), the term "coupled" or "is coupled" is used in a broad sense to mean directly or indirectly connected. Thus, if a first device is coupled to a second device, then the connection can be by a direct connection or by an indirect connection via other devices and connections. In addition, components integrated into or integrated with other components are also coupled to each other. Detailed description

[0024] The MPEG-4 AAC standard contemplates that an encoded MPEG-4 AAC bitstream contains metadata that indicates each type of high-frequency reconstruction ("HFR") processing to be applied (if any) by a decoder to decode the audio content of the bitstream and / or controls this HFR processing and / or indicates at least one characteristic or parameter of at least one HFR tool to be used to decode the audio content of the bitstream. Herein, we use the expression "SBR metadata" to denote this type of metadata described or mentioned in the MPEG-4 AAC standard for use with spectral band replication ("SBR"). Those skilled in the art will understand that SBR is a form of HFR.

[0025] SBR is preferably used as a dual-rate system, where the underlying codec operates at half of the original sampling rate, while SBR operates at the original sampling rate. The SBR encoder works in parallel with the underlying core codec, but at a higher sampling rate. Although SBR is mainly a post-process in the decoder, important parameters are extracted in the encoder to ensure the most accurate high-frequency reconstruction in the decoder. The encoder estimates the spectral envelope of the SBR range for a time and frequency range / resolution suitable for the characteristics of the current input signal segment. The spectral envelope is estimated by complex QMF analysis and subsequent energy calculation. The time and frequency resolution of the spectral envelope can be highly freely selected to ensure the most suitable time-frequency resolution for a given input segment. The envelope estimation needs to consider that the transients in the original mainly located in the high-frequency region (such as the top-hat region) will slightly exist in the high-frequency band generated by SBR before the envelope adjustment, because the high-frequency band in the decoder is based on the low-frequency band where the transients are less significant compared to the high-frequency band. This aspect poses different requirements on the time-frequency resolution of the spectral envelope data compared to the ordinary spectral envelope used in other audio coding algorithms.

[0026] In addition to the spectral envelope, several additional parameters representing the spectral characteristics of the input signal for different time and frequency regions are also extracted. Since the encoder can naturally access the original signal and the information on how the SBR unit in the decoder will generate the high-frequency band, given a specific set of control parameters, the system can handle situations where the low-frequency band constitutes a strong harmonic series and the high-frequency band to be regenerated mainly constitutes random signal components, as well as situations where strong tonal components exist in the original high-frequency band without corresponding ones in the low-frequency band (the high-frequency band region is based on which). Furthermore, the SBR encoder works closely with the underlying core codec to evaluate which frequency range should be covered by SBR at a given time. In the case of stereo signals, the SBR data is efficiently encoded before transmission by exploiting the entropy coding and the channel dependency of the control data.

[0027] The control parameter extraction algorithm usually needs to be carefully tuned to the underlying codec at a given bitrate and a given sampling rate. This is due to the fact that compared to high bitrates, lower bitrates usually imply a larger SBR range, and different sampling rates correspond to different time resolutions of the SBR frames.

[0028] An SBR decoder generally includes several different parts. The SBR decoder includes a bitstream decoding module, a high-frequency reconstruction (HFR) module, an additional high-frequency component module, and an envelope adjuster module. The system is based on a complex-valued QMF filter bank (for high-quality SBR) or a real-valued QMF filter bank (for low-power SBR). Embodiments of the present invention are applicable to both high-quality SBR and low-power SBR. In the bitstream extraction module, control data is read from the bitstream and decoded. A time-frequency grid is obtained for the current frame before reading the envelope data from the bitstream. The underlying core decoder decodes the audio signal of the current frame (although at a lower sampling rate) to produce time-domain audio samples. The resulting frame of audio data is used for high-frequency reconstruction by the HFR module. Then the decoded low-frequency band signal is analyzed using a QMF filter bank. Subsequently, high-frequency reconstruction and envelope adjustment are performed on the subband samples of the QMF filter bank. The high frequency is reconstructed from the low frequency in a flexible manner based on given control parameters. In addition, the reconstructed high-frequency band is adaptively filtered on a subband channel basis according to the control data to ensure appropriate spectral characteristics in a given time / frequency region.

[0029] The top level of an MPEG-4 AAC bitstream is a sequence data block ("raw_data_block" element), and each of these blocks is a segment of data (referred to herein as a "block") containing audio data (usually for a time period of 1024 or 960 samples), associated information, and / or other data. Herein, we use the term "block" to denote a segment of an MPEG-4 AAC bitstream that includes audio data (and corresponding metadata and optionally also other related data) that determines or indicates one (and not more than one) "raw_data_block" element.

[0030] Each block of an MPEG-4 AAC bitstream can contain several syntax elements (each of which is also materialized as a segment of data in the bitstream). Several types of these syntax elements are defined in the MPEG-4 AAC standard. Each syntax element is identified by a different value of the data element "id_syn_ele". Examples of syntax elements include "single_channel_element()", "channel_pair_element()", and "fill_element()". A single-channel element is a container for audio data containing a single audio channel (a mono audio signal). A channel pair element contains audio data for two audio channels (i.e., a stereo audio signal).

[0031] A stuffing element is a container for information that contains an identifier (e.g., the value of the element "id_syn_ele" mentioned above), followed by data (which is referred to as "stuffing data"). Stuffing elements have historically been used to adjust the instantaneous bit rate of a bit stream to be transmitted over a constant rate passband. By adding an appropriate amount of stuffing data to each block, a constant data rate can be achieved.

[0032] According to an embodiment of the present invention, the stuffing data may include one or more extended payloads that extend the types of data (e.g., metadata) that can be transmitted in the bit stream. Optionally, a decoder that receives the bit stream (e.g., a decoder) may use the decoder that receives the bit stream with stuffing data containing data of a new type to extend the functionality of the device. Therefore, those skilled in the art should understand that a stuffing element is a specific type of data structure and is different from the data structures commonly used to transmit audio data (e.g., an audio payload containing channel data).

[0033] In some embodiments of the present invention, the identifier used to identify a stuffing element may consist of an unsigned integer ("uimsbf") with a value of 0x6 that first transmits the most significant bit in a three-bit format. In a block, several examples of the same type of syntax element (e.g., several stuffing elements) may occur.

[0034] Another standard for encoding an audio bit stream is the MPEG Unified Speech and Audio Coding (USAC) standard (ISO / IEC 23003-3:2012). The MPEG USAC standard describes encoding and decoding audio content using spectral band replication processing (including SBR processing as described in the MPEG-4 AAC standard, and also including other enhanced forms of spectral band replication processing). This processing applies an extended and enhanced version of the spectral band replication tool set of the SBR tools described in the MPEG-4 AAC standard (sometimes referred to herein as the "enhanced SBR tool" or "eSBR tool"). Therefore, eSBR (as defined in the USAC standard) is an improvement over SBR (as defined in the MPEG-4 AAC standard).

[0035] Herein, we use the expression "enhanced SBR processing" (or "eSBR processing") to denote spectral band replication processing that uses at least one eSBR tool not described or mentioned in the MPEG-4 AAC standard (e.g., at least one eSBR tool described or mentioned in the MPEG USAC standard). Examples of such eSBR tools are harmonic transposition and QMF-patched additional preprocessing or "preflattening".

[0036] An integer-order harmonic transposer of order T maps a sine curve with frequency ω to a sine curve with frequency Tω while maintaining the signal duration. Usually, three orders T = 2, 3, 4 are used in sequence to generate each part of the desired output frequency range with the smallest possible transposition order. If an output above the fourth-order transposition range is needed, it can be generated by frequency shifting. When feasible, a near-critical sampled baseband time domain is generated for processing to minimize computational complexity.

[0037] The harmonic transposer can be based on QMF or DFT. When using a QMF-based harmonic transposer, the bandwidth expansion of the core encoder time domain signal is implemented entirely in the QMF domain using a modified phase vocoder structure (which performs decimation followed by time stretching for each QMF subband). Transposition using several transposition factors (e.g., T = 2, 3, 4) is implemented in the common QMF analysis / synthesis transform stage. Since the QMF-based harmonic transposer is not characterized by signal-adaptive frequency domain oversampling, the corresponding flag (sbrOversamplingFlag[ch]) in the bitstream can be ignored.

[0038] When using a DFT-based harmonic transposer, it is preferable to integrate factor 3 and 4 transposers (3rd-order and 4th-order transposers) into a factor 2 transposer (2nd-order transposer) by interpolation to reduce complexity. For each frame (which corresponds to coreCoderFrameLength core encoder samples), the nominal "full-size" transform size of the transposer is first determined by the signal-adaptive frequency domain oversampling flag (sbrOversamplingFlag[ch]) in the bitstream.

[0039] When sbrPatchingMode == 1 (which indicates that linear transposition is to be used to generate the high band), an additional step can be introduced to avoid discontinuities in the shape of the spectral envelope of the high-frequency signal input to the subsequent envelope adjuster. This improves the operation of the subsequent envelope adjustment stage, resulting in a more perceptually stable high-band signal. The operation of the additional preprocessing is beneficial for signal types where the coarse spectral envelope of the low-band signal used for high-frequency reconstruction shows large level variations. However, the value of the bitstream element can be determined in the encoder by applying any kind of signal-dependent classification. The additional preprocessing is preferably initiated by the unitary bitstream element bs_sbr_preprocessing. When bs_sbr_preprocessing is set to 1, the additional preprocessing is enabled. When bs_sbr_preprocessing is set to 0, the additional preprocessing is disabled. The additional processing preferably utilizes a pre-gain curve, which is used by the high-frequency generator to scale the low band X for each tile. Low For example, the pre-gain curve can be calculated according to the following:

[0040] preGain(k) = 10 (meanNrg-lowEnvSlope(k)) / 20 ,0 ≤ k < k0

[0041] where k0 is the first QMF subband in the main frequency band table and lowEnvSlope is calculated using a function (such as polyfit()) that calculates the best polynomial fit coefficients (in the least squares sense). For example,

[0042] polyfit(3, k0, x_lowband, lowEnv, lowEnvSlope);

[0043] can be employed (using a cubic polynomial) and where

[0044]

[0045] where x_lowband(k) = [0...k 0-1 , numTimeSlot is the number of SBR envelope time slots present in a frame, RATE is a constant indicating the number of QMF subband samples per time slot (e.g., 2), are the linear prediction filter coefficients (potentially obtained from an autocovariance method) and where

[0046]

[0047] A bitstream generated according to the MPEG USAC standard (sometimes referred to herein as a "USAC bitstream") contains encoded audio content and typically contains metadata indicating each type of spectral band replication processing to be applied by a decoder to decode the audio content of the USAC bitstream, and / or metadata controlling this spectral band replication processing and / or indicating at least one characteristic or parameter of at least one SBR tool and / or eSBR tool to be used to decode the audio content of the USAC bitstream.

[0048] In this document, we use the expression "enhanced SBR metadata" (or "eSBR metadata") to denote metadata that indicates each type of spectral band replication processing to be applied by a decoder to decode an encoded audio bitstream (such as a USAC bitstream) and / or controls this spectral band replication processing and / or indicates at least one characteristic or parameter of at least one SBR tool and / or eSBR tool to be used to decode this audio content, but is not described or mentioned in the MPEG-4 AAC standard. An example of eSBR metadata is metadata (indicating, or used to control spectral band replication processing) that is described or mentioned in the MPEG USAC standard but not described or mentioned in the MPEG-4 AAC standard. Thus, eSBR metadata in this document represents metadata that is not SBR metadata, and SBR metadata in this document represents metadata that is not eSBR metadata.

[0049] A USAC bitstream can contain both SBR metadata and eSBR metadata. More specifically, a USAC bitstream can contain eSBR metadata that controls the execution of eSBR processing through a decoder, and SBR metadata that controls the execution of SBR processing through a decoder. According to an exemplary embodiment of the present invention, eSBR metadata (such as eSBR specific configuration data) (according to the present invention) is contained in an MPEG-4 AAC bitstream (e.g., in the sbr_extension() container at the end of the SBR payload).

[0050] During decoding of an encoded bitstream using a set of eSBR tools (which includes at least one eSBR tool), the decoder performs eSBR processing to regenerate the high-frequency bands of an audio signal based on the replication of a truncated harmonic sequence during encoding. This eSBR processing typically adjusts the spectral envelope of the generated high-frequency bands and applies inverse filtering, and adds noise and sine wave components to regenerate the spectral characteristics of the original audio signal.

[0051] According to an exemplary embodiment of the present invention, eSBR metadata (e.g., containing a small number of control bits that are eSBR metadata) is contained in one or more of the metadata segments of an encoded audio bitstream (such as an MPEG-4 AAC bitstream), which also contains encoded audio data in other segments (audio data segments). Generally, at least one such metadata segment of each block of the bitstream is (or contains) a padding element (which contains an identifier indicating the start of the padding element), and the eSBR metadata is contained in the padding element after the identifier.

[0052] Figure 1Block diagram of an exemplary audio processing chain (audio data processing system) in which one or more elements of the system may be configured in accordance with embodiments of the present invention. The system includes the following elements coupled together as shown: an encoder 1, a delivery subsystem 2, a decoder 3, and a post-processing unit 4. In variants of the system shown, one or more elements are omitted, or additional audio data processing units are included.

[0053] In some embodiments, the encoder 1 (which optionally includes a preprocessing unit) is configured to accept PCM (time domain) samples including audio content as input and output an encoded audio bitstream indicative of the audio content (which has a format compatible with the MPEG-4 AAC standard). The data of the bitstream indicative of the audio content is sometimes referred to herein as “audio data” or “encoded audio data”. If the encoder is configured in accordance with typical embodiments of the present invention, the audio bitstream output from the encoder includes eSBR metadata (and usually also other metadata) as well as audio data.

[0054] One or more encoded audio bitstreams output from the encoder 1 may be delivered to the encoded audio delivery subsystem 2. The subsystem 2 is configured to store and / or deliver each encoded bitstream output from the encoder 1. The encoded audio bitstreams output from the encoder 1 may be stored by the subsystem 2 (e.g., in the form of a DVD or Blu-ray disc), or transmitted by the subsystem 2 (which may implement a transmission link or network), or both stored and transmitted by the subsystem 2.

[0055] The decoder 3 is configured to decode the encoded MPEG-4 AAC audio bitstream (produced by the encoder 1) that it receives via the subsystem 2. In some embodiments, the decoder 3 is configured to extract eSBR metadata from each block of the bitstream and decode the bitstream (including performing eSBR processing by using the extracted eSBR metadata) to produce decoded audio data (e.g., a stream of decoded PCM audio samples). In some embodiments, the decoder 3 is configured to extract SBR metadata from the bitstream (but ignoring the eSBR metadata included in the bitstream) and decode the bitstream (including performing SBR processing by using the extracted SBR metadata) to produce decoded audio data (e.g., a stream of decoded PCM audio samples). Generally, the decoder 3 includes (e.g., in a non-transitory manner) a buffer that stores segments of the encoded audio bitstream received from the subsystem 2.

[0056] Figure 1 The post-processing unit 4 is configured to accept the decoded audio data stream (e.g., decoded PCM audio samples) from the decoder 3 and perform post-processing on it. The post-processing unit may also be configured to present the post-processed audio content (or the decoded audio received from the decoder 3) for playback by one or more speakers.

[0057] Figure 2It is a block diagram of an encoder 100 which is an embodiment of the audio processing unit of the present invention. Any component or element of the encoder 100 can be implemented as one or more processes and / or one or more circuits (such as ASIC, FPGA or other integrated circuits) in hardware, software or a combination of hardware and software. The encoder 100 includes an encoder 105, a filler / formatter stage 107, a metadata generation stage 106 and a buffer memory 109 connected as shown. Generally, the encoder 100 also includes other processing elements (not shown). The encoder 100 is configured to convert an input audio bitstream into an encoded output MPEG-4 AAC bitstream.

[0058] The metadata generator 106 is coupled and configured to generate metadata (including eSBR metadata and SBR metadata) (and / or pass the metadata to stage 107) to be included in the encoded bitstream to be output from the encoder 100 through stage 107.

[0059] The encoder 105 is coupled and configured to encode the input audio data (e.g., by performing compression on it), and confirm the resulting encoded audio to stage 107 to be included in the encoded bitstream to be output from stage 107.

[0060] Stage 107 is configured to multiplex the encoded audio from the encoder 105 and the metadata from the generator 106 (including eSBR metadata and SBR metadata) to generate an encoded bitstream to be output from stage 107, preferably such that the encoded bitstream has a format specified by one of the embodiments of the present invention.

[0061] The buffer memory 109 is configured to store (e.g., in a non-transitory manner) at least one block of the encoded audio bitstream output from stage 107, and then confirm a sequence block of the encoded audio bitstream from the buffer memory 109 as the output from the encoder 100 to the delivery system.

[0062] Figure 3 It is a block diagram of a system which includes a decoder 200 which is an embodiment of the audio processing unit of the invention and optionally a post-processor 300 also coupled to the decoder 200. Any component or element of the decoder 200 and the post-processor 300 can be implemented as one or more processes and / or one or more circuits (such as ASIC, FPGA or other integrated circuits) in hardware, software or a combination of hardware and software. The decoder 200 includes a buffer memory 201, a bitstream payload de-formatter (parser) 205, an audio decoding subsystem 202 (sometimes referred to as the "core" decoding stage or "core" decoding subsystem), an eSBR processing stage 203 and a control bit generation stage 204 connected as shown. Generally, the decoder 200 also includes other processing elements (not shown).

[0063] The buffer memory (buffer) 201 stores, for example in a non-transitory manner, at least one block of the encoded MPEG-4 AAC audio bitstream received by the decoder 200. During the operation of the decoder 200, a sequence block of the bitstream is verified from the buffer 201 to the de-formatter 205.

[0064] In Figure 3 the embodiment (or the embodiment to be described Figure 4 ), in a variant of the embodiment), the APU that is not a decoder (for example, Figure 6 the APU 500 of Figure 3 or Figure 4 contains a buffer memory (for example the same buffer memory as the buffer 201), which stores, for example in a non-transitory manner, at least one block of the same type of encoded audio bitstream (for example, an MPEG-4 AAC audio bitstream) (i.e., an encoded audio bitstream containing eSBR metadata) received by

[0065] Referring again to Figure 3 , the de-formatter 205 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (including quantized envelope data) and eSBR metadata (and generally also other metadata) therefrom, verify at least the eSBR metadata and the SBR metadata to the eSBR processing stage 203, and generally also verify other extracted metadata to the decoding subsystem 202 (and optionally also to the control bit generator 204). The de-formatter 205 is also coupled and configured to extract audio data from each block of the bitstream, and verify the extracted audio data to the decoding subsystem (decoding stage) 202.

[0066] Figure 3 The system of

[0067] The audio decoding subsystem 202 of decoder 200 is configured to decode the audio data extracted by parser 205 (this decoding may be referred to as the "core" decoding operation) to produce decoded audio data, and confirm the decoded audio data to the eSBR processing stage 203. The decoding is performed in the frequency domain and generally includes inverse quantization, followed by spectral processing. Generally, the last stage of processing in subsystem 202 applies a frequency-domain to time-domain transform to the decoded frequency-domain audio data, such that the output of the subsystem is time-domain, decoded audio data. Stage 203 is configured to apply the SBR tool and eSBR tool indicated by the eSBR metadata and eSBR (extracted by parser 205) to the decoded audio data (i.e., perform SBR and eSBR processing on the output of decoding subsystem 202 using the SBR and eSBR metadata) to produce fully decoded audio data output from decoder 200 (e.g., to post-processor 300). Generally, decoder 200 includes a memory that stores the de-formatted audio data and metadata output from de-formatter 205 (accessible by subsystem 202 and stage 203), and stage 203 is configured to access the audio data and metadata (including SBR metadata and eSBR metadata) as needed during SBR and eSBR processing. The SBR processing and eSBR processing in stage 203 can be regarded as post-processing of the output of the core decoding subsystem 202. Optionally, decoder 200 further includes a final upmixing subsystem (which may apply the parametric stereo ("PS") tool defined in the MPEG-4 AAC standard, using the PS metadata extracted by de-formatter 205 and / or control bits generated in subsystem 204), the final upmixing subsystem being coupled and configured to perform upmixing on the output of stage 203 to produce fully decoded upmixed audio output from decoder 200. Alternatively, post-processor 300 is configured to perform upmixing on the output of decoder 200 (e.g., using the PS metadata extracted by de-formatter 205 and / or control bits generated in subsystem 204).

[0068] In response to metadata extracted by the de-formatter 205, the control bit generator 204 may generate control data, and the control data may be used within the decoder 200 (e.g., in the final upmixing system) and / or verified as an output of the decoder 200 (e.g., to the post-processor 300 for use in post-processing). In response to metadata extracted from the input bitstream (and optionally also in response to control data), stage 204 may generate (and verify to the post-processor 300) control bits indicating that the decoded audio data output from the eSBR processing stage 203 should undergo a particular type of post-processing. In some embodiments, the decoder 200 is configured to verify metadata extracted by the de-formatter 205 from the input bitstream to the post-processor 300, and the post-processor 300 is configured to use the metadata to perform post-processing on the decoded audio data output from the decoder 200.

[0069] Figure 4 is a block diagram of an audio processing unit (“APU”) (210) that is another embodiment of the audio processing unit of the present invention. The APU 210 is a legacy decoder that is not configured to perform eSBR processing. Any component or element of the APU 210 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware and software). The APU 210 includes a buffer memory 201, a bitstream payload de-formatter (parser) 215, an audio decoding subsystem 202 (sometimes referred to as the “core” decoding stage or “core” decoding subsystem), and an SBR processing stage 213 connected as shown. Generally, the APU 210 also contains other processing elements (not shown). The APU 210 may represent, for example, an audio encoder, decoder, or codec.

[0070] Elements 201 and 202 of the APU 210 are the same as the similarly numbered elements of the Figure 3 decoder 200 and will not repeat their description above. In operation of the APU 210, sequence blocks of the encoded audio bitstream (MPEG-4 AAC bitstream) received by the APU 210 are verified from the buffer 201 to the de-formatter 215.

[0071] According to any embodiment of the present invention, the de-formatter 215 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (including quantized envelope data) and generally also to extract other metadata therefrom, but to ignore eSBR metadata that may be included in the bitstream. The de-formatter 215 is configured to verify at least the SBR metadata to the SBR processing stage 213. The de-formatter 215 is also coupled and configured to extract audio data from each block of the bitstream, and to verify the extracted audio data to the decoding subsystem (decoding stage) 202.

[0072] The audio decoding subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the de-formatter 215 (this decoding may be referred to as the "core" decoding operation) to generate decoded audio data, and confirm the decoded audio data to the SBR processing stage 213. The decoding is performed in the time domain. Generally, the last stage of the processing in the subsystem 202 applies an inverse frequency-domain to time-domain transform to the decoded frequency-domain audio data, such that the output of the subsystem is time-domain, decoded audio data. The stage 213 is configured to apply the SBR tool (rather than the eSBR tool) indicated by the SBR metadata (extracted by the parser 215) to the decoded audio data (i.e., perform SBR processing on the output of the decoding subsystem 202 using the SBR metadata) to generate the fully decoded audio data output from the APU 210 (e.g., to the post-processor 300). Generally, the APU 210 includes a memory that stores the de-formatted audio data and metadata output from the de-formatter 215 (which can be accessed by the subsystem 202 and the stage 213), and the stage 213 is configured to access the audio data and metadata (including SBR metadata) as needed during the SBR processing. The SBR processing in the stage 213 can be regarded as post-processing of the output of the core decoding subsystem 202. Optionally, the APU 210 further includes a final upmixing subsystem (which can apply the parametric stereo ("PS") tool defined in the MPEG-4 AAC standard, using the PS metadata extracted by the de-formatter 215), and the final upmixing subsystem is coupled and configured to perform upmixing on the output of the stage 213 to generate the fully decoded upmixed audio output from the APU 210. Alternatively, the post-processor is configured to perform upmixing on the output of the APU 210 (e.g., using the PS metadata extracted by the de-formatter 215 and / or the control bits generated in the APU 210).

[0073] Various embodiments of the encoder 100, the decoder 200, and the APU 210 are configured to perform different embodiments of the method of the present invention.

[0074] According to some embodiments, eSBR metadata (e.g., a small number of control bits that are eSBR metadata) is included in an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) such that legacy decoders (which are not configured to parse eSBR metadata or use any eSBR tools associated with eSBR metadata) can ignore the eSBR metadata but still decode the bitstream as well as possible without using the eSBR metadata or any eSBR tools associated with the eSBR metadata, generally without any significant loss of decoded audio quality. However, an eSBR decoder that is configured to parse the bitstream to identify the eSBR metadata and use at least one eSBR tool in response to the eSBR metadata will enjoy the benefits of using at least one such eSBR tool. Thus, embodiments of the present invention provide a means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner.

[0075] Generally, the eSBR metadata in the bitstream indicates one or more of the following eSBR tools (which are described in the MPEG USAC standard and which may or may not be applied by the encoder during generation of the bitstream) (e.g., indicates at least one characteristic or parameter of one or more of the following eSBR tools):

[0076] · Harmonic transposition; and

[0077] · QMF-patching extra preprocessing (pre-flattening).

[0078] For example, the eSBR metadata included in the bitstream may indicate the values of parameters (described in the MPEG USAC standard and in the present invention): sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch], and bs_sbr_preprocessing.

[0079] Herein, the notation X[ch] (where X is a certain parameter) indicates that the parameter is related to the channel ("ch") of the audio content of the encoded bitstream to be decoded. For the sake of brevity, we sometimes omit the expression [ch] and assume that the relevant parameter is related to the channel of the audio content.

[0080] Herein, the notation X[ch][env] (where X is a certain parameter) indicates that the parameter is related to the SBR envelope ("env") of the channel ("ch") of the audio content of the encoded bitstream to be decoded. For the sake of brevity, we sometimes omit the expressions [env] and [ch] and assume that the relevant parameter is related to the SBR envelope of the channel of the audio content.

[0081] During the decoding of an encoded bitstream, the execution of harmonic transposition during the eSBR processing stage of decoding (for each channel "ch" of the audio content indicated by the bitstream) is controlled by the following eSBR metadata parameters: sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBinsFlag[ch], and sbrPitchInBins[ch].

[0082] The value "sbrPatchingMode[ch]" indicates the type of transposer used in eSBR: sbrPatchingMode[ch] = 1 indicates linear transposition patching as described in Section 4.6.18 of the MPEG-4 AAC standard (when used with high-quality SBR or low-power SBR); sbrPatchingMode[ch] = 0 indicates harmonic SBR patching as described in Sections 7.5.3 or 7.5.4 of the MPEG USAC standard.

[0083] The value "sbrOversamplingFlag[ch]" indicates the use of signal-adaptive frequency-domain oversampling in eSBR in combination with DFT-based harmonic SBR patching, as described in Section 7.5.3 of the MPEG USAC standard. This flag controls the size of the DFT utilized in the transposer: 1 indicates that signal-adaptive frequency-domain oversampling is enabled, as described in Section 7.5.3.1 of the MPEG USAC standard; 0 indicates that signal-adaptive frequency-domain oversampling is disabled, as described in Section 7.5.3.1 of the MPEG USAC standard.

[0084] The value "sbrPitchInBinsFlag[ch]" controls the interpretation of the sbrPitchInBins[ch] parameter: 1 indicates that the value in sbrPitchInBins[ch] is valid and greater than zero; 0 indicates that the value of sbrPitchInBins[ch] is set to zero.

[0085] The value "sbrPitchInBins[ch]" controls the addition of cross-product terms in the SBR harmonic transposer. The value sbrPitchinBins[ch] is an integer value in the range [0, 127] and represents the distance measured in frequency bins for a 1536-line DFT of the sampling frequency applied to the core encoder.

[0086] In the case of an SBR channel pair (as opposed to a single SBR channel) where the MPEG-4 AAC bitstream indicates that its channels are not coupled, the bitstream indicates two instances of the above syntax (for harmonic or non-harmonic transforms), one for each channel of sbr_channel_pair_element().

[0087] The harmonic transposition of the eSBR tool generally improves the quality of the decoded music signal at relatively low crossover frequencies. Non-harmonic transposition (i.e., the old type of spectral patching) generally improves the speech signal. Therefore, the starting point in deciding which type of transposition is preferred for encoding a particular audio content is to select a transposition method based on speech / music detection, with harmonic transposition being used for music content and spectral patching being used for speech content.

[0088] The performance of pre-flattening during eSBR processing is controlled by the value of a single eSBR metadata parameter called "bs_sbr_preprocessing", meaning that pre-flattening is performed or not depending on the value of this single bit. When using an SBR QMF-patching algorithm as described in Section 4.6.18.6.3 of the MPEG-4 AAC standard, a pre-flattening step may be performed (when indicated by the "bs_sbr_preprocessing" parameter) in an effort to avoid discontinuities in the shape of the spectral envelope of the high-frequency signal that is input to a subsequent envelope adjuster (which performs another stage of eSBR processing). Pre-flattening generally improves the operation of the subsequent envelope adjustment stage, resulting in a high-band signal that is perceived as more stable.

[0089] The overall bitrate requirement for including eSBR metadata indicating the above-mentioned eSBR tools (harmonic transposition and pre-flattening) in the MPEG-4 AAC bitstream is expected to be on the order of several hundred bits per second, because according to some embodiments of the present invention, only the differential control data required to perform the eSBR processing is transmitted. Legacy decoders can ignore this information because it is included in a backward-compatible manner (as will be explained later). Therefore, the adverse impact on bitrate associated with including eSBR metadata is negligible for several reasons, including the following:

[0090] The bitrate loss (due to the inclusion of the eSBR metadata) is a very small fraction of the total bitrate, since only the differential control data required to perform the eSBR processing is transmitted (and not the simulcast of the SBR control data); and

[0091] • The tuning of SBR-related control information does not typically depend on the details of the transposition. Examples of when control data depends on the operation of the transposer are discussed later in this application.

[0092] Accordingly, embodiments of the present invention provide a means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner. This efficient transmission of eSBR control data reduces the memory requirements in decoders, encoders, and codecs that employ aspects of the present invention, while having no tangible adverse effect on the bitrate. Additionally, the complexity and processing requirements associated with performing eSBR according to embodiments of the present invention are reduced because the SBR data only needs to be processed once and is not re-broadcast, as would be the case if eSBR were treated as a completely separate object type in MPEG-4 AAC rather than being integrated into the MPEG-4 AAC codec in a backward-compatible manner.

[0093] Next, with reference to Figure 7 , we describe the elements of a block ("raw_data_block") of an MPEG-4 AAC bitstream that includes eSBR metadata according to some embodiments of the present invention. Figure 7 is a diagram of a block ("raw_data_block") of an MPEG-4 AAC bitstream that shows some segments of the block.

[0094] A block of an MPEG-4 AAC bitstream can include at least one "single_channel_element()" (such as the single-channel element shown in Figure 7 ) and / or at least one "channel_pair_element()" (although it may be present, it is not specifically shown in Figure 7 ), which includes audio data for an audio program. The block can also include several "fill_elements" (e.g., fill element 1 and / or fill element 2 of Figure 7 ), which include program-related data (e.g., metadata). Each "single_channel_element()" includes an identifier (e.g., "ID1" of Figure 7 ) that indicates the start of the single-channel element, and can include audio data for different channels of a multichannel audio program. Each "channel_pair_element()" includes an identifier (not shown in Figure 7 ) that indicates the start of the channel-pair element, and can include audio data for two channels of the program.

[0095] An MPEG-4 AAC bitstream fill_element (referred to herein as a fill element) includes an identifier ( Figure 7The "ID2") and the padding data after the identifier. The identifier ID2 can consist of an unsigned integer in three bits with the value of 0x6, transmitting the most significant bit first ("uimsbf"). The padding data can contain an extension_payload() element (sometimes referred to as an extended payload in this article), and its syntax is shown in Table 4.57 of the MPEG-4 AAC standard. There are several types of extended payloads and they are identified by the "extension_type" parameter, which is an unsigned integer in four bits transmitting the most significant bit first ("uimsbf").

[0096] The padding data (e.g., its extended payload) can contain a header or an identifier (e.g., Figure 7 the "header 1"), and the header or identifier indicates a segment of the padding data that indicates an SBR object (i.e., the header initializes the "SBR object" type, which is called sbr_extension_data() in the MPEG-4 AAC standard). For example, the spectral band replication (SBR) extended payload is identified by having a value of '1101' or '1110' for the extension_type field in the header, where the identifier '1101' identifies the extended payload with SBR data and '1110' identifies the extended payload with SBR data that has a cyclic redundancy check (CRC) to verify the correctness of the SBR data.

[0097] When the header (e.g., the extension_type field) initializes the SBR object type, the SBR metadata (sometimes referred to as "spectral band replication data" in this article and called sbr_data() in the MPEG-4 AAC standard) follows the header, and at least one spectral band replication extension element (e.g., Figure 7 the "SBR extension element" of padding element 1) can follow the SBR metadata. This spectral band replication extension element (a segment of the bitstream) is called the "sbr_extension()" container in the MPEG-4 AAC standard. The spectral band replication extension element optionally contains a header (e.g., Figure 7 the "SBR extension header" of padding element 1).

[0098] The MPEG-4 AAC standard envisages that the spectral band replication extension element can contain PS (parametric stereo) data for the audio data of the program. The MPEG-4 AAC standard envisages that when the header of the padding element (e.g., its extended payload) initializes the SBR object type (as Figure 7When the "header 1") of and the spectral band replication extended element of the padding element contains PS data, the padding element (e.g., its extended payload) contains spectral band replication data and a "bs_extension_id" parameter, and the value of the "bs_extension_id" parameter (i.e., bs_extension_id = 2) indicates that the PS data is contained in the spectral band replication extended element of the padding element.

[0099] According to some embodiments of the present invention, eSBR metadata (e.g., a flag indicating whether enhanced spectral band replication (eSBR) processing is performed on the audio content of a block) is contained in the spectral band replication extended element of the padding element. For example, in Figure 7 this flag is indicated in padding element 1, where the flag appears after the header of the "SBR extended element" of padding element 1 (the "SBR extended header" of padding element 1). Optionally, this flag and additional eSBR metadata are contained in the spectral band replication extended element after the header of the spectral band replication extended element (e.g., in Figure 7 the SBR extended element of padding element 1 in, after the SBR extended header). According to some embodiments of the present invention, the padding element containing eSBR metadata also contains a "bs_extension_id" parameter, and the value of the "bs_extension_id" parameter (e.g., bs_extension_id = 3) indicates that the eSBR metadata is contained in the padding element and eSBR processing is to be performed on the audio content of the relevant block.

[0100] According to some embodiments of the present invention, eSBR metadata is contained in a padding element (e.g., Figure 7 padding element 2) of an MPEG-4 AAC bitstream rather than in the spectral band replication extended element (SBR extended element) of the padding element. This is because a padding element containing extension_payload() with SBR data or SBR data with CRC does not contain any other extended payload of any other extension type. Therefore, in embodiments where the eSBR metadata stores its own extended payload, a separate padding element is used to store the eSBR metadata. This padding element contains an identifier indicating the start of the padding element (e.g., Figure 7 "ID2") and the padding data after the identifier. The padding data may contain an extension_payload() element (sometimes referred to herein as an extended payload), and its syntax is shown in Table 4.57 of the MPEG-4 AAC standard. The padding data (e.g., its extended payload) contains a header indicating the eSBR object ( Figure 7("header 2" of padding element 2), i.e., header initializes enhanced spectral band replication (eSBR) object type, and padding data (e.g., its extended payload) contains eSBR metadata after the header. For example, Figure 7 padding element 2 contains this header ("header 2") and also contains eSBR metadata after the header (i.e., the "flag" in padding element 2, which indicates whether enhanced spectral band replication (eSBR) processing is performed on the audio content of the block). Optionally, in Figure padding data of padding element 2 also contains additional eSBR metadata after header 2. In the embodiments described in this paragraph, the header (e.g., ​ header 2) has an identification value that is not one of the conventional values specified in Table 4.57 of the MPEG-4 AAC standard, and instead indicates the eSBR extended payload (such that the extension_type field of the header indicates that the padding data contains eSBR metadata).

[0101] In a first category of embodiments, the present invention is an audio processing unit (e.g., a decoder) that includes:

[0102] a memory (e.g., ​ buffer 201 of or 4), which is configured to store at least one block of an encoded bitstream (e.g., at least one block of an MPEG-4 AAC bitstream);

[0103] a bitstream payload de-formatter (e.g., ​ element 205 of or ​ element 215 of), which is coupled to the memory and is configured to demultiplex at least one part of the block of the bitstream; and

[0104] a decoding subsystem (e.g., ​ elements 202 and 20'3 of or ​ elements 202 and 213 of), which is coupled and configured to decode at least one part of the audio content of the block of the bitstream, where the block includes:

[0105] a padding element, which contains an identifier indicating the start of the padding element (e.g., the "id_syn_ele" identifier with a value of 0x6 in Table 4.85 of the MPEG-4 AAC standard) and padding data after the identifier, where the padding data includes:

[0106] at least one flag, which identifies whether enhanced spectral band replication (eSBR) processing is performed on the audio content of the block (e.g., using spectral band replication data and eSBR metadata included in the block).

[0107] The flag is eSBR metadata, and an instance of the flag is the sbrPatchingMode flag. Another instance of the flag is the harmonicSBR flag. Both of these flags indicate whether to perform a basic form of spectral band replication or an enhanced form of spectral replication on the audio data of the block. The basic form of spectral replication is spectral patching, and the enhanced form of spectral band replication is harmonic transposition.

[0108] In some embodiments, the padding data further includes additional eSBR metadata (i.e., eSBR metadata other than the flag).

[0109] The memory can be (e.g., in a non-transitory manner) a buffer memory (e.g., ​ an implementation of buffer 201) that stores the at least one block of the encoded audio bitstream.

[0110] It is estimated that the complexity of performing eSBR processing during the decoding of an MPEG-4 AAC bitstream containing eSBR metadata (indicating these eSBR tools) by an eSBR decoder (using eSBR harmonic transposition and pre-flattening) will be as follows (for typical decoding using the indicated parameters):

[0111] · Harmonic transposition (16 kbps, 14400 / 28800 Hz)

[0112] ○ Based on DFT: 3.68 WMOPS (weighted million operations per second);

[0113] ○ Based on QMF: 0.98 WMOPS;

[0114] · QMF patching preprocessing (pre-flattening): 0.1 WMOPS.

[0115] It is known that DFT-based transposition generally performs better than QMF-based transposition for transients.

[0116] According to some embodiments of the present invention, the padding element (of the encoded audio bitstream) containing the eSBR metadata further contains a parameter (e.g., "bs_extension_id" parameter) whose value (e.g., bs_extension_id = 3) signals that the eSBR metadata is contained in the padding element and that eSBR processing is to be performed on the audio content of the associated block, and / or a parameter (e.g., "bs_extension_id" parameter) whose value (e.g., bs_extension_id = 2) signals that the sbr_extension() container of the padding element contains PS data. For example, as indicated in Table 1 below, this parameter with the value bs_extension_id = 2 can signal that the sbr_extension() container of the padding element contains PS data, and this parameter with the value bs_extension_id = 3 can signal that the sbr_extension() container of the padding element contains eSBR metadata:

[0117] Table 1

[0118]

[0119]

[0120] According to some embodiments of the present invention, the syntax of each spectral band replication extension element containing eSBR metadata and / or PS data is as indicated in Table 2 below (where "sbr_extension()" represents the container that is the spectral band replication extension element, "bs_extension_id" is as described in Table 1 above, "ps_data" represents PS data, and "esbr_data" represents eSBR metadata):

[0121] Table 2

[0122]

[0123] In an exemplary embodiment, the esbr_data() mentioned above in ​ indicates the values of the following metadata parameters:

[0124] 1. The bit data parameter "bs_sbr_preprocessing"; and

[0125] 2. For each channel ("ch") of the audio content of the encoded bitstream to be decoded, each of the parameters described above is: "sbrPatchingMode[ch]"; "sbrOversamplingFlag[ch]"; "sbrPitchInBinsFlag[ch]"; and "sbrPitchInBins[ch]".

[0126] For example, in some embodiments, esbr_data() may have the syntax indicated in Table 3 to indicate these metadata parameters:

[0127] Table 3

[0128]

[0129]

[0130] The above syntax enables an efficient implementation of an enhanced form of spectral band replication (e.g., harmonic transposition), as an extension to legacy decoders. Specifically, the eSBR data in Table 3 contains only those parameters required to perform the enhanced form of spectral band replication that are not already supported in the bitstream or not directly derivable from pre-existing parameters in the bitstream. All other parameters and processing data required to perform the enhanced form of spectral band replication are extracted from pre-existing parameters in positions already defined in the bitstream.

[0131] For example, an extensible MPEG-4 HE-AAC or HE-AAC v2 compliant decoder can be enhanced to include an enhanced form of spectral band replication, such as harmonic transposition. This enhanced form of spectral band replication is supplementary to the basic form of spectral band replication already supported by the decoder. In the context of an MPEG-4 HE-AAC or HE-AAC v2 compliant decoder, this basic form of spectral band replication is the QMF spectral patching SBR tool as defined in Section 4.6.18 of the MPEG-4 AAC standard.

[0132] When performing the enhanced form of spectral band replication, the extended HE-AAC decoder can reuse many bitstream parameters that are already included in the SBR extension payload of the bitstream. Specific parameters that can be reused include, for example, various parameters that determine the main frequency band table. These parameters include bs_start_freq (the parameter that determines the start of the main frequency table parameters), bs_stop_freq (the parameter that determines the stop of the main frequency table), bs_freq_scale (the parameter that determines the number of bands per octave), and bs_alter_scale (the parameter that changes the scale of the bands). The parameters that can be reused also include the parameters that determine the noise band table (bs_noise_bands) and the limiter band table parameters (bs_limiter_bands). Thus, in various embodiments, at least some equivalent parameters specified in the USAC standard are omitted from the bitstream, thereby reducing the control overhead in the bitstream. Generally, where a parameter specified in the AAC standard has an equivalent parameter specified in the USAC standard, the equivalent parameter specified in the USAC standard has the same name as the parameter specified in the AAC standard, for example, the envelope scale factor EOrigMapped However, the equivalent parameters specified in the USAC standard typically have different values, which are "tuned" for the enhanced SBR processing defined in the USAC standard rather than for the SBR processing defined in the AAC standard.

[0133] To improve the subjective quality of audio content with a harmonic frequency structure and strong tonal characteristics, especially at low bitrates, the activation of enhanced SBR is recommended. The values of the corresponding bitstream elements (i.e., esbr_data()) that control these tools can be determined in the encoder by applying a signal-dependent classification mechanism. In general, the use of the harmonic patching method (sbrPatchingMode == 1) is preferred for encoding music signals at very low bitrates, where the core codec can be significantly limited in terms of audio bandwidth. This is especially the case if these signals contain a significant harmonic structure. Conversely, the use of the conventional SBR patching method is preferred for speech and mixed signals because it provides a preferred preservation of the temporal structure in speech.

[0134] To improve the performance of the harmonic transposer, a preprocessing step (bs_sbr_preprocessing == 1) can be activated, which endeavors to avoid introducing spectral discontinuities in the signal into the subsequent envelope adjuster. The operation of the tool is beneficial for signal types where the coarse spectral envelope of the low-frequency band signal used for high-frequency reconstruction shows large level variations.

[0135] To improve the transient response of harmonic SBR patching, signal-adaptive frequency-domain oversampling (sbrOversamplingFlag == 1) can be applied. Since signal-adaptive frequency-domain oversampling increases the computational complexity of the transposer but only benefits frames containing transients, the use of this tool is controlled by a bitstream element that is transmitted once per frame and per independent SBR channel.

[0136] Decoders operating in the proposed enhanced SBR mode typically need to be able to switch between legacy SBR patching and enhanced SBR patching. Therefore, a delay can be introduced, which, depending on the decoder settings, can be as long as the duration of one core audio frame. In general, the delay for legacy SBR patching and for enhanced SBR patching will be similar.

[0137] In addition to many parameters, when performing an enhanced form of spectral band replication according to an embodiment of the present invention, the extended HE-AAC decoder can also reuse other data elements. For example, envelope data and noise data can also be extracted from the bs_data_env (envelope scale factor) and bs_noise_env (noise floor scale factor) data and used during the enhanced form of spectral band replication.

[0138] In essence, these embodiments utilize configuration parameters and envelope data already supported by legacy HE-AAC or HE-AAC v2 decoders in the SBR extension payload to implement an enhanced form of spectral band replication that requires as little additional transmitted data as possible. The metadata is initially tuned for a base form of HFR (such as the spectral translation operation of SBR), but according to embodiments, the metadata is used for an enhanced form of HFR (such as the harmonic transposition of eSBR). As previously discussed, the metadata generally represents operation parameters (such as envelope scale factors, noise floor scale factors, time / frequency grid parameters, sine addition information, variable crossover frequencies / bands, inverse filter modes, envelope resolution, smoothing modes, frequency interpolation modes) that are tuned and intended to be used with a base form of HFR (such as linear spectral translation). However, this metadata in combination with additional metadata parameters specific to an enhanced form of HFR (such as harmonic transposition) can be used to efficiently and effectively process audio data using the enhanced form of HFR.

[0139] Accordingly, an extended decoder that supports an enhanced form of spectral band replication can be generated in a very efficient manner by relying on already defined bitstream elements (such as bitstream elements in the SBR extension payload) and (in the fill element extension payload) only adding the parameters required to support the enhanced form of spectral band replication. This data reduction feature combined with placing the newly added parameters in a reserved data field (such as an extension container) substantially reduces the barrier to generating a decoder that supports an enhanced form of spectral band replication by ensuring bitstream backward compatibility with legacy decoders that do not support the enhanced form of spectral band replication. It will be appreciated that the reserved data field is a backward compatibility data field that is a data field already supported by earlier decoders (such as legacy HE-AAC or HE-AAC v2 decoders). Similarly, the extension container is backward compatible, which is an extension container already supported by earlier decoders (such as legacy HE-AAC or HE-AAC v2 decoders).

[0140] In Table 3, the numbers in the right column indicate the number of bits of the corresponding parameters in the left column.

[0141] In some embodiments, the SBR object type defined in MPEG-4 AAC is updated to include aspects of SBR tools and enhanced SBR (eSBR) tools signaled in an SBR extension element (bs_extension_id == EXTENSION_ID_ESBR). If a decoder detects this SBR extension element, then the decoder employs the signaled aspects of the enhanced SBR tools.

[0142] In some embodiments, the present invention is a method comprising the steps of: encoding audio data to produce an encoded bitstream (e.g., an MPEG-4 AAC bitstream) that includes including eSBR metadata in at least one segment of at least one block of the encoded bitstream and audio data in at least another segment of the block. In an exemplary embodiment, the method includes the step of multiplexing audio data using eSBR metadata in each block of the encoded bitstream. In a typical decoding of the encoded bitstream in an eSBR decoder, the decoder (including by parsing and demultiplexing the eSBR metadata and the audio data) extracts the eSBR metadata from the bitstream and uses the eSBR metadata to process the audio data to produce a decoded audio data stream.

[0143] Another aspect of the present invention is an eSBR decoder configured to perform eSBR processing during decoding of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that does not include eSBR metadata (e.g., using at least one of eSBR tools referred to as harmonic transposition or pre-flattening). An example of this decoder will be described with reference to ​ the eSBR decoder (400) shown in

[0144] ​ includes a buffer memory 201 connected as shown (which is the same as the memory 201 of ​ and 4 ), a bitstream payload de-formatter 215 (which is the same as the de-formatter 215 of ​ ), an audio decoding subsystem 202 (which is sometimes referred to as the "core" decoding stage or "core" decoding subsystem and is the same as the core decoding subsystem 202 of ​ ), an eSBR control data generation subsystem 401, and an eSBR processing stage 203 (which is the same as the stage 203 of ​ ). Generally, the decoder 400 also includes other processing elements (not shown).

[0145] In the operation of the decoder 400, a sequence block of the encoded audio bitstream (MPEG-4 AAC bitstream) received by the decoder 400 is verified from the buffer 201 to the de-formatter 215.

[0146] The de-formatter 215 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (including quantized envelope data) and usually also other metadata therefrom. The de-formatter 215 is also configured to verify at least the SBR metadata to the eSBR processing stage 203. The de-formatter 215 is also coupled and configured to extract audio data from each block of the bitstream and verify the extracted audio data to the decoding subsystem (decoding stage) 202.

[0147] The audio decoding subsystem 202 of the decoder 400 is configured to decode the audio data extracted by the de-formatter 215 (this decoding may be referred to as the "core" decoding operation) to generate decoded audio data, and confirm the decoded audio data to the eSBR processing stage 203. The decoding is performed in the frequency domain. Generally, the last stage of the processing in the subsystem 202 applies an inverse frequency-domain to time-domain transform to the decoded frequency-domain audio data, such that the output of the subsystem is time-domain, decoded audio data. The stage 203 is configured to apply the SBR tools (and eSBR tools) indicated by the SBR metadata (extracted by the de-formatter 215) and the eSBR metadata generated in the subsystem 401 to the decoded audio data (i.e., to perform SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata) to generate the fully decoded audio data output from the decoder 400. Generally, the decoder 400 includes a memory that stores the de-formatted audio data and metadata output from the de-formatter 215 (and optionally also the subsystem 401), which can be accessed by the subsystem 202 and the stage 203, and the stage 203 is configured to access the audio data and metadata as needed during the SBR and eSBR processing. The SBR processing in the stage 203 can be regarded as post-processing of the output of the core decoding subsystem 202. Optionally, the decoder 400 further includes a final upmixing subsystem (which can apply the parametric stereo ("PS") tool defined in the MPEG-4 AAC standard, using the PS metadata extracted by the de-formatter 215), and the final upmixing subsystem is coupled and configured to perform upmixing on the output of the stage 203 to generate the fully decoded and upmixed audio output from the APU 210.

[0148] Parametric stereo is an encoding tool that represents a stereo signal using the downmixing of the left and right channels of the stereo signal and a set of spatial parameters that describe the stereo image. Parametric stereo typically employs three types of spatial parameters: (1) the inter-channel intensity difference (IID) that describes the intensity difference between channels; (2) the inter-channel phase difference (IPD) that describes the phase difference between channels; and (3) the inter-channel coherence (ICC) that describes the coherence (or similarity) between channels. Coherence can be measured as the maximum of the cross-correlation that varies according to time or phase. These three parameters typically enable a high-quality reconstruction of the stereo image. However, the IPD parameter only specifies the relative phase difference between the channels of the stereo input signal and does not indicate the distribution of these phase differences within the left and right channels. Therefore, a fourth type of parameter that describes the overall phase shift or overall phase difference (OPD) can be used additionally. In the stereo reconstruction process, consecutive windowed segments of both the received downmixed signal s[n] and the uncorrelated version of the received downmixed d[n] are processed together with the spatial parameters to generate the left (l k (n)) and right (rk (n) Reconstructed signal:

[0149] l k (n) = H 11 (k, n)s k (n) + H 21 (k, n)d k (n)

[0150] r k (n) = H 12 (k,n)s k (n) + H 22 (k, n)d k (n)

[0151] where H 11 , H 12 , H 21 and H 22 are defined by stereo parameters. Finally, the signals l k (n) and r k (n) are transformed back to the time domain through a frequency-to-time transformation.

[0152] ​ The control data generation subsystem 401 of ​ is coupled and configured to detect at least one property of an encoded audio bitstream to be decoded and to generate eSBR control data (which may be or include any type of eSBR metadata included in an encoded audio bitstream according to other embodiments of the present invention) in response to at least one result of the detection step. The eSBR control data is confirmed to stage 203 to trigger the application of individual eSBR tools or combinations of eSBR tools and / or to control the application of these eSBR tools after detecting a particular property (or combination of properties) of the bitstream. For example, to control the performance of eSBR processing using harmonic transposition, some embodiments of the control data generation subsystem 401 will include: a music detector (e.g., a simplified version of a conventional music detector) for setting the sbrPatchingMode[ch] parameter (and confirming the set parameter to stage 203) in response to detecting whether the bitstream indicates or does not indicate music; an instantaneous detector for setting the sbrOversamplingFlag[ch] parameter (and confirming the set parameter to stage 203) in response to detecting the presence or absence of an instantaneous in the audio content indicated by the bitstream; and / or a pitch detector for setting the sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch] parameters (and confirming the set parameters to stage 203) in response to detecting the pitch of the audio content indicated by the bitstream. Other aspects of the present invention are audio bitstream decoding methods performed by any embodiment of the inventive decoder described in this paragraph and the previous paragraphs.

[0153] Aspects of the present invention include types of encoding or decoding methods that any embodiment of the APU, system, or device of the present invention is configured (e.g., programmed) to perform. Other aspects of the present invention include systems or devices configured (e.g., programmed) to perform any embodiment of the methods of the present invention, and computer-readable media (e.g., disks) that (e.g., in a non-transitory manner) store program code for implementing any embodiment of the inventive methods or their steps. For example, the system of the present invention can be or include any programmable general-purpose processor, digital signal processor, or microprocessor that is programmed with software or firmware and / or otherwise configured to perform various operations on data (including embodiments of the methods of the present invention or their steps). Such a general-purpose processor can be or include a computer system that includes an input device, a memory, and processing circuitry programmed (and / or otherwise configured) to perform an embodiment of the method of the present invention (or its steps) in response to data being confirmed to it.

[0154] Embodiments of the present invention can be implemented in hardware, firmware, software, or a combination of both (e.g., as a programmable logic array). Unless otherwise indicated, algorithms or programs that are part of the present invention are not inherently related to any particular computer or other device. Specifically, various general-purpose machines can be used with programs written in accordance with the teachings herein, or it may be more convenient to construct more specialized devices (e.g., integrated circuits) to perform the required method steps. Thus, the present invention can be implemented in one or more computer programs executed on one or more programmable computer systems (e.g., ​ any element of, or ​ encoder 100 (or its elements) of, or ​ decoder 200 (or its elements) of, or ​ decoder 210 (or its elements) of, or ​ decoder 400 (or its elements) of

[0155] Each such program can be implemented in any desired computer language (including machine, assembly, or high-level programming, logic, or object-oriented programming languages) to communicate with the computer system. In any case, the language can be a compiled or interpreted language.

[0156] For example, when implemented by a sequence of computer software instructions, the various functions and steps of the embodiments of the present invention can be implemented by a multi-threaded software instruction sequence running in suitable digital signal processing hardware. In such a case, the various devices, steps, and functions of the embodiments can correspond to portions of the software instructions.

[0157] Each such computer program is preferably stored on or downloaded to a storage medium or device (e.g., solid-state memory or medium, or magnetic or optical medium) readable by a general-purpose or special-purpose programmable computer for configuring and operating the computer to execute the programs described herein when the storage medium or device is read by a computer system. The system of the present invention can also be implemented as a computer-readable storage medium configured to have (e.g., store) a computer program, where the storage medium so configured causes the computer system to operate in a specific and predefined manner to perform the functions described herein.

[0158] Several embodiments of the present invention have been described. However, it will be understood that various modifications can be made without departing from the spirit and scope of the present invention. Many modifications and variations of the present invention are possible in light of the above teachings. For example, to facilitate an efficient implementation, phase shift can be used in combination with a composite QMF analysis and synthesis filter bank. The analysis filter bank is responsible for filtering the time-domain low-frequency band signal generated by the core decoder into multiple sub-bands (e.g., QMF sub-bands). The synthesis filter bank is responsible for combining the regenerated high-frequency band generated by a selected HFR technique (such as by the received sbrPatchingMode parameter) with the decoded low-frequency band to produce a wide-band output audio signal. However, a given filter bank implementation operating in a particular sampling rate mode (e.g., normal dual-rate operation or reduced sampling rate SBR mode) should not have bit-stream-dependent phase shift. The QMF bank used in SBR is a theoretical complex indicator extension of a cosine-modulated filter bank. It can be shown that when using a complex exponential modulation extension of a cosine-modulated filter bank, the aliasing cancellation constraint becomes obsolete. Therefore, for the SBR QMF bank, both the analysis filter h k (n) and the synthesis filter f k (n) can be defined by the following:

[0159]

[0160] where p0(n) is a real-valued symmetric or asymmetric prototype filter (generally, a low-pass prototype filter), M represents the number of channels and N is the order of the prototype filter. The number of channels used in the analysis filter bank can be different from the number of channels used in the synthesis filter bank. For example, the analysis filter bank can have 32 channels, and the synthesis filter bank can have 64 channels. When operating the synthesis filter bank in the decimation mode, the synthesis filter bank can have only 32 channels. Since the subband samples from the filter bank are complex-valued, an additional possible channel-dependent phase shift step can be added to the analysis filter bank. These additional phase shifts need to be compensated for before the synthesis filter bank. Although in principle, the phase shift terms can have arbitrary values without impairing the operation of the QMF analysis / synthesis chain, they can also be constrained to certain values for consistency verification. The SBR signal will be affected by the choice of the phase factor while the low-pass signal from the core decoder will not. The audio quality of the output signal will not be affected.

[0161] The coefficients of the prototype filter p0(n) can be defined to have a length L of 640, as shown in Table 4 below.

[0162] Table 4

[0163]

[0164]

[0165]

[0166]

[0167]

[0168]

[0169] The prototype filter p0(n) can also be derived from Table 4 through one or more mathematical operations (such as rounding, subsampling, interpolation, and decimation).

[0170] Although the tuning of SBR-related control information generally does not depend on the details of transposition (as previously discussed), in some embodiments, certain elements of control data may be multiplexed in the eSBR extension container (bs_extension_id == EXTENSION_ID_ESBR) to improve the quality of the regenerated signal. Some of the multiplexed elements may include noise floor data (e.g., noise floor scale factors and an indication of the direction of differential coding of each noise floor (in the frequency direction or the time direction)), inverse filtering data (e.g., a parameter indicating an inverse filtering mode selected from no inverse filtering, low-level inverse filtering, medium-level inverse filtering, and strong-level inverse filtering), and missing harmonic data (e.g., a parameter indicating whether a sine curve should be added to a specific band of the regenerated high-frequency band). All of these elements rely on the synthesis analog of the transposer of the decoder performed in the encoder and thus can increase the quality of the regenerated signal if properly tuned for the selected transposer.

[0171] Specifically, in some embodiments, the missing harmonic and inverse filtering control data are transmitted in the eSBR extension container (along with other bitstream parameters of Table 3) and the harmonic transposer for eSBR is tuned. The additional bitrate required to transmit these two categories of metadata for the harmonic transposer for eSBR is relatively low. Thus, transmitting the tuned missing harmonic and / or inverse filtering control data in the eSBR extension container will increase the quality of the audio produced by the transposer while only minimally affecting the bitrate. To ensure backward compatibility with legacy decoders, parameters tuned for the spectral translation operation for SBR may also be sent in the bitstream as part of the SBR control data, either implicitly or explicitly signaled.

[0172] It should be understood that within the scope of the appended claims, the invention may be practiced in ways other than those specifically described herein. Any element symbols contained in the following claims are for illustrative purposes only and should not be used to interpret or limit the claims in any way. From the following enumerated exemplary embodiments (EEEs), various aspects of the invention will be understood:

[0173] EEE1. A method for performing high-frequency reconstruction of an audio signal, the method comprising:

[0174] Receiving an encoded audio bitstream that includes audio data representing a low-frequency band portion of the audio signal and high-frequency reconstruction metadata;

[0175] Decoding the audio data to produce a decoded low-frequency band audio signal;

[0176] Extract the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata including operation parameters for a high-frequency reconstruction process, the operation parameters including a patching mode parameter located in an extension container of the encoded audio bitstream, wherein a first value of the patching mode parameter indicates spectral translation, and a second value of the patching mode parameter indicates harmonic transposition by phase vocoder frequency expansion;

[0177] Filter the decoded low-band audio signal to produce a filtered low-band audio signal;

[0178] Regenerate a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein if the patching mode parameter is the first value, then the regeneration includes spectral translation, and if the patching mode parameter is the second value, then the regeneration includes harmonic transposition by phase vocoder frequency expansion; and

[0179] Combine the filtered low-band audio signal with the regenerated high-band portion to form a wide-band audio signal.

[0180] EEE2. The method according to EEE 1, wherein the extension container includes inverse filtering control data used when the patching mode parameter is equal to the second value.

[0181] EEE3. The method according to any one of EEE 1 to 2, wherein the extension container further includes missing harmonic control data used when the patching mode parameter is equal to the second value.

[0182] EEE4. The method according to any of the foregoing EEE, wherein the encoded audio bitstream further includes a padding element having an identifier indicating the start of the padding element and padding data after the identifier, wherein the padding data includes the extension container.

[0183] EEE5. The method according to EEE 4, wherein the identifier is an unsigned integer that transmits the most significant bit first in a three-bit format and has a value of 0x6.

[0184] EEE6. The method according to EEE 4 or EEE 5, wherein the padding data includes an extended payload, the extended payload including spectral band replication extension data, and the extended payload is identified as an unsigned integer that transmits the most significant bit first in a four-bit format and has a value of '1101' or '1110', and optionally,

[0185] wherein the spectral band replication extension data includes:

[0186] Select a spectral band replication header,

[0187] Spectral band replication data, which is after the header, and

[0188] A spectral band replication extension element, which is after the spectral band replication data, and wherein a flag is included in the spectral band replication extension element.

[0189] EEE7. The method according to any one of EEE 1 to 6, wherein the high-frequency reconstruction metadata includes an envelope scale factor, a noise floor scale factor, time / frequency grid information, or a parameter indicating a crossover frequency.

[0190] EEE8. The method according to any one of EEE 1 to 7, wherein the filtering is performed by an analysis filter bank including an analysis filter h k (n) which is a modulated version of the prototype filter p0(n) according to the following:

[0191]

[0192] where p0(n) is a real-valued symmetric or asymmetric prototype filter, M is the number of channels in the analysis filter bank, and N is the order of the prototype filter.

[0193] EEE9. The method according to EEE 8, wherein the prototype filter p0(n) is derived from the coefficients in Table 4 herein.

[0194] EEE10. The method according to EEE 8, wherein the prototype filter p0(n) is derived from the coefficients in Table 4 herein by one or more mathematical operations selected from the group consisting of rounding, subsampling, interpolation, or decimation.

[0195] EEE11. The method according to any one of EEE 1 to 10, wherein a phase shift is added to the filtered low-frequency band audio signal after the filtering, and the phase shift is compensated before the combination to reduce the complexity of the method.

[0196] EEE12. The method according to any of the foregoing EEEs, wherein the extension container further includes a flag indicating whether additional preprocessing is used to avoid a discontinuity in the shape of the spectral envelope of the high-frequency band portion when the patching mode parameter is equal to the first value, wherein the first value of the flag enables the additional preprocessing and the second value of the flag disables the additional preprocessing.

[0197] EEE13. The method according to EEE 12, wherein the additional preprocessing includes calculating a pre-gain curve using linear prediction filter coefficients.

[0198] EEE14. A method according to any one of EEE 1 to 13, wherein the extended container is a backward compatible extended container.

[0199] EEE15. A method according to any one of EEE 1 to 14, wherein the encoded audio stream is encoded according to a format, and wherein the extended container is an extended container defined in at least one legacy version of the format.

[0200] EEE16. A non - transitory computer - readable medium containing instructions that, when executed by a processor, perform a method according to any one of EEE1 to 15.

[0201] EEE17. An audio processing unit for performing high - frequency reconstruction of an audio signal, the audio processing unit being configured to perform a method according to any one of EEE 1 to 15.

Claims

1. A method for performing high-frequency reconstruction of an audio signal, the method comprising: Receiving an encoded audio bitstream that includes audio data representing a low-frequency band portion of the audio signal and high-frequency reconstruction metadata; Decoding the audio data to produce a decoded low-frequency band audio signal; Extracting the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata including operation parameters for a high-frequency reconstruction process, the operation parameters including a patching mode parameter located in a backward-compatible extension container of the encoded audio bitstream, wherein a first value of the patching mode parameter indicates spectral translation, and a second value of the patching mode parameter indicates harmonic transposition by phase vocoder frequency expansion; Filtering the decoded low-frequency band audio signal to produce a filtered low-frequency band audio signal; Regenerating a high-frequency band portion of the audio signal using the filtered low-frequency band audio signal and the high-frequency reconstruction metadata, wherein if the patching mode parameter is the first value, then regenerating the high-frequency band portion of the audio signal includes spectral translation, and if the patching mode parameter is the second value, then regenerating the high-frequency band portion of the audio signal includes harmonic transposition by phase vocoder frequency expansion.

2. The method according to claim 1, wherein the backward-compatible extension container includes inverse filter control data used when the patching mode parameter is equal to the second value.

3. The method according to claim 1, wherein the backward-compatible extension container further includes missing harmonic control data used when the patching mode parameter is equal to the second value.

4. The method according to claim 1, wherein after the filtering, a phase shift is added to the filtered low-frequency band audio signal, and the phase shift is compensated to reduce the complexity of the method.

5. The method according to claim 1, wherein the backward-compatible extension container further includes a flag indicating whether additional preprocessing is used to avoid discontinuities in the shape of the spectral envelope of the high-frequency band portion when the patching mode parameter is equal to the first value, wherein a first value of the flag enables the additional preprocessing, and a second value of the flag disables the additional preprocessing.

6. The method according to claim 5, wherein the additional preprocessing includes using linear prediction filter coefficients to calculate a pre-gain curve.

7. A non-transitory computer-readable medium containing instructions that, when executed by a processor, perform the method according to claim 1.

8. An audio processing unit for performing high-frequency reconstruction of an audio signal, the audio processing unit comprising: An input interface for receiving an encoded audio bitstream that includes audio data representing a low-frequency band portion of the audio signal and high-frequency reconstruction metadata; A core audio decoder for decoding the audio data to produce a decoded low-frequency band audio signal; An inverse formatter for extracting the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata including operation parameters for a high-frequency reconstruction process, the operation parameters including a patching mode parameter located in a backward-compatible extension container of the encoded audio bitstream, wherein a first value of the patching mode parameter indicates spectral translation and a second value of the patching mode parameter indicates harmonic transposition by phase vocoder frequency expansion; An analysis filter bank for filtering the decoded low-band audio signal to produce a filtered low-band audio signal; A high-frequency regenerator for reconstructing a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein if the patching mode parameter is the first value, the reconstruction includes spectral translation, and if the patching mode parameter is the second value, the reconstruction includes harmonic transposition by phase vocoder frequency expansion.

Citation Information

Patent Citations

  • Backtracking compatible integration of high frequency reconstruction techniques for audio signals

    CN113936671A

  • Backtracking compatible integration of high frequency reconstruction techniques for audio signals

    CN113936672A

  • Backtracking compatible integration of high frequency reconstruction techniques for audio signals

    CN113936674A

  • Backward-compatible integration of high frequency reconstruction techniques for audio signals

    CN113990331A