Integration of high frequency reconstruction techniques with reduced post-processing delay

By integrating spectral shifting and harmonic transposition techniques, the patent addresses the challenge of high-frequency reconstruction in audio encoding, enhancing audio quality and reducing delays in MPEG-4 AAC systems.

HK40135067APending Publication Date: 2026-07-17DOLBY INTERNATIONAL AB

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Applications
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2026-06-04
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing audio encoding technologies, such as MPEG-4 AAC, face challenges in effectively reconstructing high-frequency components of audio signals, particularly for music content with low cross frequencies, as spectral band replication techniques like SBR may not be suitable, leading to suboptimal audio quality at low data rates.

Method used

The integration of high-frequency reconstruction techniques, including spectral shifting and harmonic transposition, is achieved through advanced metadata processing and filter banks to regenerate high-frequency bands, maintaining tone and noise-like component ratios, with a delay of 3010 samples per audio channel.

Benefits of technology

This approach enhances audio quality by accurately reconstructing high-frequency components, maintaining spectral characteristics, and reducing post-processing delays, thereby improving audio fidelity without significant computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to integration of high frequency reconstruction techniques with reduced post-processing delay, and specifically discloses a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream and decoding audio data to produce a decoded low-band audio signal. The method further includes extracting high frequency reconstruction metadata and filtering the decoded low band audio signal using an analysis filter bank to produce a filtered low band audio signal. The method also includes extracting a flag indicating whether to perform spectral translation or harmonic transpose on the audio data and regenerating a high frequency band portion of the audio signal using the filtered low frequency band audio signal and the high frequency reconstruction metadata in accordance with the flag. The high frequency regeneration is performed as a delayed post-processing operation with 3010 samples per audio channel.
Need to check novelty before this filing date? Find Prior Art

Description

(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202511677677.3 (22) Application Date 2019.04.25 (30) Priority Data 62 / 662,296 2018.04.25 US (62) Divisional Application Data 201980034811.4 2019.04.25 (71) Applicant Dolby International Address Amsterdam, Netherlands (72) Inventors K. Kjoerlin L. Wilmons H. Purnagen P. Ekstrand (74) Patent Agency Beijing Law Alliance Intellectual Property Agency Co., Ltd. 11287 Patent Attorney Liu Feng (51) Int.Cl. G10L 19 / 18 (2013.01) G10L 19 / 16 (2013.01) G10L 21 / 038 (2013.01) (54) Invention Title: Integration of High-Frequency Reconstruction Techniques with Reduced Post-Processing Delay (57) Abstract This application relates to the integration of high-frequency reconstruction techniques with reduced post-processing delay, and specifically discloses a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream and decoding audio data to generate a decoded low-frequency band audio signal. The method further includes extracting high-frequency reconstruction metadata and using an analytical filter bank to filter the decoded low-frequency band audio signal to generate a filtered low-frequency band audio signal. The method also includes extracting a flag indicating whether to perform a spectral shift or harmonic transpose on the audio data and using the filtered low-frequency band audio signal and the high-frequency reconstruction metadata to regenerate the high-frequency band portion of the audio signal according to the flag. The high-frequency regeneration is performed with a post-processing operation having a delay of 3010 samples per audio channel. Claims 1 page, Description 31 pages, Drawings 4 pages, CN 121214953 A 2025.12.26 CN 1 21 21 49 53 A 1. A method for performing high-frequency reconstruction of an audio signal, the method comprising: receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-frequency band portion of the audio signal and high-frequency reconstruction metadata, wherein the high-frequency reconstruction metadata includes a noise scaling factor; decoding the audio data to generate a decoded low-frequency band audio signal; extracting the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata including operating parameters for the high-frequency reconstruction process, the operating parameters including a patch mode parameter located in a backward-compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates a spectral shift and a second value of the patch mode parameter indicates a harmonic transpose by frequency stretching via a phase vocoder;The decoded low-frequency audio signal is filtered to generate a filtered low-frequency audio signal; and the filtered low-frequency audio signal and the high-frequency reconstruction metadata are used to regenerate the high-frequency portion of the audio signal, wherein if the patching mode parameter is the first value, the regeneration includes spectral shifting, and if the patching mode parameter is the second value, the regeneration includes harmonic transposition by frequency stretching via a phase vocoder, wherein the filtering and regeneration are performed as post-processing operations with a delay of 3010 samples per audio channel, and wherein the spectral shifting includes maintaining the ratio between the tone component and the noise-like component by adaptive inverse filtering. 2. The method of claim 1, wherein the harmonic transposition by frequency stretching via a phase vocoder is performed with an estimation complexity equal to or less than 4.5 million operations per second and equal to or less than 3,000 words of memory. 3. A non-transitory computer-readable medium containing instructions that, when executed by a processor, perform the method of claim 1. 4. A computer program product stored on a non-transitory computer-readable medium having instructions, when executed by a computing device or system, to cause the computing device or system to perform the method of claim 1. 5. An audio processing unit for performing high-frequency reconstruction of an audio signal, the audio processing unit comprising: an input interface for receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-frequency band portion of the audio signal and high-frequency reconstruction metadata, wherein the high-frequency reconstruction metadata includes a noise scaling factor; a core audio decoder for decoding the audio data to generate a decoded low-frequency band audio signal; and a deformatter for extracting the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata including operating parameters for the high-frequency reconstruction process, the operating parameters including a patch mode parameter located in a backward-compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates a spectral shift and a second value of the patch mode parameter indicates a harmonic transpose through frequency stretching by a phase vocoder; An analysis filter bank for filtering the decoded low-frequency band audio signal to produce a filtered low-frequency band audio signal; and a high-frequency regenerator for reconstructing the high-frequency band portion of the audio signal using the filtered low-frequency band audio signal and the high-frequency reconstruction metadata, wherein the reconstruction includes spectral shifting if the patching mode parameter is the first value, and the reconstruction includes harmonic transposition via frequency extension by a phase vocoder if the patching mode parameter is the second value, wherein the analysis filter bank and the high-frequency regenerator are executed in a post-processor with a delay of 3010 samples per audio channel, and wherein the spectral shifting includes maintaining the ratio between the tone component and the noise-like component through adaptive inverse filtering. Claim 1 / 1Page 2 CN 121214953 A Integration of High-Frequency Reconstruction Techniques with Reduced Post-Processing Delay

[0001] Information about the Divisional Application

[0002] This application is a divisional application. The parent application of this divisional application is the invention patent application filed on April 25, 2019, with application number 202111584446.X, entitled "Integration of High-Frequency Reconstruction Techniques with Reduced Post-Processing Delay"; the parent application is the invention patent application filed on April 25, 2019, with application number 201980034811.4, entitled "Integration of High-Frequency Reconstruction Techniques with Reduced Post-Processing Delay".

[0003] Cross-Reference to Related Applications

[0004] This application claims priority to U.S. Provisional Patent Application No. 62 / 662,296, filed on April 25, 2018, the entire contents of which are incorporated herein by reference. Technical Field

[0005] Embodiments relate to audio signal processing, and more specifically, embodiments relate to encoding, decoding, or transcoding audio bitstreams using control data that specifies a basic form or an enhanced form of HFR to perform high-frequency reconstruction (“HFR”) on audio data. Background Art

[0006] A typical audio bitstream contains both audio data (e.g., encoded audio data) indicating one or more channels of audio content and metadata indicating at least one characteristic of the audio data or audio content. A well-known format for generating encoded audio bitstreams is the MPEG-4 Advanced Audio Coding (AAC) format described in the MPEG standard ISO / IEC 14496-3:2009. In the MPEG-4 standard, AAC stands for “Advanced Audio Coding” and HE-AAC stands for “High-Efficiency Advanced Audio Coding”.

[0007] The MPEG-4 AAC standard defines several audio profiles that determine which objects and encoding tools are present in a compatible encoder or decoder. Three of these audio profiles are (1) the AAC profile, (2) the HE-AAC profile, and (3) the HE-AAC v2 profile. AAC profiles contain AAC Low Complexity (or "AAC-LC") object types. AAC-LC objects are the counterpart to the MPEG-2 AAC Low Complexity profile, with some modifications, and do not include either the Spectral Band Replication ("SBR") object type or the Parametric Stereo ("PS") object type. HE-AAC profiles are a superset of AAC profiles and additionally contain the SBR object type. HE-AAC v2 profiles are a superset of HE-AAC profiles and additionally contain the PS object type.

[0008] The SBR object type contains a Spectral Band Replication tool, which significantly improves the compression efficiency of perceptual audio codecs.SBR is an important high-frequency reconstruction (“HFR”) encoding tool. SBR reconstructs the high-frequency components of the audio signal on the receiver side (e.g., in the decoder). Therefore, the encoder only needs to encode and transmit the low-frequency components to allow for much higher audio quality at low data rates. SBR replicates a harmonic sequence previously truncated to reduce the data rate based on the limited available bandwidth signal obtained from the encoder and control data. The ratio between the tone components and noise-like components is maintained through adaptive inverse filtering and optionally the addition of noise and sine curves. In the MPEG-4 AAC standard, the SBR tool performs spectral patching (also known as linear shifting or spectrum shifting), where several consecutive quadrature mirror filter (QMF) subbands are copied (or “patched”) from the transmitted low-frequency portion of the audio signal to the high-frequency portion of the audio signal (which is generated in the decoder).

[0009] Spectral patching or linear shifting may not be suitable for certain audio types (e.g., music content with relatively low cross frequencies). Therefore, techniques for improving spectral band replication are needed. Specification 1 / 31 Page 3 CN 121214953 A Summary of the Invention

[0010] A first type of embodiment relates to a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream and decoding the audio data to generate a decoded low-frequency band audio signal. The method further includes extracting high-frequency reconstruction metadata and using an analysis filter bank to filter the decoded low-frequency band audio signal to generate a filtered low-frequency band audio signal. The method further includes extracting a marker indicating whether to perform a spectral shift or harmonic transpose on the audio data and using the filtered low-frequency band audio signal and the high-frequency reconstruction metadata to regenerate a high-frequency band portion of the audio signal according to the marker. Finally, the method includes combining the filtered low-frequency band audio signal and the regenerated high-frequency band portion to form a broadband audio signal.

[0011] A second type of embodiment relates to an audio decoder for decoding an encoded audio bitstream. The decoder includes: an input interface for receiving the encoded audio bitstream, wherein the encoded audio bitstream contains audio data representing the low-frequency band portion of an audio signal; and a core decoder for decoding the audio data to generate a decoded low-frequency band audio signal. The decoder also includes: a demultiplexer for extracting high-frequency reconstruction metadata from the encoded audio bitstream, wherein the high-frequency reconstruction metadata contains operating parameters for a high-frequency reconstruction process that linearly shifts several consecutive sub-bands from the low-frequency band portion of the audio signal to the high-frequency band portion of the audio signal; and an analysis filter bank for filtering the decoded low-frequency band audio signal to generate a filtered low-frequency band audio signal. The decoder further includes: a demultiplexer for extracting an indication from the encoded audio bitstream that is related to the audio...The data is marked as either linearly shifted or harmonic transposed; and a high-frequency regenerator is used to regenerate the high-frequency band portion of the audio signal using the filtered low-frequency band audio signal and the high-frequency reconstruction metadata according to the marked value. Finally, the decoder includes a synthesis filter bank for combining the filtered low-frequency band audio signal and the regenerated high-frequency band portion to form a broadband audio signal.

[0012] Other class of embodiments relate to encoding and transcoding audio bitstreams containing metadata that identifies whether enhanced spectrum band replication (eSBR) processing is performed. Brief Description of the Drawings

[0013] FIG1 is a block diagram of an embodiment of a system that can be configured to perform an embodiment of the inventive method.

[0014] FIG2 is a block diagram of an encoder, which is an embodiment of the inventive audio processing unit.

[0015] FIG3 is a block diagram of a system that includes a decoder (which is an embodiment of the inventive audio processing unit) and also optionally includes a post-processor coupled to the decoder.

[0016] FIG4 is a block diagram of a decoder, which is an embodiment of the inventive audio processing unit.

[0017] FIG5 is a block diagram of a decoder, which is another embodiment of the audio processing unit of the invention.

[0018] FIG6 is a block diagram of another embodiment of the audio processing unit of the invention.

[0019] FIG7 is a block diagram of an MPEG-4 AAC bitstream, including its division into several segments.

[0020] Symbols and Terms

[0021] In this invention (included in the claims), the expression “to” a signal or data to perform an operation (e.g., filtering, scaling, transforming, or applying gain to a signal or data) is used broadly to mean performing an operation directly on the signal or data or on a processed version of the signal or data (e.g., on a version of a signal that has undergone preliminary filtering or preprocessing before the operation is performed on it).

[0022] In this invention (included in the claims), the expression “audio processing unit” or “audio processor” is used broadly to mean a system, apparatus, or device configured to process audio data. Examples of audio processing units include (but are not limited to) encoders, transcoders, decoders, codecs, preprocessing systems, post-processing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools). Almost all consumer electronics products (e.g., mobile phones, televisions, laptops, and tablets) contain audio processing units or audio processors.

[0023] In this invention (included in the claims), the term "coupled" or "via coupling" is used broadly to mean a direct or indirect connection. Thus, if a first device is coupled to a second device, the connection can be either a direct connection or an indirect connection via other devices and connections. Furthermore, components integrated into or with other components are also coupled to each other. Detailed Description

[0024] MPEG-4The AAC standard anticipates that encoded MPEG-4 AAC bitstreams contain metadata indicating each type of High Frequency Reconstruction (“HFR”) processing applied (if to be applied) by the decoder to decode the audio content of the bitstream, and / or controlling this HFR processing, and / or indicating at least one characteristic or parameter of at least one HFR tool used to decode the audio content of the bitstream. In this document, we use the term “SBR metadata” to refer to this type of metadata used in conjunction with Spectral Band Replication (“SBR”), as described or mentioned in the MPEG-4 AAC standard. Those skilled in the art will understand that SBR is a form of HFR.

[0025] SBR is preferably used as a dual-rate system, where the basic codec operates at half the original sampling rate, while the SBR operates at the original sampling rate. Despite the higher sampling rate, the SBR encoder operates in parallel with the basic core codec. Although SBR is primarily post-processing in the decoder, important parameters are extracted from the encoder to ensure the most accurate high-frequency reconstruction in the decoder. The encoder estimates the spectral envelope of the SBR range suitable for the time and frequency range / resolution of the current input signal segment characteristics. The spectral envelope is estimated through complex QMF analysis and subsequent energy calculation. The time and frequency resolution of the spectral envelope can be chosen with great freedom to ensure the most suitable time and frequency resolution for a given input region. Envelope estimation needs to take into account that transients of the original source, which are mainly located in the high-frequency region (e.g., the high cap), will appear to a lesser extent in the high-frequency band generated by the SBR before envelope adjustment, because the high-frequency band in the decoder is based on the low-frequency band in which transients are much less pronounced than the high-frequency band. This aspect places different requirements on the time and frequency resolution of the spectral envelope data compared to general spectral envelope estimation used in other audio coding algorithms.

[0026] In addition to the spectral envelope, several additional parameters representing the spectral characteristics of the input signal in different time and frequency regions are also extracted. Since the encoder naturally has access to the original signal and information about how the SBR unit in the decoder will generate the high-frequency band, the system can handle, given a specific set of control parameters, cases where the low-frequency band constitutes a strong harmonic series and the regenerated high-frequency band mainly constitutes random signal components, and cases where strong tone components exist in the original high-frequency band but have no counterpart in the low-frequency band (the high-frequency band region is based on this). Furthermore, the SBR encoder works closely in conjunction with the basic core codec to evaluate which frequency range should be covered by the SBR at a given time. For stereo signals, the SBR data is efficiently encoded prior to transmission by utilizing entropy coding and the channel dependence of the control data.

[0027] Typically, the basic codec is required to carefully tune the control parameter extraction algorithm for a given bit rate and sampling rate. This is due to the fact that lower bit rates generally imply a larger SBR range than higher bit rates, and different sampling rates correspond to different temporal resolutions of the SBR frames.

[0028] An SBR decoder typically comprises several distinct parts. These include a bitstream decoding module, a high-frequency reconstruction (HFR) module, an additional high-frequency component module, and an envelope adjuster module. The system is based on either a complex-valued QMF filter bank (for high-quality SBR) or a real-valued QMF filter bank (for low-power SBR). Embodiments of this invention are applicable to both high-quality and low-power SBR. In the bitstream extraction module, control data is read from and decoded from the bitstream. Before reading the envelope data from the bitstream, the time-frequency grid of the current frame is obtained (see page 3 / 31 of the specification, CN 121214953 A). The basic core decoder decodes the audio signal of the current frame (albeit at a lower sampling rate) to generate time-domain audio samples. The resulting frame is used by the HFR module to perform high-frequency reconstruction using the audio data. Next, a QMF filter bank is used to analyze the decoded low-frequency band signal. Subsequently, high-frequency reconstruction and envelope adjustment are performed on the sub-band samples of the QMF filter bank. High frequencies are reconstructed from the low-frequency band in a flexible manner based on given control parameters. Furthermore, based on control data, adaptive filtering of sub-band channels reconstructs high-frequency bands to ensure appropriate spectral characteristics for a given time / frequency region.

[0029] The top layer of the MPEG-4 AAC bitstream is a sequence of data blocks (“raw_data_block” elements), each of which is a data segment (referred to herein as a “block”) containing audio data (typically within a time period of 1024 or 960 samples) and related and / or other data. In this document, we use the term “block” to refer to a segment of the MPEG-4 AAC bitstream that includes audio data (and corresponding metadata and optionally other related data), which identifies or indicates one (but not more than one) “raw_data_block” element.

[0030] Each block of the MPEG-4 AAC bitstream may contain several syntax elements (each of which is also embodied as a data segment in the bitstream). Seven types of these syntax elements are defined in the MPEG-4 AAC standard. Each syntax element is identified by a distinct value of the data element “id_syn_ele”. Instances of syntax elements include “single_channel_element()”, “channel_pair_element()”, and “fill_element()”. A single-channel element is a container for audio data containing a single audio channel (mono audio signal). A channel-pair element contains audio data for two audio channels (i.e., stereo audio signals).

[0031] A fill element is an information container containing an identifier (e.g., the value of the element “id_syn_ele” mentioned above) and subsequent data (which is referred to as “fill data”). Fill elements have historically been used to adjust the instantaneous bit rate of a bitstream that will be transmitted through a constant-rate channel. A constant data rate can be achieved by adding an appropriate amount of fill data to each block.

[0032] According to embodiments of the invention, the padding data may include one or more extended payloads that extend the data types (e.g., metadata) that can be transmitted in the bitstream. A decoder receiving a bitstream with padding data containing new data types may optionally be used by the receiving bitstream device (e.g., a decoder) to extend the functionality of the device. Therefore, those skilled in the art will understand that padding elements are a special type of data structure and differ from data structures typically used for transmitting audio data (e.g., audio payloads containing channel data).

[0033] In some embodiments of the invention, the identifier used to identify padding elements may consist of a 3-bit unsigned integer (“uimsbf”) with a value of 0×6, transmitting the most significant bit first. Within a block, several instances of the same type of syntax elements (e.g., several padding elements) may appear.

[0034] Another standard for encoding audio bitstreams is the MPEG Unified Speech and Audio Coding (USAC) standard (ISO / IEC 23003-3:2012). The MPEG USAC standard describes the use of spectrum band copy processing (including the SBR processing described in the MPEG-4 AAC standard and other enhancements to spectrum band copy processing) to encode and decode audio content. This processing applies an extended and enhanced version of the SBR toolset described in the MPEG-4 AAC standard (sometimes referred to herein as “enhanced SBR tools” or “eSBR tools”). Therefore, eSBR (as defined in the USAC standard) is an improvement upon SBR (as defined in the MPEG-4 AAC standard).

[0035] In this document, we use the term “enhanced SBR processing” (or “eSBR processing”) to refer to a spectral band copying process using at least one eSBR tool not described or mentioned in the MPEG-4 AAC standard (e.g., at least one eSBR tool described or mentioned in the MPEG USAC standard). Examples of these eSBR tools are harmonic transpose and QMF patching additional preprocessing or “pre-flattening”.

[0036] An integer-order T harmonic transpose maps a sine curve with frequency ω to a sine curve with frequency Tω, while preserving the signal duration. Typically, three orders T=2, 3, 4 are used sequentially to generate each portion of the desired output frequency range using the minimum possible transpose order (page 4 / 31, CN 121214953 A). If an output range higher than the 4th order transpose is required, it can be generated by frequency shifting. The fundamental frequency time domain, as close to the critical sampling as possible, is generated for processing to minimize computational complexity.

[0037] The harmonic transposer can be based on QMF or DFT. When using a QMF-based harmonic transposer, a modified phase vocoder structure is used in the QMF domain to fully implement the bandwidth extension of the core encoder time-domain signal for each QMF subband.Sampling and subsequent time extension. Transposition using several transposition factors (e.g., T = 2, 3, 4) is implemented at the common QMF analysis / synthesis transform stage. Since the QMF-based harmonic transposer does not feature signal-adaptive frequency-domain oversampling, the corresponding flag (sbrOversamplingFlag[ch]) in the bitstream can be ignored.

[0038] When using a DFT-based harmonic transposer, factor 3 and 4 transposers (3rd and 4th order transposers) are preferably integrated into the factor 2 transposer (2nd order converter) by interpolation to reduce complexity. For each frame (corresponding to coreCoderFrameLength core encoder samples), the nominal "full-size" transform size of the transposer is first determined by the signal-adaptive frequency-domain oversampling flag (sbrOversamplingFlag[ch]) in the bitstream.

[0039] When sbrPatchingMode == 1 to indicate that linear transposition will be used to generate the high-frequency band, an additional step can be introduced to avoid shape discontinuities in the spectral envelope of the high-frequency signal being input into the subsequent envelope adjuster. This improves the operation of the subsequent envelope adjustment stage to result in a high-frequency band signal that is perceived as more stable. The operation of the additional preprocessing benefits signal types where the rough spectral envelope of the low-frequency band signal used for high-frequency reconstruction exhibits a large level of variation. However, the value of the bitstream element can be determined in the encoder by applying any kind of signal-dependent classification. Preferably, the additional preprocessing is initiated by a 1-bit bitstream element bs_sbr_preprocessing. When bs_sbr_preprocessing is set to 1, the additional processing is enabled. When bs_sbr_preprocessing is set to 0, the additional preprocessing is disabled. The additional processing preferably uses the pre-gain curve used by the high-frequency generator to scale each patched low-frequency band XLow. For example, the pre-gain curve can be calculated according to the following equation:

[0040] , 0 ≤ k < k0

[0041] where k0 is the first QMF sub-band in the main frequency band table and lowEnvSlope is calculated using a function (e.g., polyfit()) that calculates the coefficients of the best-fit polynomial (in the least squares sense). For example, (using a cubic polynomial) can be employed

[0042] ;

[0043] and where

[0044] , 0 ≤ k < k0

[0045] where x_lowband(k) = [0...k0 - 1], numTimeSlot is the number of SBR envelope time slots present in the frame, RATE is a constant indicating the number of QMF sub-band samples per time slot (e.g., 2), φk is the linear prediction filter coefficient (obtainable from the covariance method) and where

[0046] .

[0047] According to MPEGThe bitstream generated by the USAC standard (sometimes referred to herein as “USAC bitstream”) contains encoded audio content and typically contains metadata indicating each type of spectral band copying process applied by the decoder to decode the audio content of the USAC bitstream and / or metadata controlling this spectral band copying process and / or indicating at least one feature or parameter of at least one SBR tool and / or eSBR tool used to decode the audio content of the USAC bitstream.

[0048] In this document, we use the term “enhanced SBR metadata” (or “eSBR metadata”) to refer to metadata, which refers to each type of spectral band copying process applied by the decoder to decode the audio content of the encoded audio bitstream (e.g., USAC bitstream), and / or controlling this spectral band copying process, and / or indicating at least one feature or parameter of at least one SBR tool and / or eSBR tool used to decode this audio content but not described or mentioned in the MPEG-4 AAC standard. An instance of eSBR metadata is metadata (indicating or used to control spectrum band copying processing) described or mentioned in the MPEG USAC standard but not described or mentioned in the MPEG-4 AAC standard. Therefore, eSBR metadata in this document refers to metadata that is not SBR metadata, and SBR metadata in this document refers to metadata that is not eSBR metadata.

[0049] The USAC bitstream may contain both SBR metadata and eSBR metadata. More specifically, the USAC bitstream may contain eSBR metadata that controls eSBR processing performed by the decoder and SBR metadata that controls SBR processing performed by the decoder. According to a typical embodiment of the invention, eSBR metadata (e.g., eSBR-specific configuration data) is contained (according to the invention) in the MPEG-4 AAC bitstream (e.g., in the sbr_extension() container at the end of the SBR payload).

[0050] During decoding of the encoded bitstream using the eSBR toolset (including at least one eSBR tool), the decoder performs eSBR processing to regenerate the high-frequency band of the audio signal based on the copying of harmonic sequences truncated during encoding. This eSBR processing typically adjusts the spectral envelope of the generated high-frequency band and applies inverse filtering, and adds noise and sinusoidal components to regenerate the spectral characteristics of the original audio signal.

[0051] According to a typical embodiment of the invention, eSBR metadata (e.g., a small number of control bits of eSBR metadata) is contained in one or more metadata segments of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream), which also contains encoded audio data in other segments (audio data segments). Typically, at least one of these metadata bits is present in each block of the bitstream.The data segment is (or contains) a padding element (containing an identifier indicating the start of the padding element), and the eSBR metadata is contained in the padding element following the identifier.

[0052] Figure 1 is a block diagram of an exemplary audio processing chain (audio data processing system) in which one or more elements of the system can be configured according to embodiments of the invention. The system includes the following elements coupled together as shown: encoder 1, transmission subsystem 2, decoder 3, and post-processing unit 4. In variations of the system shown, one or more elements are omitted, or additional audio data processing units are included.

[0053] In some embodiments, encoder 1 (which optionally includes a preprocessing unit) is configured to accept PCM (temporal domain) samples including audio content as input and output an encoded audio bitstream (with a format conforming to the MPEG-4 AAC standard) indicating the audio content. The data of the bitstream indicating the audio content is sometimes referred to herein as “audio data” or “encoded audio data”. If the encoder is configured according to a typical embodiment of the invention, the audio bitstream output from the encoder contains eSBR metadata (and typically also other metadata) as well as audio data.

[0054] One or more encoded audio bitstreams output from encoder 1 can be asserted to encoded audio transmission subsystem 2. Subsystem 2 is configured to store and / or transmit each encoded bitstream output from encoder 1. The encoded audio bitstreams output from encoder 1 may be stored by subsystem 2 (e.g., in the form of DVD or Blu-ray disc), or transmitted by subsystem 2 (which may implement a transmission link or network), or may be both stored and transmitted by subsystem 2.

[0055] Decoder 3 is configured to decode the encoded MPEG-4 AAC audio bitstream (generated by encoder 1) received via subsystem 2. In some embodiments, decoder 3 is configured to extract eSBR metadata from each block of the bitstream and decode the bitstream (including performing eSBR processing using the extracted eSBR metadata) to produce decoded audio data (e.g., a decoded PCM audio sample stream). In some embodiments, decoder 3 is configured to extract SBR metadata from the bitstream (but ignore eSBR metadata contained in the bitstream) and decode the bitstream (including performing SBR processing using the extracted SBR metadata) to produce decoded audio data (e.g., a stream of decoded PCM audio samples). Typically, decoder 3 includes a buffer that stores (e.g., in a non-transitory manner) segments of the encoded audio bitstream received from subsystem 2. Specification 6 / 31 pages 8 CN 121214953 A

[0056] The post-processing unit 4 of FIG1 is configured to receive the decoded audio data stream (e.g., decoded PCM audio samples) from decoder 3 and perform post-processing on it. The post-processing unit may also be configured to render the post-processed audio content (orThe decoded audio received from decoder 3 is played by one or more speakers.

[0057] FIG2 is a block diagram of encoder 100, which is an embodiment of the audio processing unit of the invention. Any component or element of encoder 100 may be implemented in hardware, software or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA or other integrated circuits). Encoder 100 includes encoder 105, filler / formatter stage 107, metadata generation stage 106 and buffer memory 109 connected as shown. Typically, encoder 100 also includes other processing elements (not shown). Encoder 100 is configured to convert the input audio bitstream into an encoded output MPEG-4 AAC bitstream.

[0058] Metadata generator 106 is coupled and configured to generate metadata (including eSBR metadata and SBR metadata) (and / or pass it to stage 107) to be included in the encoded bitstream output from encoder 100 by stage 107.

[0059] Encoder 105 is coupled and configured to encode input audio data (e.g., by performing compression on it) and asserts the resulting encoded audio to stage 107 for inclusion in the encoded bitstream output from stage 107.

[0060] Stage 107 is configured to multiplex the encoded audio from encoder 105 and metadata (including eSBR metadata and SBR metadata) from generator 106 to produce the encoded bitstream output from stage 107, preferably such that the encoded bitstream has a format specified by an embodiment of the invention.

[0061] Buffer memory 109 is configured to store (e.g., in a non-transitory manner) at least one block of the encoded audio bitstream output from stage 107, and then asserts the block sequence of the encoded audio bitstream from buffer memory 109 as the output from encoder 100 to the transmission system.

[0062] FIG3 is a block diagram of a system including a decoder 200 (which is an embodiment of the audio processing unit of the invention) and optionally also including a post-processor 300 coupled to the decoder 200. Any component or element of the decoder 200 and the post-processor 300 may be implemented in hardware, software or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA or other integrated circuits). The decoder 200 includes a buffer memory 201, a bitstream payload deformatter (parser) 205, an audio decoding subsystem 202 (sometimes referred to as the “core” decoding stage or “core” decoding subsystem), an eSBR processing stage 203, and a control bit generation stage 204 connected as shown. Typically, the decoder 200 also includes other processing elements (not shown).

[0063] The buffer memory (buffer) 201 stores (e.g., in a non-transitory manner) the encoded MPEG-4 received by the decoder 200.At least one block of an AAC audio bitstream. In the operation of decoder 200, the block sequence of the bitstream is asserted from buffer 201 to deformatter 205.

[0064] In a variant of the embodiment of FIG3 (or the embodiment of FIG4 to be described), the APU (which is not a decoder) (e.g., APU 500 of FIG6) includes a buffer memory (e.g., a buffer memory identical to buffer 201) that stores (e.g., in a non-transitory manner) at least one block of an encoded audio bitstream of the same type (e.g., an MPEG-4 AAC audio bitstream) received by buffer 201 of FIG3 or FIG4 (i.e., an encoded audio bitstream containing eSBR metadata).

[0065] Referring again to FIG3, the deformatter 205 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (containing quantization envelope data) and eSBR metadata (and typically also other metadata) to assert at least the eSBR metadata and SBR metadata to the eSBR processing stage 203 and typically also to the decoding subsystem 202 (and optionally also to the control bit generator 204). The deformatter 205 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.

[0066] The system of FIG3 also optionally includes a post-processor 300. The post-processor 300 includes a buffer memory (buffer) 301 and other processing elements (not shown), the processing elements including at least one processing element coupled to the buffer 301. Buffer 301 stores (e.g., in a non-transitory manner) at least one block (or frame) of decoded audio data received by post-processor 300 from decoder 200. Processing elements of post-processor 300 are coupled and configured to receive and use metadata output from decoding subsystem 202 (and / or deformatter 205) and / or control bits output from stage 204 of decoder 200 to adapt the processing of the block (or frame) sequence of decoded audio output from buffer 301.

[0067] Audio decoding subsystem 202 of decoder 200 is configured to decode the audio data extracted by parser 205 (this decoding may be referred to as the "core" decoding operation) to produce decoded audio data and assert the decoded audio data to eSBR processing stage 203. Decoding is performed in the frequency domain and typically involves inverse quantization followed by spectral processing. Typically, the final processing stage in subsystem 202 applies a frequency-domain to time-domain transform to the decoded frequency-domain audio data, resulting in the subsystem output being time-domain decoded audio data. Stage 203 is configured to integrate the SBR tools and eSBR tools indicated by the eSBR metadata and eSBR (extracted by parser 205).The decoder is equipped with a mechanism to perform SBR and eSBR processing on the output of the decoding subsystem 202 using SBR and eSBR metadata to produce fully decoded audio data from the output of the decoder 200 (e.g., to the post-processor 300). Typically, the decoder 200 includes memory (accessible by subsystem 202 and stage 203) storing deformatted audio data and metadata from the deformatter 205 output, and stage 203 is configured to access audio data and metadata (including SBR metadata and eSBR metadata) as needed during SBR and eSBR processing. The SBR and eSBR processing in stage 203 can be considered as post-processing of the output of the core decoding subsystem 202. Decoder 200 also optionally includes a final upmixing subsystem (which can apply the parametric stereo (“PS”) tool defined in the MPEG-4 AAC standard using PS metadata extracted by deformatter 205 and / or control bits generated in subsystem 204), which is coupled and configured to perform upmixing on the output of stage 203 to produce fully decoded upmixed audio output from decoder 200. Alternatively, postprocessor 300 is configured to perform upmixing on the output of decoder 200 (e.g., using PS metadata extracted by deformatter 205 and / or control bits generated in subsystem 204).

[0068] In response to the metadata extracted by deformatter 205, control bit generator 204 can generate control data, which can be used within decoder 200 (e.g., for the final upmixing subsystem) and / or asserted as the output of decoder 200 (e.g., to postprocessor 300 for post-processing). In response to metadata extracted from the input bitstream (and optionally also in response to control data), stage 204 may generate control bits (and assert control bits to postprocessor 300) to indicate that the decoded audio data output from eSBR processing stage 203 should undergo a specific type of postprocessing. In some embodiments, decoder 200 is configured to assert metadata extracted from the input bitstream by deformatter 205 to postprocessor 300, and postprocessor 300 is configured to perform postprocessing on the decoded audio data output from decoder 200 using the metadata.

[0069] FIG4 is a block diagram of audio processing unit (“APU”) 210, which is another embodiment of the inventive audio processing unit. APU 210 is a conventional decoder not configured to perform eSBR processing. Any component or element of APU 210 may be implemented in hardware, software, or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuits). The APU 210 includes a buffer memory 201 connected as shown, a bitstream payload deformatter (parser) 215, an audio decoding subsystem 202 (sometimes referred to as the "core" decoding stage or "core decoding subsystem"), and an SBR processor.Processing stage 213. Typically, APU 210 also includes other processing elements (not shown). APU 210 may represent, for example, an audio encoder, decoder, or transcoder.

[0070] Elements 201 and 202 of APU 210 are identical to the same numbered elements of decoder 200 (FIG 3), and their above description will not be repeated. In operation of APU 210, a sequence of blocks of encoded audio bitstream (MPEG-4 AAC bitstream) received by APU 210 is asserted from buffer 201 to deformatter 215.

[0071] Deformatter 215 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (containing quantization envelope data) from it and typically also extracts other metadata from it, but ignores eSBR metadata that may be included in the bitstream according to any embodiment of the invention. Deformatter 215 is configured to assert at least SBR metadata to SBR processing stage 213. The deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.

[0072] The audio decoding subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the deformatter 215 (this decoding may be referred to as the "core" decoding operation) to produce decoded audio data and assert the decoded audio data to the SBR processing stage 213. Decoding is performed in the frequency domain. Typically, the final processing stage in subsystem 202 applies a frequency-to-time domain transformation to the decoded frequency-domain audio data, such that the output of the subsystem is time-domain decoded audio data. Stage 213 is configured to apply an SBR tool (but not an eSBR tool) indicated by SBR metadata (extracted by deformatter 215) to the decoded audio data (i.e., to perform SBR processing on the output of decoding subsystem 202 using SBR metadata) to produce fully decoded audio data output from APU 210 (e.g., to post-processor 300). Typically, APU 210 includes memory (accessible by subsystem 202 and stage 213) storing the deformatted audio data and metadata output from deformatter 215, and stage 213 is configured to access the audio data and metadata (including SBR metadata) as needed during SBR processing. The SBR processing in stage 213 can be considered as post-processing of the output of core decoding subsystem 202. The APU 210 also optionally includes a final upmixing subsystem (which can apply the parametric stereo “PS” tool defined in the MPEG-4 AAC standard using PS metadata extracted by the deformatter 215), which is coupled and configured to perform upmixing on the output of stage 213 to produce fully decoded upmixed audio output from the APU 210. Alternatively, a post-processor is configured to...The output of 210 is upmixed (e.g., using PS metadata extracted by deformatter 215 and / or control bits generated in APU 210).

[0073] Various embodiments of encoder 100, decoder 200, and APU 210 are configured to perform different embodiments of the inventive method.

[0074] According to some embodiments, eSBR metadata (e.g., a small number of control bits of eSBR metadata) is included in the encoded audio bitstream (e.g., an MPEG-4 AAC bitstream), such that a conventional decoder (which is not configured to parse eSBR metadata or use any eSBR tool associated with eSBR metadata) can ignore the eSBR metadata but still decode the bitstream as much as possible without using eSBR metadata or any eSBR tool associated with eSBR metadata, typically without significant loss of decoded audio quality. However, an eSBR decoder (which is configured to parse the bitstream to identify eSBR metadata and uses at least one eSBR tool in response to eSBR metadata) will benefit from using at least one of these eSBR tools. Therefore, embodiments of the present invention provide a method for efficiently transmitting enhanced spectrum band replication (eSBR) control data or metadata in a backward-compatible manner.

[0075] Typically, eSBR metadata in a bitstream indicates one or more of the following eSBR tools (e.g., indicating at least one of their characteristics or parameters) (the eSBR tools are described in the MPEG USAC standard and may or may not be applied by the encoder during bitstream generation):

[0076] ● Harmonic transpose; and

[0077] ● QMF patching additional preprocessing (pre-flattening).

[0078] For example, eSBR metadata included in a bitstream may indicate the values ​​of parameters (as described in the MPEG USAC standard and in this invention): sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch], and bs_sbr_preprocessing.

[0079] In this document, the symbol X[ch] (where X is a parameter) indicates that the parameter relates to the channel (“ch”) of the audio content of the encoded bitstream to be decoded. For simplicity, we sometimes omit the expression [ch] and assume that the relevant parameter relates to the channel of the audio content.

[0080] In this document, the symbol X[ch][env] (where X is a parameter) indicates that the parameter relates to the SBR envelope (“env”) of the channel (“ch”) of the audio content of the encoded bitstream to be decoded. For simplicity, we sometimes omit the expressions [env] and [ch] and assume that the relevant parameter relates to the SBR envelope of the channel of the audio content.

[0081] During the decoding of the encoded bitstream, harmonic transposition is performed during the eSBR processing stage of the decoded data (for each channel "ch" of the audio content indicated by page 9 / 31 of the bitstream specification, CN 121214953 A) and is controlled by the following eSBR metadata parameters: sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBinsFlag[ch], and sbrPitchInBins[ch].

[0082] The value "sbrPatchingMode[ch]" indicates the transpose type used in the eSBR: sbrPatchingMode[ch] = 1 indicates linear transpose patching as described in section 4.6.18 of the MPEG-4 AAC standard (used with high-quality SBR or low-power SBR); sbrPatchingMode[ch] = 0 indicates harmonic SBR patching as described in section 7.5.3 or 7.5.4 of the MPEG USAC standard.

[0083] The value “sbrOversamplingFlag[ch]” indicates that adaptive frequency domain oversampling in the eSBR is used in combination with DFT-based harmonic SBR patching as described in section 7.5.3 of the MPEG USAC standard. This flag controls the size of the DFT used in the transposer: 1 indicates that adaptive frequency domain oversampling is enabled as described in section 7.5.3.1 of the MPEG USAC standard; 0 indicates that adaptive frequency domain oversampling is disabled as described in section 7.5.3.1 of the MPEG USAC standard.

[0084] The value “sbrPitchInBinsFlag[ch]” controls the interpretation of the sbrPitchInBins[ch] parameter: 1 indicates that the value of sbrPitchInBins[ch] is valid and greater than 0; 0 indicates that the value of sbrPitchInBins[ch] is set to 0.

[0085] The value “sbrPitchInBins[ch]” controls the addition of cross-product terms in the SBR harmonic transposer. The value sbrPitchinBins[ch] is an integer value in the range [0, 127] and represents the distance measured in a frequency grid of a 1536-line DFT acting on the sampling frequency of the core encoder.

[0086] If the MPEG-4 AAC bitstream indicates a pair of uncoupled SBR channels (rather than a single SBR channel), then the bitstream indicates two instances of the above syntax (for harmonic or non-harmonic transpose), one instance sbr_channel_pair_element() for each channel.

[0087] Harmonic transpose of the eSBR tool generally improves the quality of decoded music signals at relatively low cross frequencies. Non-harmonicWave transposition (i.e., conventional spectral patching) typically improves speech signals. Therefore, the starting point for determining which type of transposition is preferred for encoding specific audio content is to select the transposition method based on speech / music detection, where harmonic transposition is used for music content and spectral patching is used for tempo content.

[0088] Pre-flattening during eSBR processing is controlled by the value of a 1-bit eSBR metadata parameter called “bs_sbr_preprocessing”, which, in a sense, determines whether or not pre-flattening is performed based on the value of this single bit. When using the SBR QMF patching algorithm described in section 4.6.18.6.3 of the MPEG-4 AAC standard, a pre-flattening step (when indicated by the “bs_sbr_preprocessing” parameter) may be performed to attempt to avoid shape discontinuities in the spectral envelope of the high-frequency signal input to the subsequent envelope adjuster (another stage of eSBR processing performed by the envelope adjuster). Pre-flattening typically improves the operation of the subsequent envelope adjustment stage, resulting in a perceived more stable high-frequency band signal.

[0089] According to some embodiments of the invention, the total bit rate requirement included in the MPEG-4 AAC bitstream eSBR metadata indicating the aforementioned eSBR tools (harmonic transposition and pre-flattening) is expected to be approximately several hundred bits per second, since only differential control data required to perform eSBR processing is transmitted. Conventional decoders can ignore this information because it is included in a backward-compatible manner (as explained later). Therefore, the adverse effects on bit rate associated with including eSBR metadata are negligible due to several reasons, including:

[0090] ● Bit rate loss (attributed to including eSBR metadata) accounts for a very small percentage of the total bit rate, since only differential control data required to perform eSBR processing (and cascading of non-SBR control data) is transmitted; and

[0091] ● Tuning of SBR-related control information generally does not depend on the details of the transposition. Examples of control data depending on the operation of the transposer will be discussed later in this application.

[0092] Therefore, embodiments of the invention provide a method for efficiently transmitting enhanced spectrum band replication (eSBR) control data or metadata in a backward-compatible manner. This efficient transmission of eSBR control data reduces the memory requirements in decoders, encoders, and transcoders using aspects of the invention, without significantly adversely affecting the bit rate. Furthermore, it reduces the complexity and processing requirements associated with implementing eSBR according to embodiments of the invention, since the SBR data only needs to be processed once and is not rebroadcast, as is the case when eSBR is treated as a completely independent object type in MPEG-4 AAC rather than being integrated into the MPEG-4 AAC codec in a backward-compatible manner.

[0093] Next, referring to FIG7, we describe the elements of a block (“raw_data_block”) of an MPEG-4 AAC bitstream (containing eSBR metadata) according to some embodiments of the present invention. FIG7 is a diagram of a block (“raw_data_block”) of an MPEG-4 AAC bitstream, showing some segments of the MPEG-4 AAC bitstream.

[0094] A block of an MPEG-4 AAC bitstream may contain at least one “single_channel_element()” (e.g., a single-channel element shown in FIG7) and / or at least one “channel_pair_element()” (not explicitly shown in FIG7, but may exist), which contains audio data of an audio program. The block may also contain several “fill_element” (e.g., fill element 1 and / or fill element 2 in FIG7), which contain program-related data (e.g., metadata). Each “single_channel_element()” contains an identifier indicating the start of a single-channel element (e.g., “ID1” in FIG7) and may contain audio data indicating different channels of a multi-channel audio program. Each “channel_pair_element()” contains an identifier indicating the start of the channel pair element (not shown in Figure 7) and may contain audio data indicating the two channels of the program.

[0095] The fill_element of the MPEG-4 AAC bitstream (referred to herein as the fill element) contains an identifier indicating the start of the fill element (“ID2” in Figure 7) and fill data following the identifier. The identifier ID2 may consist of a 3-bit unsigned integer (“uimsbf”) with a value of 0×6, transmitting the most significant bit first. The fill data may contain extension_payload() elements (sometimes referred to herein as extended payloads) whose syntax is shown in Table 4.57 of the MPEG-4 AAC standard. Several types of extended payloads exist and are identified by the “extension_type” parameter, which is a 4-bit unsigned integer (“uimsbf”) transmitting the most significant bit first.

[0096] The padding data (e.g., its extended payload) may contain a header or identifier (e.g., “Header 1” in FIG7) indicating a segment of the padding data (which indicates an SBR object) (i.e., the header initializes the “SBR object” type referred to as sbr_extension_data() in the MPEG-4 AAC standard). For example, the value of “1101” or “1110” in the extension_type field of the header is used to identify the Spectral Band Replication (SBR) extended payload, where the identifier “1101” identifies a payload with SBR data.The extended payload and “1110” identify the extended payload containing SBR data with cyclic redundancy check (CRC) to verify the correctness of the SBR data.

[0097] When the header (e.g., the extension_type field) initializes the SBR object type, the SBR metadata (sometimes referred to herein as “spectral band copy data”, and referred to herein as sbr_data() in the MPEG-4 AAC standard) follows the header, and at least one spectral band copy extension element (e.g., the “SBR extension element” of padding element 1 in FIG7) may follow the SBR metadata. This spectral band copy extension element (a segment of the bitstream) is referred to as the “sbr_extension()” container in the MPEG-4 AAC standard. The spectral band copy extension element optionally contains a header (e.g., the “SBR extension header” of padding element 1 in FIG7).

[0098] The MPEG-4 AAC standard anticipates that the spectral band copy extension element may contain PS (parametric stereo) data for the audio data of the program. The MPEG-4 AAC standard anticipates that when the header of a padding element (e.g., its extended payload) initializes the SBR object type (as shown in "Header 1" in FIG7) and the spectral band copy extension element of the padding element contains PS data, the padding element (e.g., its extended payload) contains spectral band copy data and a "bs_extension_id" parameter, the value of which (i.e., bs_extension_id=2) indicates that PS data is included in the spectral band copy extension element of the padding element.

[0099] According to some embodiments of the invention, eSBR metadata (e.g., a flag indicating whether enhanced spectral band copy (eSBR) processing is performed on the audio content of the block) is included in the spectral band copy extension element of the padding element. For example, this flag is indicated in padding element 1 of FIG7, where the flag appears after the header of the "SBR extension element" of padding element 1 (page 11 / 31 of the specification of padding element 1, 13 CN 121214953 A "SBR extension header"). This tag and additional eSBR metadata are optionally included in the spectrum band copy extension element after the header of the spectrum band copy extension element (e.g., in the SBR extension element of padding element 1 in FIG. 7 after the SBR extension header). According to some embodiments of the invention, the padding element containing eSBR metadata also includes a “bs_extension_id” parameter, the value of which (e.g., bs_extension_id=3) indicates that the eSBR metadata is included in the padding element and that eSBR processing is performed on the audio content of the relevant block.

[0100] According to some embodiments of the invention, the eSBR metadata is included in the padding element of the MPEG-4 AAC bitstream (e.g.,In Figure 7, padding element 2) is not a padding element but a Spectral Band Replication Extension (SBR) Extension element. This is because padding elements containing extension_payload() (which has SBR data or SBR data with CRC) do not contain any other extension payloads of any other extension type. Therefore, in embodiments where the eSBR metadata stores its own extension payload, a separate padding element is used to store the eSBR metadata. This padding element contains an identifier indicating the start of the padding element (e.g., “ID2” in Figure 7) and padding data following the identifier. The padding data may contain extension_payload() elements (sometimes referred to herein as extension payloads) whose syntax is shown in Table 4.57 of the MPEG-4 AAC standard. The padding data (e.g., its extension payload) contains a header indicating the eSBR object (e.g., “Header 2” in padding element 2 of Figure 7) (i.e., the header initializes the Enhanced Spectral Band Replication (eSBR) object type), and the padding data (e.g., its extension payload) contains the eSBR metadata following the header. For example, padding element 2 of Figure 7 contains this header (“Header 2”) and also contains eSBR metadata following the header (i.e., “tags” in padding element 2 that indicate whether Enhanced Spectral Band Replication (eSBR) processing is performed on the audio content of the block). Additional eSBR metadata is also optionally included in the padding data of padding element 2 of Figure 7 following Header 2. In the embodiments described in this paragraph, the header (e.g., Header 2 of Figure 7) has an identification value that is not the conventional value specified in Table 4.57 of the MPEG-4 AAC standard, but instead indicates the eSBR extended payload (such that the extension_type field of the header indicates that the padding data contains eSBR metadata).

[0101] In a first-type embodiment, the present invention is an audio processing unit (e.g., a decoder) comprising:

[0102] a memory (e.g., buffer 201 of FIG. 3 or 4) configured to store at least one block of an encoded audio bitstream (e.g., at least one block of an MPEG-4 AAC bitstream);

[0103] a bitstream payload deformatter (e.g., element 205 of FIG. 3 or element 215 of FIG. 4) coupled to the memory and configured to demultiplex at least a portion of the block of the bitstream; and

[0104] a decoding subsystem (e.g., elements 202 and 203 of FIG. 3 or elements 202 and 213 of FIG. 4) coupled and configured to decode at least a portion of the audio content of the block of the bitstream, wherein the block comprises:

[0105] a padding element comprising an identifier indicating the start of the padding element (e.g., having an MPEG-4 AAC standard).The value of 0×6 in Table 4.85 for the “id_syn_ele” identifier and the padding data following the identifier, wherein the padding data includes:

[0106] at least one flag that identifies whether enhanced spectral band copying (eSBR) processing is performed on the audio content of the block (e.g., using spectral band copying data and eSBR metadata contained in the block).

[0107] The flag is eSBR metadata, and an example of the flag is the sbrPatchingMode flag. Another example of the flag is the harmonicSBR flag. Both of these flags indicate whether the basic form of spectral band copying or the enhanced form of spectral copying is performed on the audio data of the block. The basic form of spectral copying is spectral patching, and the enhanced form of spectral band copying is harmonic transpose.

[0108] In some embodiments, the padding data also includes additional eSBR metadata (i.e., eSBR metadata other than the flags).

[0109] The memory may be a buffer memory (e.g., an embodiment of buffer 201 of FIG4) that stores (e.g., in a non-transitory manner, on page 12 / 31 of the specification, CN 121214953 A) at least one block of the encoded audio bitstream.

[0110] It is estimated that the complexity of performing eSBR processing (using eSBR harmonic transpose and pre-flattening) by the eSBR decoder during the decoding of an MPEG-4 AAC bitstream containing eSBR metadata (indicating these eSBR tools) will be as follows (for typical decoding with indicative parameters):

[0111] ● Harmonic transpose (16 kbps, 14400 / 28800 Hz)

[0112] ○ DFT-based: 3.68 WMOPS (weighted million operations per second);

[0113] ○ QMF-based: 0.98 WMOPS;

[0114] ● QMF patching preprocessing (pre-flattening): 0.1 WMOPS.

[0115] It is well known that, for transients, DFT-based transpose generally performs better than QMF-based transpose.

[0116] According to some embodiments of the present invention, the padding element containing eSBR metadata (encoded audio bitstream) also includes a parameter whose value (e.g., bs_extension_id=3) indicates that eSBR metadata is contained in the padding element and that eSBR processing is performed on the audio content of the relevant block (e.g., the "bs_extension_id" parameter) and / or its value (e.g., bs_extension_id=2) indicates that the sbr_extension() container of the padding element contains PS data (e.g., the same "bs_extension_id" parameter).The parameter is “id”. For example, as indicated in Table 1 below, this parameter with the value bs_extension_id=2 indicates that the sbr_extension() container of the padding element contains PS data, and this parameter with the value bs_extension_id=3 indicates that the sbr_extension() container of the padding element contains eSBR metadata:

[0117] Table 1

[0118]

[0119] According to some embodiments of the present invention, the syntax of each spectrum band replication extension element containing eSBR metadata and / or PS data is indicated in Table 2 below (where “sbr_extension()” indicates that it is a container of spectrum band replication extension element, “bs_extension_id” is as described in Table 1 above, “ps_data” indicates PS data, and “esbr_data” indicates eSBR metadata):

[0120] Table 2 Specification 13 / 31 pages 15 CN 121214953 A

[0121]

[0122] In an exemplary embodiment, the esbr_data( The following metadata parameters are indicated:

[0123] 1. A 1-bit metadata parameter “bs_sbr_preprocessing”; and

[0124] 2. For each channel (“ch”) of the audio content of the encoded bitstream to be decoded, each of the above parameters is “sbrPatchingMode[ch]”, “SbrOversamplingFlag[ch]”, “SbrPitchInBinsFlag[ch]”, and “sbrPitchInBins[ch]”.

[0125] For example, in some embodiments, esbr_data() may have the syntax indicated in Table 3 to indicate these metadata parameters:

[0126] Table 3

[0127] Specification 14 / 31 pages 16 CN 121214953 A

[0128] Specification 15 / 31 pages 17 CN 121214953 A

[0129]

[0130] The above syntax enables the efficient implementation of enhanced forms of spectral band copying (e.g., harmonic transpose) as an extension of conventional decoders. Specifically, the eSBR data in Table 3 contains only the parameters required to perform the enhanced form of spectral band copying, which are no longer supported in the bitstream and cannot be directly derived from parameters already supported in the bitstream. All other parameters and processing data required to perform the enhanced form of spectral band copying are extracted from readily available parameters in defined locations in the bitstream.

[0131] For example, an MPEG-4 HE-AAC or HE-AAC v2 compatible decoder can be extended to include enhanced forms of spectral band copying.For example, harmonic transpose. This enhanced form of spectrum band copying is an addition to the basic form of spectrum band copying already supported by the decoder. In the context of MPEG-4 HE-AAC or HE-AAC v2 compatible decoders, this basic form of spectrum band copying is the QMF spectrum patching SBR tool, as defined in section 4.6.18 of the MPEG-4 AAC standard.

[0132] When performing the enhanced form of spectrum band copying, the extended HE-AAC decoder can reuse many bitstream parameters already included in the SBR extension payload of the bitstream. Specific parameters that can be reused include, for example, various parameters that determine the main frequency table. These parameters include bs_start_freq (a parameter that determines the start of the main frequency table parameters), bs_stop_freq (a parameter that determines the stop of the main frequency table), bs_freq_scale (a parameter that determines the number of bands per octave), and bs_alter_scale (a parameter that modifies the proportion of the bands). Reusable parameters also include parameters that determine the noise band table (bs_noise_bands) and limiter band table parameters (bs_limiter_bands). Therefore, in various embodiments, at least some equivalent parameters specified in the USAC standard are omitted from the bit stream specification on page 16 / 31 of 18 CN 121214953 A to reduce the control burden of the bit stream. Typically, when the parameters specified in the AAC standard have equivalent parameters specified in the USAC standard, the equivalent parameters specified in the USAC standard have the same names as the parameters specified in the AAC standard, such as the envelope scaling factor EOrigMapped. However, the equivalent parameters specified in the USAC standard typically have different values, which are “tuned” according to the enhanced SBR processing defined in the USAC standard rather than the SBR processing defined in the AAC standard.

[0133] It is recommended to enable enhanced SBR to improve the subjective quality of audio content with harmonic frequency structure and strong tonal characteristics, especially at low bit rates. The value of the corresponding bitstream element (i.e., esbr_data()) controlling these tools can be determined in the encoder by applying a signal dependency classification mechanism. Generally, the use of the harmonic patching method (sbrPatchingMode==1) is preferred for encoding music signals at very low bit rates, where the audio bandwidth of the core codec is greatly limited. This is particularly prominent when these signals contain significant harmonic structures. Conversely, the use of the conventional SBR patching method is preferred for speech and mixed signals because it provides better preservation of the temporal structure of speech.

[0134] To improve the performance of the harmonic transposer, a preprocessing step (bs_sbr_preprocessing==1) can be initiated, which attempts toThe figure avoids introducing spectral discontinuities of the signal into the subsequent envelope adjuster. The operation of the tool is beneficial for signal types in which the coarse spectral envelope of the low-frequency band signal used for high-frequency reconstruction shows a large level of variation.

[0135] To improve the transient response of harmonic SBR patching, adaptive frequency domain oversampling (sbrOversamplingFlag==1) can be applied. Since adaptive frequency domain oversampling increases the computational complexity of the transposer and only benefits frames containing transients, the use of this tool is controlled by the bitstream element, which is transmitted once per frame and per independent SBR channel.

[0136] Decoders operating in the proposed enhanced SBR mode typically need to be able to switch between conventional SBR patching and enhanced SBR patching. Therefore, a delay that can be as long as the duration of a core audio frame can be introduced depending on the decoder settings. Typically, the delays for conventional SBR patching and enhanced SBR patching will be similar.

[0137] In addition to many parameters, other data elements can also be reused by the extended HE-AAC decoder when performing the enhanced form of spectral band copying according to an embodiment of the invention. For example, envelope data and noise floor data can also be extracted from bs_data_env (envelope scaling factor) and bs_noise_env (noise floor scaling factor) data and used during the enhanced form of spectrum band replication.

[0138] Essentially, these embodiments utilize configuration parameters and envelope data already supported by conventional HE-AAC or HE-AAC v2 decoders in the SBR extended payload to enable the enhanced form of spectrum band replication, which requires as little additional transmission data as possible. The metadata is initially tuned according to the basic form of HFR (e.g., spectrum shift operation of SBR), but according to embodiments, it is used for the enhanced form of HFR (e.g., harmonic transpose of eSBR). As previously discussed, the metadata generally represents the operating parameters (e.g., envelope scaling factor, noise floor scaling factor, time / frequency grid parameters, sine wave addition information, variable cross frequency / band, inverse filtering mode, envelope resolution, smoothing mode, frequency interpolation mode) tuned and designed to be used with the basic form of HFR (e.g., linear spectrum shift). However, this metadata can be combined with additional metadata parameters specifically for HFR enhancements (e.g., harmonic transpose) to efficiently and effectively process audio data using HFR enhancements.

[0139] Therefore, an extended decoder supporting spectral band replication can be generated very efficiently by relying on defined bitstream elements (e.g., bitstream elements in the SBR extended payload) and only adding the parameters required for the enhanced form supporting spectral band replication (in the padding element extended payload). This data reduction feature is achieved by placing the newly added parameters inThe combination of reserved data fields (e.g., extended containers) largely reduces the barriers to generating a decoder that supports the enhanced form of spectral band replication by ensuring backward compatibility of the bitstream with a legacy decoder that does not support spectral band replication. Specification 17 / 31 pages 19 CN 121214953 A

[0140] In Table 3, the numbers in the right row indicate the number of bits of the corresponding parameter in the left row.

[0141] In some embodiments, the SBR object type defined in MPEG-4 AAC is updated to include aspects of the SBR tool and the enhanced SBR (eSBR) tool, as indicated in the SBR extension element (bs_extension_id==EXTENSION_ID_ESBR). If the decoder detects and supports this SBR extension element, then the decoder adopts the indicated aspect of the enhanced SBR tool. The SBR object type updated in this way is called SBR enhancement.

[0142] In some embodiments, the present invention is a method comprising the step of encoding audio data to produce an encoded bitstream (e.g., an MPEG-4 AAC bitstream), comprising including eSBR metadata in at least one segment of at least one block of the encoded bitstream and audio data in at least another segment of said block. In a typical embodiment, the method comprises the step of multiplexing the audio data and eSBR metadata in each block of the encoded bitstream. In typical decoding of an encoded bitstream in an eSBR decoder, the decoder extracts eSBR metadata from the bitstream (including parsing and demultiplexing the eSBR metadata and audio data) and uses the eSBR metadata to process the audio data to produce a decoded audio data stream.

[0143] Another aspect of the present invention is an eSBR decoder configured to perform eSBR processing (e.g., using at least one of eSBR tools called harmonic transpose or pre-flattening) during the decoding of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that does not contain eSBR metadata. An example of this decoder will be described with reference to FIG5.

[0144] The eSBR decoder 400 of FIG. 5 includes a buffer memory 201 (which is the same as the memory 201 of FIG. 3 and 4) connected as shown, a bitstream payload deformatter 215 (which is the same as the deformatter 215 of FIG. 4), an audio decoding subsystem 202 (sometimes referred to as the “core” decoding stage or “core” decoding subsystem, and is the same as the core decoding subsystem 202 of FIG. 3), an eSBR control data generation subsystem 401, and an eSBR processing stage 203 (which is the same as stage 203 of FIG. 3). Typically, the decoder 400 also includes other processing elements (not shown).

[0145] In the operation of the decoder 400, the encoded audio bitstream (MPEG-4 AAC bitstream) received by the decoder 400 is processed...The block sequence is asserted from buffer 201 to deformatter 215.

[0146] Deformatter 215 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (containing quantization envelope data) from it and typically also extract other metadata from it. Deformatter 215 is configured to assert at least the SBR metadata to eSBR processing stage 203. Deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to decoding subsystem (decoding stage) 202.

[0147] Audio decoding subsystem 202 of decoder 400 is configured to decode the audio data extracted by deformatter 215 (this decoding may be referred to as the “core” decoding operation) to produce decoded audio data and assert the decoded audio data to eSBR processing stage 203. Decoding is performed in the frequency domain. Typically, the final processing stage in subsystem 202 applies a frequency-domain to time-domain transformation to the decoded frequency-domain audio data, resulting in time-domain decoded audio data as the subsystem output. Stage 203 is configured to apply SBR tools (and eSBR tools) indicated by SBR metadata (extracted by deformatter 215) and eSBR metadata generated in subsystem 401 to the decoded audio data (i.e., performing SBR and ESBR processing on the output of decoding subsystem 202 using SBR and eSBR metadata) to produce fully decoded audio data output from decoder 400. Typically, decoder 400 includes memory (accessible by subsystem 202 and stage 203) storing deformatted audio data and metadata output from deformatter 215 (and optionally subsystem 401), and stage 203 is configured to access audio data and metadata as needed during SBR and eSBR processing. The SBR processing in stage 203 can be considered as post-processing of the output of core decoding subsystem 202. Decoder 400 also optionally includes a final upmixing subsystem (which can use PS metadata extracted by deformatter 215 to apply the parametric stereo “PS” tool as defined in the MPEG-4 AAC standard), which is coupled and configured to perform upmixing on the output of stage 203 to produce fully decoded upmixed audio output from APU 210.

[0148] Parametric stereo is an encoding tool that uses linear downmixing of the left and right channels of a stereo signal and a set of spatial parameters describing the stereo picture specification 18 / 31 pages 20 CN 121214953 A to represent a stereo signal. Parametric stereo typically uses three types of spatial parameters: (1) inter-channel intensity difference (IID), which describes the intensity difference between channels; (2) inter-channel phase difference (IPD), which describes the phase difference between channels; and (3) inter-channel coherence (ICC), which describes the coherence (or similarity) between channels. Coherence can beThe measurement is the maximum value of the cross-correlation that varies with time or phase. These three parameters typically enable high-quality reconstruction of stereo images. However, the IPD parameter only specifies the relative phase difference between channels of the stereo input signal and does not indicate the distribution of these phase differences on the left and right channels. Therefore, a fourth type of parameter describing the total phase shift or total phase difference (OPD) can be used. During stereo reconstruction, consecutive window segments of both the received downmixed signal s[n] and the received decorrelated version d[n] of the downmixed signal are processed together with spatial parameters to generate left (lk(n)) and right (rk(n)) reconstructed signals according to the following equations:

[0149]

[0150]

[0151] where H11, H12, H21, and H22 are defined by stereo parameters. Finally, the signals lk(n) and rk(n) are transformed back to the time domain by frequency-to-time transformation.

[0152] The control data generation subsystem 401 of FIG5 is coupled and configured to detect at least one property of the encoded audio bitstream to be decoded and to generate eSBR control data (which may be or contain any type of eSBR metadata contained in the encoded audio bitstream according to other embodiments of the invention) in response to at least one result of the detection step. The eSBR control data is asserted to level 203 to trigger the application of individual eSBR tools or combinations of eSBR tools and / or control the application of these eSBR tools after a specific property (or combination of properties) of the bitstream is detected. For example, to control the execution of eSBR processing using harmonic transposition, some embodiments of the control data generation subsystem 401 would include: a music detector (e.g., a simplified version of a conventional music detector) for setting the sbrPatchingMode[ch] parameter in response to detecting whether the bitstream indicates music (and asserting the setting parameter to level 203); a transient detector for setting the sbrOversamplingFlag[ch] parameter in response to detecting the presence or absence of transients in the audio content indicated by the bitstream (and asserting the setting parameter to level 203); and / or a spacing detector for setting the sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch] parameters in response to detecting spacing in the audio content indicated by the bitstream (and asserting the setting parameter to level 203). Other aspects of the invention are audio bitstream decoding methods performed by any embodiment of the inventive decoder described in this and previous paragraphs.

[0153] Aspects of the invention include encoding or decoding method types configured (e.g., programmed) to perform any embodiment of the inventive APU, system, or apparatus. Other aspects of the invention include methods configured (e.g., programmed) to perform the inventive methods.The system or apparatus of any embodiment and computer-readable medium (e.g., optical disc) storing (e.g., in a non-transitory manner) code for implementing any embodiment of the inventive method or steps thereof. For example, the inventive system may be or include a programmable general-purpose processor, digital signal processor, or microprocessor that is programmed with software or firmware and / or otherwise configured to perform various operations on data (including embodiments of the inventive method or steps thereof). This general-purpose processor may be or include a computer system that includes input devices, memory, and processing circuitry programmed (and / or otherwise configured) to perform embodiments of the inventive method (or steps thereof) in response to data asserted thereto.

[0154] Embodiments of the invention may be implemented in hardware, firmware, or software, or a combination of both (e.g., as a programmable logic array). Unless otherwise stated, the algorithms or processes included as part of the invention are not inherently associated with any particular computer or other device. Specifically, various general-purpose machines may be used with programs written in accordance with the teachings herein, or may be more readily adapted to construct more specialized devices (e.g., integrated circuits) to perform the desired method steps. Therefore, the present invention can be implemented in one or more computer programs that execute on one or more programmable computer systems (e.g., embodiments of any of the elements of FIG1, or encoder 100 (or elements thereof) of FIG2, or decoder 200 (or elements thereof) of FIG3, or decoder 210 (or elements thereof) of FIG4, or decoder 400 (or elements thereof) of FIG5), each of the one or more programmable computer systems including at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and to produce output information. The output information is applied to one or more output devices in a known manner.

[0155] Each of these programs can be implemented in any desired computer language (including machine, assembly or high-level programming, logic or object-oriented programming languages) to communicate with the computer system. In any case, the language can be a compiled or interpreted language.

[0156] For example, when implemented by a sequence of computer software instructions, the various functions and steps of embodiments of the present invention can be implemented by a multi-threaded sequence of software instructions that run in suitable digital signal processing hardware, in which case the various means, steps, and functions of the embodiments may correspond to portions of the software instructions.

[0157] Each computer program is preferably stored on or downloaded to a storage medium or device (e.g., solid-state memory or media or magnetic or optical media) that can be read by a general-purpose or special-purpose programmable computer to configure and operate the computer when the storage medium or device is read by a computer system to execute the program described herein. The system of the present invention can also be implemented via...A computer-readable storage medium configured (i.e., storing) a computer program, wherein such a storage medium causes a computer system to operate in a particular and predefined manner to perform the functions described herein.

[0158] Many embodiments of the invention have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the invention. Many modifications and variations of the invention can be made in light of the foregoing teachings. For example, to facilitate efficient implementation, phase shifts can be combined with complex QMF analysis and synthesis filter banks. The analysis filter bank is responsible for filtering the time-domain low-frequency signal generated by the core decoder into multiple sub-bands (e.g., QMF sub-bands). The synthesis filter bank is responsible for combining the regenerated high-frequency band generated by a selected HFR technique (as indicated by the received sbrPatchingMode parameter) with the decoded low-frequency band to produce a wideband output audio signal. However, a given filter bank implementation operating in a certain sampling rate mode (e.g., normal dual-rate operation or downsampled SBR mode) should not have a phase shift dependent on the bitstream. The QMF bank used in SBR is a theoretical complex exponential extension of the cosine modulation filter bank. It can be shown that when using complex exponential modulation to extend the cosine modulation filter bank, the frequency overlap cancellation constraint becomes obsolete. Therefore, for the SBR QMF bank, both the analysis filter hk(n) and the synthesis filter fk(n) can be defined by the following equations:

[0159] , 0≤n≤N; 0≤k≤M (1)

[0160] where p0(n) is a real-valued symmetric or asymmetric prototype filter (usually a low-pass prototype filter), M represents the number of channels, and N is the order of the prototype filter. The number of channels used in the analysis filter bank can be different from the number of channels used in the synthesis filter bank. For example, the analysis filter bank can have 32 channels and the synthesis filter bank can have 64 channels. When operating the synthesis filter bank in downsampling mode, the synthesis filter bank can have only 32 channels. Since the subband samples from the filter bank are complex values, an additive feasible channel-dependent phase shift step can be added to the analysis filter bank. These additional phase shifts need to be compensated before the synthesis filter bank. Although the phase shift term can theoretically have any value without disrupting the operation of the QMF analysis / synthesis chain, it can also be constrained to certain values ​​for consistency verification. The SBR signal is affected by the choice of the phase factor, while the low-pass signal from the core decoder is not. The audio quality of the output signal is unaffected.

[0161] The coefficients p0(n) of the prototype filter can be defined as a length L of 640, as shown in Table 4 below.

[0162] Table 4 Specification 20 / 31 pages 22 CN 121214953 A

[0163] Specification 21 / 31 pages 23 CN 121214953 A

[0164] Specification 22 / 31 pages 24 CN 121214953 A

[0165] Specification 23 / 31 pages 25 CN 121214953 A

[0166] Specification 24 / 31 pages 26 CN 121214953 A

[0167] Specification 25 / 31 pages 27 CN 121214953 A

[0168] Specification 26 / 31 pages 28 CN 121214953 A

[0169]

[0170] The prototype filter p0(n) can also be derived from Table 4 by one or more mathematical operations, such as rounding, subsampling, interpolation and sampling.

[0171] Although the tuning of SBR-related control information generally does not depend on the details of the transpose (as previously discussed), in some embodiments, certain elements of the control data may be cascaded in the eSBR extension container (bs_extension_id==EXTENSION_ID_ [Specification 27 / 31 pages 29 CN 121214953 A ESBR]) to improve the quality of the regenerated signal. Some cascaded elements may include noise floor data (e.g., noise floor scaling factor and parameters indicating the direction (frequency or time direction) of differential encoding for each noise floor), inverse filtering data (e.g., parameters indicating inverse filtering modes selected from no inverse filtering, low inverse filtering, moderate inverse filtering, and strong inverse filtering), and missing harmonic data (e.g., parameters indicating whether a sine wave should be added to a specific frequency band of the regenerated high-frequency band). All these elements depend on the synthetic simulation of the transpose of the decoder performed in the encoder and thus can improve the quality of the regenerated signal after proper tuning according to the selected transpose.

[0172] Specifically, in some embodiments, missing harmonic and inverse filter control data (along with other bitstream parameters in Table 3) are transmitted in the eSBR extension container and tuned according to the harmonic transposer of the eSBR. The additional bit rate required to transmit these two types of metadata of the harmonic converter of the eSBR is relatively low. Therefore, sending the tuning missing harmonic and / or inverse filter control data in the eSBR extension container will improve the quality of the audio produced by the transposer while only slightly affecting the bit rate. To ensure backward compatibility with conventional decoders, parameters tuned for the SBR spectral shift operation can also be sent as part of the SBR control data in the bitstream using implicit or explicit signaling.

[0173] The complexity of the decoder with SBR enhancement described in this application must be limited so as not to significantly increase the overall computational complexity of the implementation. Preferably, when using eSBR tools, the PCU (MOP) of the SBR object type is equal to or less than 4.5, and the RCU of the SBR object type is equal to or less than 3. Approximate processing power is given in processor complexity units (PCUs) (specified by an integer number of MOPS). Approximate RAM usage is given in RAM complexity units (RCUs) (specified by kWords).(The integer number of 1000 words) is given. The number of RCUs does not include working buffers that can be shared between different objects and / or channels. In addition, the PCU is proportional to the sampling frequency. The PCU value is given in MOPS (millions of operations per second) per channel and the RCU value is given in kilowords per channel.

[0174] Special attention needs to be paid to compressed data, such as HE-AAC encoded audio that can be decoded by different decoder configurations. In this case, decoding can be performed in a backward compatible manner (AAC only) as well as in an enhanced manner (AAC+SBR). If the compressed data allows both backward compatible and enhanced decoding, and if the decoder operates in an enhanced manner so that it uses a post-processor that inserts some additional delay (e.g., an SBR post-processor in HE-AAC), then it must be ensured that this additional time delay caused by the backward compatible mode is taken into account when presenting the combined unit, as described by the corresponding value n. To ensure proper handling of the combined timestamp (so that the audio is synchronized with other media), when the decoder operating mode includes the SBR enhancement described in this application (including eSBR), the additional delay introduced by post-processing, given the number of samples (per audio channel) at the output sampling rate, is 3010. Therefore, for the audio combining unit, when the decoder operating mode includes the SBR enhancement described in this application, the combining time is applied to the 3011th audio sample within the combining unit.

[0175] SBR enhancement should be enabled to improve the subjective quality of audio content with harmonic frequency structures and strong tonal characteristics, especially at low bit rates. The value of the corresponding bitstream element (i.e., esbr_data()) controlling these tools can be determined in the encoder by applying a signal dependency classification mechanism.

[0176] Generally, the use of the harmonic patching method (sbrPatchingMode==0) is preferred for encoding music signals at very low bit rates, where the audio bandwidth of the core codec is greatly limited. This is particularly prominent when these signals contain significant harmonic structures. Conversely, the use of conventional SBR patching methods is preferred for speech and mixed signals because it provides better preservation of the temporal structure of speech.

[0177] To improve the performance of the MPEG-4 SBR transposer, a preprocessing step (bs_sbr_preprocessing== 1) can be initiated, which avoids introducing spectral discontinuities of the signal into the subsequent envelope adjuster. The operation of the tool is beneficial for signal types in which the coarse spectral envelope of the low-frequency band signal used for high-frequency reconstruction shows a large level of variation.

[0178] To improve the transient response of harmonic SBR patching (sbrPatchingMode== 0), the signal adaptive frequency domain specification 28 / 31 pages 30 CN 121214953 A can be applied.Oversampling (sbrOversamplingFlag==1). Since adaptive frequency domain oversampling increases the computational complexity of the transposer but only benefits frames containing transients, the use of this tool is controlled by the bitstream element, transmitted once per frame and per independent SBR channel.

[0179] Typical bitrate settings for HE-AACv2 with SBR enhancement (i.e., a harmonic transposer with eSBR tool enabled) are suggested to correspond to 20 kbp to 32 kbp for stereo audio content at sampling rates of 44.1 kHz or 48 kHz. The relative subjective quality gain of SBR enhancement increases toward lower bitrate boundaries, and a properly configured encoder allows this range to be extended to even lower bitrates. The bitrates provided above are merely suggestions and may be applicable to specific service requirements.

[0180] Decoders operating in the suggested enhanced SBR mode typically need to be able to switch between conventional SBR patching and enhanced SBR patching. Therefore, a delay as long as the duration of a core audio frame can be introduced depending on the decoder settings. Typically, the delays of conventional SBR repair and enhanced SBR repair will be similar.

[0181] It should be understood that the invention may be practiced in ways other than those specifically described herein, within the scope of the appended claims. Any element symbols contained in the following claims are for illustrative purposes only and should in no way be used to interpret or limit the claims.

[0182] Various aspects of the invention can be understood from the following enumerated example embodiments (EEE):

[0183] EEE 1. A method for performing high-frequency reconstruction of an audio signal, the method comprising:

[0184] receiving an encoded audio bitstream, the encoded audio bitstream comprising audio data representing a low-frequency band portion of the audio signal and high-frequency reconstruction metadata;

[0185] decoding the audio data to generate a decoded low-frequency band audio signal;

[0186] extracting the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata comprising operating parameters for a high-frequency reconstruction process, the operating parameters comprising a patch mode parameter located in a backward-compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates a spectral shift and a second value of the patch mode parameter indicates a harmonic transpose by frequency extension via a phase vocoder;

[0187] filtering the decoded low-frequency band audio signal to generate a filtered low-frequency band audio signal;

[0188] The high-frequency band portion of the audio signal is regenerated using the filtered low-frequency audio signal and the high-frequency reconstruction metadata, wherein if the patching mode parameter is the first value, the regeneration includes spectral shifting, and if the patching mode parameter is the second value, the regeneration includes harmonic transposition via frequency extension by a phase vocoder; and

[0189] The filtered low-frequency audio signal is combined with the regenerated high-frequency portion to form a wideband audio signal,

[0190] wherein the filtering, regeneration, and combination are performed as a post-processing operation with a delay of 3010 samples or less per audio channel, and wherein the spectral shift includes maintaining the ratio between the tone component and the noise-like component by adaptive inverse filtering.

[0191] EEE 2. The method according to EEE 1, wherein the encoded audio bitstream further includes a padding element having an identifier indicating the start of the padding element and padding data following the identifier, wherein the padding data includes the backward-compatible extension container.

[0192] EEE 3. The method according to EEE 2, wherein the identifier is a 3-bit unsigned integer having a value of 0×6 and transmitting the most significant bit first.

[0193] EEE 4. The method according to EEE 2 or EEE 3, wherein the padding data comprises an extended payload, the extended payload comprises spectral band copy extended data, and the extended payload is identified by a 4-bit unsigned integer having a value of “1101” or “1110” transmitted first, and optionally, Specification 29 / 31 pages 31 CN 121214953 A

[0194] wherein the spectral band copy extended data comprises:

[0195] optionally a spectral band copy header,

[0196] spectral band copy data, which is located after the header, and

[0197] a spectral band copy extended element, which is located after the spectral band copy data, and wherein the marker is included in the spectral band copy extended element.

[0198] EEE 5. The method according to any one of EEE 1 to 4, wherein the high-frequency reconstruction metadata comprises an envelope scaling factor, a noise floor scaling factor, time / frequency grid information, or a parameter indicating a cross frequency.

[0199] EEE 6. The method according to any one of EEE 1 to 5, wherein the backward compatible extension container further includes a flag indicating whether additional preprocessing is used to avoid shape discontinuities in the spectral envelope of the high-frequency band portion when the patch mode parameter is equal to the first value, wherein a first value of the flag enables the additional preprocessing and a second value of the flag disables the additional preprocessing.

[0200] EEE 7. The method according to EEE 6, wherein the additional preprocessing includes using linear prediction filter coefficients to compute a pre-gain curve.

[0201] EEE 8. The method according to any one of EEE 1 to 5, wherein the backward compatible extension container further includes a flag indicating whether signal adaptive frequency domain oversampling is applied when the patch mode parameter is equal to the second value, wherein a first value of the flag enables the signal adaptive frequency domain oversampling and a second value of the flag disables the signal adaptive frequency domain oversampling.

[0202] EEE 9. The method according to EEE 8, wherein the signal adaptive frequency domain oversampling is applied only to frames containing transients.

[0203] EEE 10. The method according to any one of the preceding EEEs, wherein the harmonic transpose by frequency extension via a phase vocoder is performed with an estimation complexity equal to or less than 4.5 million operations per second and 3,000 words of memory.

[0204] EEE 11. A non-transitory computer-readable medium containing instructions that, when executed by a processor, execute the method according to any one of EEEs 1 to 10.

[0205] EEE 12. A computer program product having instructions that, when executed by a computing device or system, cause the computing device or system to execute the method according to any one of EEEs 1 to 10.

[0206] EEE 13. An audio processing unit for performing high-frequency reconstruction of an audio signal, the audio processing unit comprising:

[0207] an input interface for receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-frequency band portion of the audio signal and high-frequency reconstruction metadata;

[0208] a core audio decoder for decoding the audio data to generate a decoded low-frequency band audio signal;

[0209] a deformatter for extracting the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata including operating parameters for the high-frequency reconstruction process, the operating parameters including a patch mode parameter located in a backward-compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates a spectral shift and a second value of the patch mode parameter indicates a harmonic transpose by frequency extension via a phase vocoder;

[0210] an analysis filter bank for filtering the decoded low-frequency band audio signal to generate a filtered low-frequency band audio signal;

[0211] A high-frequency regenerator for reconstructing a high-frequency portion of the audio signal using the filtered low-frequency audio signal and the high-frequency reconstruction metadata, wherein if the patching mode parameter is the first value, the reconstruction includes spectral shifting, and if the patching mode parameter is the second value, the reconstruction includes harmonic transposition via frequency extension by a phase vocoder; and

[0212] a synthesis filter bank for combining the filtered low-frequency audio signal with the regenerated high-frequency portion to form a broadband audio signal,

[0213] wherein the analysis filter bank, the high-frequency regenerator, and the synthesis filter bank are executed in a post-processor with a delay of 3010 samples or less per audio channel, and wherein the spectral shifting includes maintaining the ratio between the tone component and the noise-like component by adaptive inverse filtering.

[0214] EEE 14. According to EEEThe audio processing unit of 13, wherein the harmonic transpose via frequency extension by a phase vocoder is performed with an estimated complexity of 4.5 million operations per second and 3,000 words of memory, equal to or less than 4.5 million operations per second. Instruction Manual 31 / 31 Page 33 CN 121214953 A Figure 1 Figure 2 Figure 3 Figure 4 Instruction Manual Appendix 1 / 4 Page 34 CN 121214953 A Figure 5 Instruction Manual Appendix 2 / 4 Page 35 CN 121214953 A Figure 6 Instruction Manual Appendix 3 / 4 Page 36 CN 121214953 A Figure 7 Instruction Manual Appendix 4 / 4 Page 37 CN 121214953 A INTEGRATION OF HIGH FREQUENCY RECONSTRUCTION TECHNIQUES WITH REDUCED POST-PROCESSING DELAY Abstract The present disclosure is related to the integration of high frequency reconstruction techniques with reduced post-processing delay, and specifically discloses a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream and decoding the audio data to generate a decoded lowband audio signal. The method further includes extracting high frequency reconstruction metadata and filtering the decoded lowband audio signal with an analysis filterbank to generate a filtered lowband audio signal. The method also includes extracting a flag indicating whether either spectral translation or harmonictransposition is to be performed on the audio data and regenerating a highband portion of the audio signal using the filtered lowband audio signal and the high frequency reconstruction metadata in accordance with the flag. The high frequency regeneration is performed as a post-processing operation with a delay of 3010 samples per audio channel.

Claims

1. A method for performing high frequency reconstruction of an audio signal, the method comprising: receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low frequency band portion of the audio signal and high frequency reconstruction metadata, wherein the high frequency reconstruction metadata includes a noise scale factor; decoding the audio data to produce a decoded low frequency band audio signal; extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata including an operating parameter of a high frequency reconstruction process, the operating parameter including a patch mode parameter positioned in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates spectral panning and a second value of the patch mode parameter indicates harmonic transposition by phase vocoder frequency spreading; filtering the decoded low frequency band audio signal to produce a filtered low frequency band audio signal; and regenerating a high frequency band portion of the audio signal using the filtered low frequency band audio signal and the high frequency reconstruction metadata, wherein if the patch mode parameter is the first value, the regenerating includes spectral panning, and if the patch mode parameter is the second value, the regenerating includes harmonic transposition by phase vocoder frequency spreading, wherein the filtering and regenerating are performed as a post-processing operation with a delay of 3010 samples per audio channel, and wherein the spectral panning includes maintaining a ratio between tonal components and noise-like components by adaptive inverse filtering.

2. The method of claim 1, wherein the harmonic transposition by phase vocoder frequency spreading is performed with an estimated complexity equal to or below 45 million operations per second and equal to or below 3 kilowords of memory.

3. A non-transitory computer readable medium containing instructions that, when executed by a processor, perform the method of claim 1.

4. A computer program product stored on a non-transitory computer readable medium having instructions that, when executed by a computing device or system, cause the computing device or system to perform the method of claim 1.

5. An audio processing unit for performing high frequency reconstruction of an audio signal, the audio processing unit comprising: an input interface for receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low frequency band portion of the audio signal and high frequency reconstruction metadata, wherein the high frequency reconstruction metadata includes a noise scale factor; a core audio decoder for decoding the audio data to produce a decoded low frequency band audio signal; a deformatter for extracting the high frequency reconstruction metadata from the encoded audio bitstream, the high frequency reconstruction metadata including an operating parameter for a high frequency reconstruction process, the operating parameter including a patch mode parameter positioned in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates spectral panning and a second value of the patch mode parameter indicates harmonic transposition by phase vocoder frequency spreading; ​ an analysis filterbank for filtering the decoded low-band audio signal to produce a filtered low-band audio signal; and a high-frequency regenerator for regenerating a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein if the patching mode parameter is the first value, the regenerating includes spectral warping, and if the patching mode parameter is the second value, the regenerating includes harmonic transposition by phase vocoder frequency stretching, wherein the analysis filterbank and high-frequency regenerator are performed in a post-processor having a delay of 3010 samples per audio channel, and wherein the spectral warping comprises maintaining a ratio between tonal components and noise-like components by adaptive inverse filtering.