Integration of high frequency reconstruction techniques with reduced post-processing delay

By using analytical filter banks and high-frequency reconstruction metadata in audio signal processing, the problem of audio quality degradation in low cross-frequency music content caused by spectral band copying technology is solved, achieving more efficient audio signal reconstruction and encoding/decoding.

CN114242089BActive Publication Date: 2025-10-28DOLBY INTERNATIONAL AB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111585701.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-04-25
Filing Date
2019-04-25
Publication Date
2025-10-28
Estimated Expiration
2039-04-25

AI Technical Summary

Technical Problem

Existing spectrum band copying techniques are not effective when processing certain audio types, especially music with low cross frequencies, resulting in a decrease in audio quality.

Method used

By using an analytical filter bank to filter the decoded low-frequency audio signal and selectively performing spectral shifting or harmonic transpose based on high-frequency reconstruction metadata and tags, the high-frequency band portion of the audio signal is reconstructed, forming a broadband audio signal.

Benefits of technology

It improves the reconstruction quality of audio signals, especially in low cross-frequency music content, and enhances the compression efficiency and sound quality of audio codecs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114242089B_ABST
    Figure CN114242089B_ABST
Patent Text Reader

Abstract

This application relates to the integration of high-frequency reconstruction techniques with reduced post-processing latency, and specifically discloses a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream and decoding audio data to generate a decoded low-frequency band audio signal. The method further includes extracting high-frequency reconstruction metadata and using an analytical filter bank to filter the decoded low-frequency band audio signal to generate a filtered low-frequency band audio signal. The method also includes extracting a flag indicating whether to perform a spectral shift or harmonic transpose on the audio data and regenerating the high-frequency band portion of the audio signal according to the flag using the filtered low-frequency band audio signal and the high-frequency reconstruction metadata. The high-frequency regeneration is then subjected to a post-processing operation with a delay of 3010 samples per audio channel.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Information related to divisional application

[0002] This case is a divisional application. The parent application of this divisional application is the invention patent application filed on April 25, 2019, with application number 201980034811.4 and titled "Integration of High-Frequency Reconstruction Technology with Reduced Post-Processing Delay".

[0003] Cross-reference of related applications

[0004] This application claims priority to U.S. Provisional Patent Application No. 62 / 662,296, filed April 25, 2018, the entire contents of which are incorporated herein by reference. Technical Field

[0005] The embodiments relate to audio signal processing, and more specifically, to encoding, decoding, or transcoding audio bitstreams using control data that specifies a basic form or an enhanced form of performing high-frequency reconstruction (“HFR”) on audio data. Background Technology

[0006] A typical audio bitstream contains both audio data (e.g., encoded audio data) indicating one or more channels of audio content and metadata indicating at least one characteristic of the audio data or audio content. A well-known format used to generate encoded audio bitstreams is the MPEG-4 Advanced Audio Coding (AAC) format, as described in the MPEG standard ISO / IEC 14496-3:2009. In the MPEG-4 standard, AAC stands for "Advanced Audio Coding," and HE-AAC stands for "High-Efficiency Advanced Audio Coding."

[0007] The MPEG-4 AAC standard defines several audio profiles that determine which objects and encoding tools are present in a compatible encoder or decoder. Three of these audio profiles are (1) the AAC profile, (2) the HE-AAC profile, and (3) the HE-AAC v2 profile. The AAC profile contains the AAC Low Complexity (or “AAC-LC”) object type. The AAC-LC object is the counterpart to the MPEG-2 AAC Low Complexity profile, with some modifications, and does not include either the Spectral Band Replication (“SBR”) object type or the Parametric Stereo (“PS”) object type. The HE-AAC profile is a superset of the AAC profile and additionally contains the SBR object type. The HE-AAC v2 profile is a superset of the HE-AAC profile and additionally contains the PS object type.

[0008] The SBR object type contains a spectrum patching tool, an important high-frequency reconstruction (“HFR”) coding tool that can significantly improve the compression efficiency of perceptual audio codecs. SBR reconstructs the high-frequency components of the audio signal at the receiver side (e.g., in the decoder). Therefore, the encoder only needs to encode and transmit the low-frequency components to allow for much higher audio quality at lower data rates. SBR replicates the harmonic sequence previously truncated to reduce the data rate based on the limited available bandwidth signal obtained from the encoder and control data. The ratio between the tone components and noise-like components is maintained through adaptive inverse filtering and optionally the addition of noise and sine waves. In the MPEG-4 AAC standard, the SBR tool performs spectrum patching (also known as linear shifting or spectrum shifting), where several consecutive quadrature mirror filter (QMF) subbands are copied (or “patched”) from the transmitted low-frequency portion of the audio signal to the high-frequency portion of the audio signal (which is generated in the decoder).

[0009] Spectral patching or linear shifting may not be suitable for certain audio types (e.g., music content with relatively low cross frequencies). Therefore, techniques for improving spectral band replication are needed. Summary of the Invention

[0010] The first type of embodiment relates to a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream and decoding the audio data to generate a decoded low-frequency band audio signal. The method further includes extracting high-frequency reconstruction metadata and using an analytical filter bank to filter the decoded low-frequency band audio signal to generate a filtered low-frequency band audio signal. The method further includes extracting a marker indicating whether to perform a spectral shift or harmonic transpose on the audio data and using the filtered low-frequency band audio signal and the high-frequency reconstruction metadata to regenerate a high-frequency band portion of the audio signal according to the marker. Finally, the method includes combining the filtered low-frequency band audio signal and the regenerated high-frequency band portion to form a broadband audio signal.

[0011] The second type of embodiment relates to an audio decoder for decoding an encoded audio bitstream. The decoder includes: an input interface for receiving the encoded audio bitstream, wherein the encoded audio bitstream contains audio data representing a low-frequency band portion of an audio signal; and a core decoder for decoding the audio data to generate a decoded low-frequency band audio signal. The decoder also includes: a demultiplexer for extracting high-frequency reconstruction metadata from the encoded audio bitstream, wherein the high-frequency reconstruction metadata contains operating parameters for a high-frequency reconstruction process that linearly shifts several consecutive sub-bands from the low-frequency band portion of the audio signal to the high-frequency band portion of the audio signal; and an analysis filter bank for filtering the decoded low-frequency band audio signal to generate a filtered low-frequency band audio signal. The decoder further includes: a demultiplexer for extracting from the encoded audio bitstream a flag indicating whether a linear shift or harmonic transpose is performed on the audio data; and a high-frequency regenerator for regenerating the high-frequency band portion of the audio signal using the filtered low-frequency band audio signal and the high-frequency reconstruction metadata according to the flag. Finally, the decoder includes a synthesis filter bank for combining the filtered low-frequency audio signal and the regenerated high-frequency portion to form a broadband audio signal.

[0012] Other embodiments involve encoding and transcoding audio bitstreams containing metadata that identifies whether enhanced spectrum band replication (eSBR) processing has been performed. Attached Figure Description

[0013] Figure 1 This is a block diagram of an embodiment of a system that can be configured to perform an embodiment of the inventive method.

[0014] Figure 2 This is a block diagram of an encoder, which is an embodiment of the inventive audio processing unit.

[0015] Figure 3 It is a block diagram of a system that includes a decoder (which is an embodiment of the inventive audio processing unit) and optionally also includes a post-processor coupled to the decoder.

[0016] Figure 4 This is a block diagram of a decoder, which is an embodiment of the inventive audio processing unit.

[0017] Figure 5 This is a block diagram of a decoder, which is another embodiment of the inventive audio processing unit.

[0018] Figure 6 This is a block diagram of another embodiment of the audio processing unit of the invention.

[0019] Figure 7It is a block diagram of the MPEG-4 AAC bitstream, which includes its division into several segments.

[0020] Symbols and terms

[0021] In this invention (included in the claims), the expression “to” perform an operation on a signal or data (e.g., filtering, scaling, transforming, or applying gain to a signal or data) is used broadly to mean performing an operation directly on the signal or data or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or preprocessing before the operation is performed on it).

[0022] In this invention (as defined in the claims), the term "audio processing unit" or "audio processor" is used broadly to refer to a system, apparatus, or device configured to process audio data. Examples of audio processing units include (but are not limited to) encoders, transcoders, decoders, codecs, preprocessing systems, post-processing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools). Almost all consumer electronics products (e.g., mobile phones, televisions, laptops, and tablets) contain audio processing units or audio processors.

[0023] In this invention (as defined in the claims), the terms "coupled" or "via coupling" are used broadly to mean a direct or indirect connection. Thus, if a first device is coupled to a second device, the connection can be either a direct connection or an indirect connection via other devices and connections. Furthermore, components integrated into or with other components are also coupled to each other. Detailed Implementation

[0024] The MPEG-4 AAC standard anticipates that encoded MPEG-4 AAC bitstreams contain metadata indicating each type of High Frequency Reconstruction (“HFR”) processing applied (if to be applied) by the decoder to decode the audio content of the bitstream, and / or controlling this HFR processing, and / or indicating at least one characteristic or parameter of at least one HFR tool used to decode the audio content of the bitstream. In this document, we use the term “SBR metadata” to refer to this type of metadata used in conjunction with Spectral Band Replication (“SBR”), as described or mentioned in the MPEG-4 AAC standard. Those skilled in the art will understand that SBR is a form of HFR.

[0025] SBR is preferably used as a dual-rate system, where the basic codec operates at half the original sampling rate, while the SBR operates at the original sampling rate. Despite the higher sampling rate, the SBR encoder works in parallel with the basic core codec. Although SBR is primarily post-processing in the decoder, it extracts important parameters from the encoder to ensure the most accurate high-frequency reconstruction in the decoder. The encoder estimates the spectral envelope of the SBR range, suitable for the time and frequency range / resolution of the current input signal segment. The spectral envelope is estimated through complex QMF analysis and subsequent energy calculation. The time and frequency resolution of the spectral envelope can be chosen with considerable freedom to ensure the most suitable time-frequency resolution for a given input segment. Envelope estimation needs to take into account that transients from the original source, primarily located in the high-frequency region (e.g., the high hat), will appear to a lesser extent in the high-frequency band generated by the SBR before envelope adjustment, because the high-frequency band in the decoder is based on the low-frequency band where transients are much less pronounced than in the high-frequency band. This aspect places different requirements on the time-frequency resolution of the spectral envelope data compared to general spectral envelope estimation used in other audio coding algorithms.

[0026] In addition to the spectral envelope, several additional parameters representing the spectral characteristics of the input signal in different time and frequency regions are extracted. Since the encoder naturally has access to the original signal and information about how the SBR unit in the decoder will generate the high-frequency band, given a specific set of control parameters, the system can handle cases where the low-frequency band constitutes a strong harmonic series and the regenerated high-frequency band mainly consists of random signal components, as well as cases where strong tone components exist in the original high-frequency band but have no corresponding counterpart in the low-frequency band (the high-frequency band region is based on this). Furthermore, the SBR encoder works closely with the basic core codec to evaluate which frequency range should be covered by the SBR at a given time. For stereo signals, the SBR data is efficiently encoded before transmission using entropy coding and the channel dependence of the control data.

[0027] Typically, the algorithm for extracting control parameters needs to be carefully tuned based on the base codec to provide a given bitrate and sampling rate. This is due to the fact that lower bitrates usually mean a larger SBR range than higher bitrates, and different sampling rates correspond to different temporal resolutions of SBR frames.

[0028] An SBR decoder typically comprises several distinct parts. These include a bitstream decoding module, a high-frequency reconstruction (HFR) module, an additional high-frequency component module, and an envelope adjuster module. The system is based on either a complex-valued QMF filter bank (for high-quality SBR) or a real-valued QMF filter bank (for low-power SBR). Embodiments of this invention are applicable to both high-quality and low-power SBR. In the bitstream extraction module, control data is read from and decoded from the bitstream. Before reading the envelope data from the bitstream, the time-frequency grid of the current frame is obtained. The basic core decoder decodes the audio signal of the current frame (albeit at a lower sampling rate) to generate time-domain audio samples. The resulting frame from the audio data is used by the HFR module for high-frequency reconstruction. Next, a QMF filter bank is used to analyze the decoded low-frequency band signal. Subsequently, high-frequency reconstruction and envelope adjustment are performed on the sub-band samples of the QMF filter bank. High frequencies are reconstructed from the low-frequency band in a flexible manner based on given control parameters. Furthermore, the reconstructed high-frequency band is adaptively filtered based on the sub-band channels according to the control data to ensure appropriate spectral characteristics for a given time / frequency region.

[0029] The top layer of an MPEG-4 AAC bitstream is a sequence of data blocks (“raw_data_block” elements), each of which is a data segment (referred to herein as a “block”) containing audio data (typically within a time period of 1024 or 960 samples) and related and / or other data. In this document, we use the term “block” to denote a segment of the MPEG-4 AAC bitstream that includes audio data (and corresponding metadata and optionally other related data), identifying or indicating one (but not more than one) “raw_data_block” element.

[0030] Each block of an MPEG-4 AAC bitstream can contain several syntax elements (each of which is also specified as a data segment in the bitstream). Seven types of these syntax elements are defined in the MPEG-4 AAC standard. Each syntax element is identified by a distinct value of the data element "id_syn_ele". Instances of syntax elements include "single_channel_element()", "channel_pair_element()", and "fill_element()". A single-channel element is a container for audio data containing a single audio channel (mono audio signal). A channel-pair element contains audio data for two audio channels (i.e., stereo audio signal).

[0031] A padding element is an information container that contains an identifier (such as the value of the element "id_syn_ele" mentioned above) followed by data (called "padding data"). Padding elements have historically been used to adjust the instantaneous bit rate of a bit stream to be transmitted over a constant-rate channel. A constant data rate can be achieved by adding an appropriate amount of padding data to each block.

[0032] According to embodiments of the invention, the padding data may include one or more extended payloads that extend the data type (e.g., metadata) that can be transmitted in the bitstream. A decoder receiving a bitstream with padding data containing the new data type may optionally be used by the receiving device (e.g., a decoder) to extend the functionality of the device. Therefore, those skilled in the art will understand that the padding element is a special type of data structure and differs from data structures typically used for transmitting audio data (e.g., audio payloads containing channel data).

[0033] In some embodiments of the invention, the identifier used to identify padding elements may consist of a 3-bit unsigned integer (“uimsbf”) with a value of 0×6, where the most significant bit is transmitted first. Within a block, several instances of the same type of syntax element (e.g., several padding elements) may appear.

[0034] Another standard used for encoding audio bitstreams is the MPEG Unified Speech and Audio Coding (USAC) standard (ISO / IEC 23003-3:2012). The MPEG USAC standard describes the use of spectrum band copying processing (including the SBR processing described in the MPEG-4 AAC standard and other enhancements to spectrum band copying processing) to encode and decode audio content. This processing applies an extended and enhanced version of the spectrum band copying toolkit described in the MPEG-4 AAC standard (sometimes referred to herein as the “enhanced SBR tool” or “eSBR tool”). Therefore, eSBR (as defined in the USAC standard) is an improvement upon SBR (as defined in the MPEG-4 AAC standard).

[0035] In this paper, we use the term "enhanced SBR processing" (or "eSBR processing") to refer to spectral band copying processing using at least one eSBR tool not described or mentioned in the MPEG-4 AAC standard (such as at least one eSBR tool described or mentioned in the MPEG USAC standard). Examples of these eSBR tools are harmonic transpose and QMF patching additional preprocessing or "pre-flattening".

[0036] An integer-order T harmonic transposer maps a sine wave with frequency ω to a sine wave with frequency Tω, while preserving the signal duration. Typically, three orders T = 2, 3, 4 are used sequentially to produce each portion of the desired output frequency range using the minimum possible transpose order. If an output with a transpose range higher than 4 is required, it can be generated by frequency shifting. The fundamental frequency time domain, as close to the critical sample as possible, is used for processing to minimize computational complexity.

[0037] Harmonic transposers can be based on QMF or DFT. When using a QMF-based harmonic transposer, a modified phase vocoder structure is used in the QMF domain to fully implement bandwidth expansion of the core encoder's time-domain signal to perform sampling and subsequent time extension for each QMF subband. Transposition using several transpose factors (e.g., T = 2, 3, 4) is implemented in the common QMF analysis / synthesis transform stage. Since QMF-based harmonic transposers do not have the characteristic of adaptive frequency domain oversampling, the corresponding flag (sbrOversamplingFlag[ch]) in the bit stream can be ignored.

[0038] When using a DFT-based harmonic transposer, factor 3 and 4 transposers (3rd and 4th order transposers) are preferably integrated into a factor 2 transposer (2nd order transposer) via interpolation to reduce complexity. For each frame (corresponding to the coreCoderFrameLength core encoder sample), the nominal “full-size” transform size of the transposer is first determined by the signal adaptive frequency domain oversampling flag (sbrOversamplingFlag[ch]) in the bitstream.

[0039] When sbrPatchingMode == 1 indicates that a linear transpose will be used to generate the high-frequency band, an additional step can be introduced to avoid discontinuities in the shape of the spectral envelope of the high-frequency signal input to the subsequent envelope adjuster. This improves the operation of the subsequent envelope adjustment stage, resulting in a perceived more stable high-frequency band signal. The operation of this additional preprocessing is beneficial for signal types where the coarse spectral envelope of the low-frequency band signal used for high-frequency reconstruction exhibits a large level of variation. However, the value of the bitstream element can be determined in the encoder by applying any kind of signal dependency classification. Preferably, the additional preprocessing is initiated by a 1-bit bitstream element, bs_sbr_preprocessing. Additional processing is enabled when bs_sbr_preprocessing is set to 1. Additional preprocessing is disabled when bs_sbr_preprocessing is set to 0. The additional processing preferably utilizes a pre-gain curve used by the high-frequency generator to scale each patched low-frequency band X. Low For example, the pregain curve can be calculated using the following equation:

[0040] preGain(k) = 10 (meanNrg-lowEnvSlope(k)) / 20 , 0≤k <k0

[0041] Where k0 is the first QMF sub-band in the main band table and lowEnvSlope is calculated using a function (e.g., polyfit()) that computes the coefficients of the best-fit polynomial (in the least-squares sense). For example, it could be done using a cubic polynomial.

[0042] polyfit(3,k0,x_lowband,lowEnv,lowEnvSlope);

[0043] And among them

[0044]

[0045] Where x_lowband(k) = [0...k0-1], numTimeSlot is the number of SBR envelope time slots existing within the frame, and RATE is a constant (e.g., 2) indicating the number of QMF sub-band samples in each time slot. These are linear predictive filter coefficients (obtainable from the covariance method) and where

[0046]

[0047] The bitstream generated according to the MPEG USAC standard (sometimes referred to herein as the “USAC bitstream”) contains encoded audio content and typically contains metadata indicating each type of spectrum band copying process applied by the decoder to decode the audio content of the USAC bitstream and / or metadata controlling this spectrum band copying process and / or indicating at least one feature or parameter of at least one SBR tool and / or eSBR tool used to decode the audio content of the USAC bitstream.

[0048] In this document, we use the term "enhanced SBR metadata" (or "eSBR metadata") to refer to metadata that indicates, and / or controls, each type of spectrum band copying process applied by the decoder to decode audio content of an encoded audio bitstream (e.g., a USAC bitstream), and / or indicates at least one characteristic or parameter of at least one SBR tool and / or eSBR tool used to decode this audio content but not described or mentioned in the MPEG-4 AAC standard. An instance of eSBR metadata is metadata (indicating or used to control spectrum band copying processes) described or mentioned in the MPEG USAC standard but not described or mentioned in the MPEG-4 AAC standard. Therefore, eSBR metadata in this document is not metadata of SBR metadata, and SBR metadata in this document is not metadata of eSBR metadata.

[0049] The USAC bitstream may contain both SBR metadata and eSBR metadata. More specifically, the USAC bitstream may contain eSBR metadata that controls the eSBR processing performed by the decoder and SBR metadata that controls the SBR processing performed by the decoder. According to a typical embodiment of the invention, the eSBR metadata (e.g., eSBR-specific configuration data) is contained (according to the invention) in the MPEG-4 AAC bitstream (e.g., in the sbr_extension() container at the end of the SBR payload).

[0050] During the decoding of an encoded bitstream using the eSBR toolset (including at least one eSBR tool), the decoder performs eSBR processing to regenerate the high-frequency band of the audio signal based on a copy of the harmonic sequence truncated during encoding. This eSBR processing typically adjusts the spectral envelope of the resulting high-frequency band and applies inverse filtering, and adds noise and sinusoidal components to regenerate the spectral characteristics of the original audio signal.

[0051] According to a typical embodiment of the invention, eSBR metadata (e.g., a small number of control bits of eSBR metadata) is contained in one or more metadata segments of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream), which also contains encoded audio data in other segments (audio data segments). Typically, at least one of these metadata segments in each block of the bitstream is (or contains) a padding element (containing an identifier indicating the start of the padding element), and the eSBR metadata is contained in the padding element following the identifier.

[0052] Figure 1 This is a block diagram of an exemplary audio processing chain (audio data processing system), in which one or more elements of the system can be configured according to embodiments of the invention. The system includes the following elements coupled together as shown: encoder 1, transmission subsystem 2, decoder 3, and post-processing unit 4. In variations of the system shown, one or more elements are omitted, or additional audio data processing units are included.

[0053] In some implementations, encoder 1 (which optionally includes a preprocessing unit) is configured to accept PCM (temporal domain) samples including audio content as input and output an encoded audio bitstream (in a format conforming to the MPEG-4 AAC standard) indicating the audio content. The data in the bitstream indicating the audio content is sometimes referred to herein as "audio data" or "encoded audio data." If the encoder is configured according to a typical embodiment of the invention, the audio bitstream output from the encoder includes eSBR metadata (and typically also other metadata) as well as the audio data.

[0054] One or more encoded audio bitstreams output from encoder 1 can be asserted to encoded audio transmission subsystem 2. Subsystem 2 is configured to store and / or transmit each encoded bitstream output from encoder 1. The encoded audio bitstreams output from encoder 1 can be stored by subsystem 2 (e.g., in the form of DVD or Blu-ray disc), or transmitted by subsystem 2 (which may implement a transmission link or network), or can be both stored and transmitted by subsystem 2.

[0055] Decoder 3 is configured to decode the encoded MPEG-4 AAC audio bitstream (generated by encoder 1) received via subsystem 2. In some embodiments, decoder 3 is configured to extract eSBR metadata from each block of the bitstream and decode the bitstream (including performing eSBR processing using the extracted eSBR metadata) to produce decoded audio data (e.g., a decoded PCM audio sample stream). In some embodiments, decoder 3 is configured to extract SBR metadata from the bitstream (but ignore the eSBR metadata contained in the bitstream) and decode the bitstream (including performing SBR processing using the extracted SBR metadata) to produce decoded audio data (e.g., a decoded PCM audio sample stream). Typically, decoder 3 includes a buffer that stores (e.g., in a non-transitory manner) segments of the encoded audio bitstream received from subsystem 2.

[0056] Figure 1 The post-processing unit 4 is configured to receive the decoded audio data stream (e.g., a decoded PCM audio sample) from the decoder 3 and perform post-processing on it. The post-processing unit may also be configured to render the post-processed audio content (or the decoded audio received from the decoder 3) for playback by one or more speakers.

[0057] Figure 2 This is a block diagram of encoder 100, an embodiment of the inventive audio processing unit. Any component or element of encoder 100 may be implemented in hardware, software, or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuits). Encoder 100 includes encoder 105, filler / formatter stage 107, metadata generation stage 106, and buffer memory 109 connected as shown. Typically, encoder 100 also includes other processing elements (not shown). Encoder 100 is configured to convert an input audio bitstream into an encoded output MPEG-4 AAC bitstream.

[0058] Metadata generator 106 is coupled and configured to generate metadata (including eSBR metadata and SBR metadata) (and / or pass it to stage 107) to be included in the encoded bitstream output from encoder 100 by stage 107.

[0059] Encoder 105 is coupled and configured to encode input audio data (e.g., by performing compression on it) and asserts the resulting encoded audio to stage 107 to be included in the encoded bitstream output from stage 107.

[0060] Stage 107 is configured to multiplex encoded audio from encoder 105 and metadata (including eSBR metadata and SBR metadata) from generator 106 to produce an encoded bitstream output from stage 107, preferably such that the encoded bitstream has a format specified by an embodiment of the present invention.

[0061] The buffer memory 109 is configured to store (e.g., in a non-transitory manner) at least one block of the encoded audio bitstream output from stage 107, and then the block sequence of the encoded audio bitstream is asserted from the buffer memory 109 as the output from encoder 100 to the transmission system.

[0062] Figure 3 This is a block diagram of a system including decoder 200 (which is an embodiment of the inventive audio processing unit) and optionally also including post-processor 300 coupled to decoder 200. Any component or element of decoder 200 and post-processor 300 may be implemented in hardware, software, or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuits). Decoder 200 includes a buffer memory 201, a bitstream payload deformatter (parser) 205, an audio decoding subsystem 202 (sometimes referred to as the “core” decoding stage or “core decoding subsystem”), an eSBR processing stage 203, and a control bit generation stage 204, all connected as shown. Typically, decoder 200 also includes other processing elements (not shown).

[0063] Buffer memory (buffer) 201 stores (e.g., in a non-transitory manner) at least one block of the encoded MPEG-4 AAC audio bitstream received by decoder 200. During the operation of decoder 200, the block sequence of the bitstream is asserted from buffer 201 to deformatter 205.

[0064] exist Figure 3 Examples (or those to be described) Figure 4 In a variant of the embodiment, the APU (which is not a decoder) (e.g.) Figure 6 The APU 500 includes a buffer memory (e.g., a buffer memory identical to buffer 201), the storage of which (e.g., in a non-transitory manner) is provided by Figure 3 or Figure 4 The buffer 201 receives at least one block of the same type of encoded audio bitstream (e.g., an MPEG-4 AAC audio bitstream) (i.e., an encoded audio bitstream containing eSBR metadata).

[0065] Refer again Figure 3 The deformatter 205 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (containing quantization envelope data) and eSBR metadata (and typically other metadata) to assert at least the eSBR metadata and SBR metadata to the eSBR processing stage 203 and typically also to the decoding subsystem 202 (and optionally to the control bit generator 204). The deformatter 205 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.

[0066] Figure 3 The system also optionally includes a post-processor 300. The post-processor 300 includes a buffer memory (buffer) 301 and other processing elements (not shown), including at least one processing element coupled to the buffer 301. The buffer 301 stores (e.g., in a non-transitory manner) at least one block (or frame) of decoded audio data received by the post-processor 300 from the decoder 200. The processing elements of the post-processor 300 are coupled and configured to receive and use metadata output from the decoding subsystem 202 (and / or deformatter 205) and / or control bits output from stage 204 of the decoder 200 to adapt the processing of the block (or frame) sequence of decoded audio output from the buffer 301.

[0067] Audio decoding subsystem 202 of decoder 200 is configured to decode audio data extracted by parser 205 (this decoding may be referred to as the "core" decoding operation) to produce decoded audio data and assert the decoded audio data to eSBR processing stage 203. Decoding is performed in the frequency domain and typically includes inverse quantization followed by spectral processing. Typically, the final processing stage in subsystem 202 applies a frequency-to-time domain transform to the decoded frequency-domain audio data, such that the output of the subsystem is time-domain decoded audio data. Stage 203 is configured to apply SBR tools and eSBR tools indicated by eSBR metadata and eSBR (extracted by parser 205) to the decoded audio data (i.e., perform SBR and eSBR processing on the output of decoding subsystem 202 using SBR and eSBR metadata) to produce fully decoded audio data from the output of decoder 200 (e.g., to post-processor 300). Typically, decoder 200 includes memory (accessible by subsystem 202 and stage 203) storing deformatted audio data and metadata output from deformatter 205, and stage 203 is configured to access audio data and metadata (including SBR metadata and eSBR metadata) as needed during SBR and eSBR processing. SBR and eSBR processing in stage 203 can be considered as post-processing of the output of core decoding subsystem 202. Decoder 200 also optionally includes a final upmixing subsystem (which can apply parametric stereo (“PS”) tools as defined in the MPEG-4 AAC standard using PS metadata extracted by deformatter 205 and / or control bits generated in subsystem 204), coupled and configured to perform upmixing on the output of stage 203 to produce fully decoded upmixed audio output from decoder 200. Alternatively, the postprocessor 300 is configured to perform upmixing on the output of the decoder 200 (e.g., using PS metadata extracted by the deformatter 205 and / or control bits generated in the subsystem 204).

[0068] In response to metadata extracted by deformatter 205, control bit generator 204 can generate control data, which can be used within decoder 200 (e.g., for the final upmixing subsystem) and / or asserted as the output of decoder 200 (e.g., to post-processor 300 for post-processing). In response to metadata extracted from the input bitstream (and optionally also in response to the control data), stage 204 can generate control bits (and assert the control bits to post-processor 300) to indicate that the decoded audio data output from eSBR processing stage 203 should undergo a specific type of post-processing. In some embodiments, decoder 200 is configured to assert metadata extracted from the input bitstream by deformatter 205 to post-processor 300, and post-processor 300 is configured to perform post-processing on the decoded audio data output from decoder 200 using the metadata.

[0069] Figure 4 This is a block diagram of an audio processing unit (“APU”) 210, which is another embodiment of the inventive audio processing unit. APU 210 is a conventional decoder not configured to perform eSBR processing. Any component or element of APU 210 may be implemented in hardware, software, or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuits). APU 210 includes a buffer memory 201, a bitstream payload deformatter (parser) 215, an audio decoding subsystem 202 (sometimes referred to as the “core” decoding stage or “core” decoding subsystem), and an SBR processing stage 213, connected as shown. Typically, APU 210 also includes other processing elements (not shown). APU 210 may represent, for example, an audio encoder, decoder, or transcoder.

[0070] Components 201 and 202 of APU 210 are the same as ( Figure 3 The same numbered elements of the decoder 200 will not be repeated in their above description. In the operation of the APU 210, the block sequence of the encoded audio bitstream (MPEG-4 AAC bitstream) received by the APU 210 is asserted from the buffer 201 to the deformatter 215.

[0071] Deformatter 215 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (including quantization envelope data) from it, and typically also extracts other metadata from it, but ignores eSBR metadata that may be included in the bitstream according to any embodiment of the invention. Deformatter 215 is configured to assert at least SBR metadata to SBR processing stage 213. Deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to decoding subsystem (decoding stage) 202.

[0072] The audio decoding subsystem 202 of decoder 200 is configured to decode the audio data extracted by deformatter 215 (this decoding may be referred to as the "core" decoding operation) to produce decoded audio data and assert the decoded audio data to SBR processing stage 213. Decoding is performed in the frequency domain. Typically, the final processing stage in subsystem 202 applies a frequency-to-time domain transformation to the decoded frequency-domain audio data, such that the output of the subsystem is time-domain decoded audio data. Stage 213 is configured to apply an SBR tool (but not an eSBR tool) indicated by SBR metadata (extracted by deformatter 215) to the decoded audio data (i.e., to perform SBR processing on the output of decoding subsystem 202 using SBR metadata) to produce fully decoded audio data output from APU 210 (e.g., to post-processor 300). Typically, APU 210 includes memory (accessible by subsystem 202 and stage 213) storing deformatted audio data and metadata output from deformatter 215, and stage 213 is configured to access audio data and metadata (including SBR metadata) as needed during SBR processing. SBR processing in stage 213 can be viewed as post-processing of the output of core decoding subsystem 202. APU 210 also optionally includes a final upmixing subsystem (which can apply the parametric stereo “PS” tool defined in the MPEG-4 AAC standard using PS metadata extracted by deformatter 215), coupled and configured to perform upmixing on the output of stage 213 to produce fully decoded upmixed audio output from APU 210. Alternatively, a postprocessor is configured to perform upmixing on the output of APU 210 (e.g., using PS metadata extracted by deformatter 215 and / or control bits generated in APU 210).

[0073] Various embodiments of the encoder 100, decoder 200, and APU 210 are configured to perform different embodiments of the inventive method.

[0074] According to some embodiments, eSBR metadata (e.g., a small number of control bits of eSBR metadata) is included in the encoded audio bitstream (e.g., an MPEG-4 AAC bitstream), allowing a conventional decoder (not configured to parse eSBR metadata or use any eSBR tools associated with eSBR metadata) to ignore the eSBR metadata but still decode the bitstream as much as possible without using the eSBR metadata or any eSBR tools associated with it, typically without a significant loss of decoded audio quality. However, an eSBR decoder (configured to parse the bitstream to identify eSBR metadata and use at least one eSBR tool in response to the eSBR metadata) will benefit from using at least one of these eSBR tools. Therefore, embodiments of the present invention provide a method for efficiently transmitting enhanced spectrum band replication (eSBR) control data or metadata in a backward-compatible manner.

[0075] Typically, the eSBR metadata in the bitstream indicates one or more of the following eSBR tools (e.g., indicating at least one of their characteristics or parameters) (the eSBR tools are described in the MPEG USAC standard and may or may not be applied by the encoder during the generation of the bitstream):

[0076] ●Harmonic transposition; and

[0077] ●QMF patching additional preprocessing (pre-flattening).

[0078] For example, eSBR metadata contained in the bitstream can indicate the values ​​of parameters (as described in the MPEG USAC standard and in this invention): sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch], and bs_sbr_preprocessing.

[0079] In this paper, the symbol X[ch] (where X is a parameter) indicates that the parameter is related to the channel (“ch”) of the audio content of the encoded bitstream to be decoded. For simplicity, we sometimes omit the expression [ch] and assume that the relevant parameter is related to the channel of the audio content.

[0080] In this paper, the symbol X[ch][env] (where X is a parameter) indicates that the parameter is related to the SBR envelope (“env”) of the channel (“ch”) of the audio content of the encoded bitstream to be decoded. For simplicity, we sometimes omit the expressions [env] and [ch], and assume that the relevant parameter is related to the SBR envelope of the channel of the audio content.

[0081] During the decoding of the encoded bitstream, harmonic transposition (for each channel "ch" of the audio content indicated by the bitstream) is performed during the eSBR processing stage of the decoder and is controlled by the following eSBR metadata parameters: sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBinsFlag[ch], and sbrPitchInBins[ch].

[0082] The value “sbrPatchingMode[ch]” indicates the transpose type used in eSBR: sbrPatchingMode[ch] = 1 indicates linear transpose patching as described in section 4.6.18 of the MPEG-4 AAC standard (used with high-quality SBR or low-power SBR); sbrPatchingMode[ch] = 0 indicates harmonic SBR patching as described in section 7.5.3 or 7.5.4 of the MPEG USAC standard.

[0083] The value “sbrOversamplingFlag[ch]” indicates that adaptive frequency domain oversampling in the eSBR is used in combination with DFT-based harmonic SBR patching as described in section 7.5.3 of the MPEG USAC standard. This flag controls the size of the DFT used in the transposer: 1 indicates that adaptive frequency domain oversampling is enabled as described in section 7.5.3.1 of the MPEG USAC standard; 0 indicates that adaptive frequency domain oversampling is disabled as described in section 7.5.3.1 of the MPEG USAC standard.

[0084] The value “sbrPitchInBinsFlag[ch]” controls the interpretation of the sbrPitchInBins[ch] parameter: 1 indicates that the value of sbrPitchInBins[ch] is valid and greater than 0; 0 indicates that the value of sbrPitchInBins[ch] is set to 0.

[0085] The value “sbrPitchInBins[ch]” controls the addition of cross-product terms in the SBR harmonic transposer. The value sbrPitchinBins[ch] is an integer value in the range [0, 127] and represents the distance measured in the frequency grid of a 1536-line DFT applied to the sampling frequency of the core encoder.

[0086] If the MPEG-4 AAC bitstream indicates an uncoupled SBR channel pair (rather than a single SBR channel), then the bitstream indicates two instances of the above syntax (for harmonic or non-harmonic transpose), with one instance sbr_channel_pair_element() for each channel.

[0087] Harmonic transposition by eSBR tools typically improves the quality of decoded music signals at relatively low crossover frequencies. Non-harmonic transposition (i.e., conventional spectral patching) typically improves speech signals. Therefore, the starting point for determining which type of transposition is preferred for encoding specific audio content is to select the transposition method based on speech / music detection, where harmonic transposition is used for music content and spectral patching is used for tempo content.

[0088] The pre-flattening performed during eSBR processing is controlled by the value of a 1-bit eSBR metadata parameter called “bs_sbr_preprocessing”, which, in a sense, determines whether pre-flattening is performed or not based on the value of this single bit. When using the SBR QMF patching algorithm described in section 4.6.18.6.3 of the MPEG-4 AAC standard, a pre-flattening step (when indicated by the “bs_sbr_preprocessing” parameter) may be performed to attempt to avoid shape discontinuities in the spectral envelope of the high-frequency signal input to the subsequent envelope adjuster (another stage of eSBR processing performed by the envelope adjuster). Pre-flattening generally improves the operation of the subsequent envelope adjustment stage, resulting in a perceived more stable high-frequency band signal.

[0089] According to some embodiments of the invention, the total bit rate requirement included in the MPEG-4 AAC bitstream eSBR metadata indicating the aforementioned eSBR tools (harmonic transpose and pre-flattening) is expected to be approximately several hundred bits per second, since only differential control data required to perform eSBR processing is transmitted. Conventional decoders can ignore this information because it is included in a backward-compatible manner (as explained later). Therefore, the adverse effects on bit rate associated with including eSBR metadata are negligible due to several reasons, including:

[0090] ● Bit rate loss (attributed to the inclusion of eSBR metadata) accounts for a very small percentage of the total bit rate because only differential control data required to perform eSBR processing is transmitted (and non-SBR control data is cascaded); and

[0091] ● Tuning of SBR-related control information typically does not depend on transpose details. Examples where control data depends on transpose operation will be discussed later in this application.

[0092] Therefore, embodiments of the present invention provide a method for efficiently transmitting enhanced spectrum band replication (eSBR) control data or metadata in a backward-compatible manner. This efficient transmission of eSBR control data reduces memory requirements in decoders, encoders, and transcoders employing aspects of the present invention, without significantly adversely affecting bit rate. Furthermore, it reduces the complexity and processing requirements associated with performing eSBR according to embodiments of the present invention, since SBR data only needs to be processed once and is not cascaded, as is the case when eSBR is treated as a completely independent object type in MPEG-4 AAC rather than being integrated into the MPEG-4 AAC codec in a backward-compatible manner.

[0093] Next, refer to Figure 7 We describe the elements of a block (“raw_data_block”) of an MPEG-4 AAC bitstream (containing eSBR metadata) according to some embodiments of the present invention. Figure 7 This is a graph of the MPEG-4 AAC bitstream (“raw_data_block”), which shows some segments of the MPEG-4 AAC bitstream.

[0094] A block of an MPEG-4 AAC bitstream may contain at least one "single_channel_element()" (e.g. Figure 7 The single-channel element shown in the image) and / or at least one "channel_pair_element()" Figure 7 (Not explicitly shown, but it may exist), it contains the audio data of the audio program. A block may also contain several "fill_element" (e.g., ... Figure 7 The padding element 1 and / or padding element 2 contain program-related data (e.g., metadata). Each "single_channel_element()" contains an identifier indicating the start of a single-channel element (e.g., ...). Figure 7 The "ID1" field can contain audio data indicating different channels in a multi-channel audio program. Each "channel_pair_element()" contains an identifier indicating the start of the channel pair element. Figure 7 (Not shown in the text), and may contain audio data indicating two channels of the program.

[0095] The fill_element (referred to as the fill element in this document) of an MPEG-4 AAC bitstream contains an identifier indicating the start of the fill element. Figure 7 The identifier ID2 is followed by padding data. ID2 can consist of a 3-bit unsigned integer (“uimsbf”) with a value of 0×6, where the most significant bit is transmitted first. The padding data can contain the extension_payload() element (sometimes referred to herein as the extended payload) whose syntax is shown in Table 4.57 of the MPEG-4 AAC standard. Several types of extended payloads exist and are identified by the “extension_type” parameter, which is a 4-bit unsigned integer (“uimsbf”) where the most significant bit is transmitted first.

[0096] The padding data (e.g., its extended payload) may contain headers or identifiers (e.g., indicating the padding data section, which in turn indicates the SBR object) that define the padding data. Figure 7The header (i.e., the header initialization of the "SBR object" type, referred to as sbr_extension_data() in the MPEG-4 AAC standard) is used. For example, the value of "1101" or "1110" in the extension_type field of the header is used to identify the Spectral Band Replication (SBR) extended payload, where the identifier "1101" identifies an extended payload with SBR data and "1110" identifies an extended payload containing SBR data with cyclic redundancy check (CRC) to verify the correctness of the SBR data.

[0097] When the header (e.g., the extension_type field) initializes the SBR object type, the SBR metadata (sometimes referred to in this document as "spectral band copy data," and referred to as sbr_data() in the MPEG-4 AAC standard) follows the header, and at least one spectral band copy extension element (e.g. Figure 7 The padding element 1, designated as the "SBR extension element," may be followed by SBR metadata. This spectral band copy extension element (a segment of the bitstream) is referred to as the "sbr_extension()" container in the MPEG-4 AAC standard. The spectral band copy extension element optionally includes a header (e.g., ...). Figure 7 The "SBR extended header" of fill element 1).

[0098] The MPEG-4 AAC standard anticipates that spectral band copy extension elements can contain PS (parametric stereo) data for the audio data used in a program. The MPEG-4 AAC standard also anticipates that when the header of a padding element (such as its extended payload) initializes an SBR object type (such as... Figure 7 When the "header 1" of the padding element contains PS data, the padding element (e.g., its extended payload) contains the spectral band copy data and the "bs_extension_id" parameter, the value of which (i.e., bs_extension_id = 2) indicates that the PS data is contained in the spectral band copy extension element of the padding element.

[0099] According to some embodiments of the invention, eSBR metadata (e.g., a tag indicating whether enhanced spectral band replication (eSBR) processing is performed on the audio content of a block) is included in the spectral band replication extension element of the padding element. For example, this tag is in Figure 7 The padding element 1 indicates that the marker appears after the header of the "SBR extension element" (the "SBR extension header" of padding element 1). This marker, along with additional eSBR metadata, is optionally included in the spectral band replication extension element after the header of the spectral band replication extension element (e.g., after the SBR extension header). Figure 7(In the SBR extension element of padding element 1 in the present invention). According to some embodiments of the present invention, the padding element containing eSBR metadata also includes a "bs_extension_id" parameter, the value of which (e.g., bs_extension_id = 3) indicates that the eSBR metadata is included in the padding element and that eSBR processing is performed on the audio content of the relevant block.

[0100] According to some embodiments of the present invention, eSBR metadata is included in the padding elements of the MPEG-4 AAC bitstream (e.g. Figure 7 The padding element 2) is in the spectral band copy extension element (SBR extension element) that is not a padding element. This is because the padding element containing extension_payload() (which has SBR data or SBR data with CRC) does not contain any other extension payload of any other extension type. Therefore, in embodiments where the eSBR metadata stores its own extension payload, a separate padding element is used to store the eSBR metadata. This padding element contains an identifier indicating the start of the padding element (e.g., ...). Figure 7 The padding data follows the identifier “ID2”. The padding data may contain the extension_payload() element (sometimes referred to herein as the extended payload) whose syntax is shown in Table 4.57 of the MPEG-4 AAC standard. The padding data (e.g., its extended payload) contains a header indicating the eSBR object (e.g., ...). Figure 7 The padding element 2 is “Header 2” (i.e., the header initializes the Enhanced Spectrum Band Replication (eSBR) object type), and the padding data (e.g., its extended payload) contains the eSBR metadata following the header. For example, Figure 7 Padding element 2 contains this header (“Header 2”) and also contains eSBR metadata following the header (i.e., the “tags” in padding element 2 that indicate whether Enhanced Spectral Band Replication (eSBR) processing is performed on the audio content of the block). Additional eSBR metadata may also be included after Header 2. Figure 7 The padding data in padding element 2. In the embodiments described in this paragraph, the header (e.g. Figure 7 The header 2) has an identification value that is not the regular value specified in Table 4.57 of the MPEG-4 AAC standard, but instead indicates the eSBR extension payload (so that the extension_type field of the header indicates that the padding data contains eSBR metadata).

[0101] In a first-type embodiment, the present invention is an audio processing unit (e.g., a decoder) comprising:

[0102] memory (e.g.) Figure 34 or 4 buffers 201), which are configured to store at least one block of an encoded audio bitstream (e.g., at least one block of an MPEG-4 AAC bitstream);

[0103] Bitstream payload deformatter (e.g.) Figure 3 Component 205 or Figure 4 Element 215), which is coupled to the memory and configured to demultiplex at least a portion of the block of the bit stream; and

[0104] Decoding subsystem (e.g.) Figure 3 Components 202 and 203 or Figure 4 Elements 202 and 213), coupled and configured to decode at least a portion of the audio content of the block of the bitstream, wherein the block comprises:

[0105] A padding element comprising an identifier indicating the start of the padding element (e.g., an "id_syn_ele" identifier with value 0x6 from Table 4.85 of the MPEG-4 AAC standard) and padding data following the identifier, wherein the padding data comprises:

[0106] At least one flag identifies whether enhanced spectrum band copying (eSBR) processing is performed on the audio content of the block (e.g., using spectrum band copying data and eSBR metadata contained in the block).

[0107] The tag is eSBR metadata, and an instance of the tag is the sbrPatchingMode tag. Another instance of the tag is the harmonicSBR tag. Both of these tags indicate whether the basic form of spectral band copying or an enhanced form of spectral copying is performed on the audio data of the block. The basic form of spectral copying is spectral patching, and the enhanced form of spectral band copying is harmonic transpose.

[0108] In some embodiments, the padding data also includes additional eSBR metadata (i.e., eSBR metadata other than the tags).

[0109] The memory may be a buffer memory (e.g.) Figure 4 An embodiment of buffer 201 stores (e.g., in a non-transitory manner) at least one block of the encoded audio bitstream.

[0110] It is estimated that the complexity of performing eSBR processing (using eSBR harmonic transpose and pre-flattening) by the eSBR decoder during the decoding of an MPEG-4 AAC bitstream containing eSBR metadata (indicating these eSBR tools) will be as follows (for typical decoding with indicator parameters):

[0111] ● Harmonic transpose (16kbps, 14400 / 28800Hz)

[0112] ○ Based on DFT: 3.68W MOPS (weighted million operations per second);

[0113] ○ Based on QMF: 0.98WMOPS;

[0114] ●QMF patch preprocessing (pre-flattening): 0.1 WMOPS.

[0115] It is well known that, for transients, DFT-based transpose typically performs better than QMF-based transpose.

[0116] According to some embodiments of the present invention, a padding element (of the encoded audio bitstream) containing eSBR metadata also includes a parameter (e.g., the "bs_extension_id" parameter) whose value (e.g., bs_extension_id = 3) indicates that eSBR metadata is contained in the padding element and that eSBR processing is performed on the audio content of the relevant block, and / or a parameter (e.g., the same "bs_extension_id" parameter) whose value (e.g., bs_extension_id = 2) indicates that the sbr_extension() container of the padding element contains PS data. For example, as indicated in Table 1 below, this parameter having a value of bs_extension_id = 2 indicates that the sbr_extension() container of the padding element contains PS data, and this parameter having a value of bs_extension_id = 3 indicates that the sbr_extension() container of the padding element contains eSBR metadata:

[0117] Table 1

[0118] bs_extension_id meaning 0 reserve 1 reserve 2 EXTENSION_ID_PS 3 EXTENSION_ID_ESBR

[0119] According to some embodiments of the present invention, the syntax of each spectrum band replication extension element containing eSBR metadata and / or PS data is indicated in Table 2 below (where “sbr_extension()” indicates a container for spectrum band replication extension elements, “bs_extension_id” is as described in Table 1 above, “ps_data” indicates PS data, and “esbr_data” indicates eSBR metadata):

[0120] Table 2

[0121]

[0122]

[0123] In the exemplary embodiment, esbr_data() mentioned in Table 2 above indicates the value of the following metadata parameters:

[0124] 1.1 Bit metadata parameter “bs_sbr_preprocessing”; and

[0125] 2. For each channel (“ch”) of the audio content of the encoded bitstream to be decoded, each of the above parameters is “sbrPatchingMode[ch]”, “SbrOversamplingFlag[ch]”, “SbrPitchInBinsFlag[ch]”, and “sbrPitchInBins[ch]”.

[0126] For example, in some embodiments, esbr_data() may have the syntax indicated in Table 3 to indicate these metadata parameters:

[0127] Table 3

[0128]

[0129]

[0130]

[0131] The above syntax enables the efficient implementation of enhanced forms of spectral band replication (e.g., harmonic transpose) as an extension of traditional decoders. Specifically, the eSBR data in Table 3 contains only the parameters required to perform the enhanced forms of spectral band replication, which are no longer supported in the bitstream and cannot be directly derived from the parameters already supported in the bitstream. All other parameters and processing data required to perform the enhanced forms of spectral band replication are extracted from readily available parameters at defined locations in the bitstream.

[0132] For example, MPEG-4 HE-AAC or HE-AAC v2 compatible decoders can be extended to include enhanced forms of spectral band copying, such as harmonic transpose. This enhanced form of spectral band copying is an addition to the basic form of spectral band copying already supported by the decoder. In the context of MPEG-4 HE-AAC or HE-AAC v2 compatible decoders, this basic form of spectral band copying is the QMF Spectral Patching SBR tool, as defined in section 4.6.18 of the MPEG-4 AAC standard.

[0133] When performing the enhanced form of spectrum band replication, the extended HE-AAC decoder can reuse many bitstream parameters already included in the SBR extended payload of the bitstream. Specific parameters that can be reused include, for example, various parameters that determine the master band table. These parameters include bs_start_freq (a parameter that determines the start of the master band table parameters), bs_stop_freq (a parameter that determines the stop of the master band table), bs_freq_scale (a parameter that determines the number of bands per octave), and bs_alter_scale (a parameter that modifies the scale of the bands). Reusable parameters also include parameters that determine the noise band table (bs_noise_bands) and the limiter band table (bs_limiter_bands). Therefore, in various embodiments, at least some equivalent parameters specified in the USAC standard are omitted from the bitstream to reduce the control burden on the bitstream. Typically, when a parameter specified in the AAC standard has an equivalent parameter specified in the USAC standard, the equivalent parameter specified in the USAC standard has the same name as the parameter specified in the AAC standard, such as the envelope scaling factor E. OrigMapped However, the equivalent parameters specified in the USAC standard typically have different values, and are "tuned" according to the enhanced SBR processing defined in the USAC standard rather than the SBR processing defined in the AAC standard.

[0134] It is recommended to enable enhanced SBR to improve the subjective quality of audio content with harmonic frequency structures and strong tonal characteristics, especially at low bit rates. The values ​​of the corresponding bitstream elements (i.e., esbr_data()) controlling these tools can be determined in the encoder by applying a signal dependency classification mechanism. Generally, the use of the harmonic patching method (sbrPatchingMode == 1) is preferred for encoding music signals at very low bit rates, where the audio bandwidth of the core codec is significantly limited. This is particularly evident when these signals contain significant harmonic structures. Conversely, the use of the conventional SBR patching method is preferred for speech and mixed signals because it provides better preservation of the temporal structure of speech.

[0135] To improve the performance of the harmonic transposer, a preprocessing step (bs_sbr_preprocessing == 1) can be initiated, which attempts to avoid introducing spectral discontinuities of the signal into the subsequent envelope adjuster. The tool's operation is beneficial for signal types where the coarse spectral envelope of the low-frequency band signal used for high-frequency reconstruction reveals large levels of variation.

[0136] To improve the transient response of harmonic SBR patching, adaptive frequency domain oversampling (sbrOversamplingFlag == 1) can be applied. Since adaptive frequency domain oversampling increases the computational complexity of the transposer and only benefits frames containing transients, the use of this tool is controlled by the bitstream element, which is transmitted once per frame and per independent SBR channel.

[0137] Decoders operating in the proposed enhanced SBR mode typically need to be able to switch between traditional SBR patching and enhanced SBR patching. Therefore, a delay potentially as long as the duration of a core audio frame can be introduced, depending on the decoder settings. Generally, the delays for traditional and enhanced SBR patching will be similar.

[0138] In addition to many parameters, other data elements may also be reused by the extended HE-AAC decoder when performing the enhanced form of spectrum band replication according to embodiments of the present invention. For example, envelope data and noise floor data may also be extracted from bs_data_env (envelope scaling factor) and bs_noise_env (noise floor scaling factor) data and used during the enhanced form of spectrum band replication.

[0139] Essentially, these embodiments leverage configuration parameters and envelope data already supported by traditional HE-AAC or HE-AAC v2 decoders in the SBR extended payload to enable an enhanced form of spectral band replication, requiring as little additional data transmission as possible. The metadata is initially tuned based on the basic form of HFR (e.g., spectral shift operation of SBR), but according to embodiments, it is used for an enhanced form of HFR (e.g., harmonic transpose of eSBR). As previously discussed, the metadata generally represents operating parameters tuned and designed for use with the basic form of HFR (e.g., linear spectral shift) (e.g., envelope scaling factor, noise floor scaling factor, time / frequency grid parameters, sine wave addition information, variable crossover frequency / band, inverse filtering mode, envelope resolution, smoothing mode, frequency interpolation mode). However, this metadata can be combined with additional metadata parameters specifically for an enhanced form of HFR (e.g., harmonic transpose) to efficiently and effectively process audio data using the enhanced form of HFR.

[0140] Therefore, an extended decoder supporting spectral band replication can be generated very efficiently by relying on predefined bitstream elements (e.g., bitstream elements in the SBR extended payload) and adding only the parameters required to support the enhanced form (in the padding element extended payload). This data reduction feature, combined with placing the newly added parameters in reserved data fields (e.g., extended containers), largely reduces the barriers to generating a decoder that supports the enhanced form by ensuring backward compatibility of the bitstream with conventional decoders that do not support spectral band replication.

[0141] In Table 3, the numbers in the right row indicate the number of bits in the corresponding parameter in the left row.

[0142] In some embodiments, the SBR object type defined in the MPEG-4 AAC is updated to include aspects of the SBR tool and the enhanced SBR (eSBR) tool, as indicated in the SBR extension element (bs_extension_id == EXTENSION_ID_ESBR). If the decoder detects and supports this SBR extension element, then the decoder adopts the indicated aspect of the enhanced SBR tool. The SBR object type updated in this manner is referred to as SBR enhancement.

[0143] In some embodiments, the present invention is a method comprising the steps of encoding audio data to produce an encoded bitstream (e.g., an MPEG-4 AAC bitstream), including including eSBR metadata in at least one segment of at least one block of the encoded bitstream and audio data in at least another segment of said block. In a typical embodiment, the method includes the step of multiplexing the audio data and eSBR metadata in each block of the encoded bitstream. In typical decoding of the encoded bitstream in an eSBR decoder, the decoder extracts eSBR metadata from the bitstream (including parsing and demultiplexing the eSBR metadata and audio data) and uses the eSBR metadata to process the audio data to produce a decoded audio data stream.

[0144] Another aspect of the invention is an eSBR decoder configured to perform eSBR processing (e.g., using at least one of eSBR tools called harmonic transpose or pre-flattening) during the decoding of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that does not contain eSBR metadata. (See also...) Figure 5 To describe an instance of this decoder.

[0145] Figure 5 The eSBR decoder 400 includes a buffer memory 201 connected as shown (which is identical to...) Figure 3 and 4 The memory 201), bit stream payload deformatter 215 (which is the same as) Figure 4 The deformatter 215), audio decoding subsystem 202 (sometimes referred to as the "core" decoding level or "core" decoding subsystem, and is identical to...) Figure 3 The core decoding subsystem 202), the eSBR control data generation subsystem 401, and the eSBR processing stage 203 (which is the same as...) Figure 3 (Level 203). Typically, the decoder 400 also includes other processing elements (not shown).

[0146] In the operation of decoder 400, the block sequence of the encoded audio bitstream (MPEG-4 AAC bitstream) received by decoder 400 is asserted from buffer 201 to deformatter 215.

[0147] Deformatter 215 is coupled and configured to demultiplex each block of the bitstream to extract SBR metadata (including quantization envelope data) and typically other metadata from it. Deformatter 215 is configured to assert at least the SBR metadata to the eSBR processing level 203. Deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding level) 202.

[0148] Audio decoding subsystem 202 of decoder 400 is configured to decode audio data extracted by deformatter 215 (this decoding may be referred to as the "core" decoding operation) to produce decoded audio data and assert the decoded audio data to eSBR processing stage 203. Decoding is performed in the frequency domain. Typically, the final processing stage in subsystem 202 applies a frequency-to-time domain transformation to the decoded frequency-domain audio data, such that the output of the subsystem is time-domain decoded audio data. Stage 203 is configured to apply SBR tools (and eSBR tools) indicated by SBR metadata (extracted by deformatter 215) and eSBR metadata generated in subsystem 401 to the decoded audio data (i.e., using SBR and eSBR metadata to perform SBR and ESBR processing on the output of decoding subsystem 202) to produce fully decoded audio data output from decoder 400. Typically, decoder 400 includes a memory (accessible by subsystem 202 and stage 203) storing deformatted audio data and metadata output from deformatter 215 (and optionally subsystem 401), and stage 203 is configured to access audio data and metadata as needed during SBR and eSBR processing. SBR processing in stage 203 can be considered as post-processing of the output of core decoding subsystem 202. Decoder 400 also optionally includes a final upmixing subsystem (which can apply the parametric stereo “PS” tool defined in the MPEG-4 AAC standard using PS metadata extracted by deformatter 215), coupled and configured to perform upmixing on the output of stage 203 to produce fully decoded upmixed audio output from APU 210.

[0149] Parametric stereo is an encoding tool that uses linear downmixing of the left and right channels of a stereo signal and a set of spatial parameters describing the stereo image to represent the stereo signal. Parametric stereo typically employs three types of spatial parameters: (1) Inter-channel intensity difference (IID), which describes the intensity difference between channels; (2) Inter-channel phase difference (IPD), which describes the phase difference between channels; and (3) Inter-channel coherence (ICC), which describes the coherence (or similarity) between channels. Coherence can be measured as the maximum value of the cross-correlation that varies with time or phase. These three parameters typically enable high-quality reconstruction of the stereo image. However, the IPD parameter only specifies the relative phase difference between the channels of the stereo input signal and does not indicate the distribution of these phase differences on the left and right channels. Therefore, a fourth type of parameter describing the total phase shift or total phase difference (OPD) can be used. During stereo reconstruction, consecutive window segments of both the received downmixed signal s[n] and its decorrelational version d[n] are processed along with spatial parameters to generate the left (l) according to the following equation. k (n)) and right (r) k (n)) Reconstructed signal:

[0150] I k (n)=H 11 (k, n)s k (n)+H 21 (k, n)d k (n)

[0151] r k (n)=H 12 (k, n)s k (n)+H 22 (k, n)d k (n)

[0152] Among them H 11 、H 12 、H 21 and H 22 Defined by stereo parameters. Finally, the signal l is transformed from frequency to time. k (n) and r k (n) Transform back to the time domain.

[0153] Figure 5 The control data generation subsystem 401 is coupled and configured to detect at least one property of the encoded audio bitstream to be decoded and to generate eSBR control data (which may be or contain any type of eSBR metadata contained in the encoded audio bitstream according to other embodiments of the invention) in response to at least one result of the detection step. The eSBR control data is asserted at level 203 to trigger the application of individual eSBR tools or combinations of eSBR tools and / or control the application of these eSBR tools after a specific property (or combination of properties) of the bitstream is detected. For example, to control the execution of eSBR processing using harmonic transposition, some embodiments of the control data generation subsystem 401 would include: a music detector (e.g., a simplified version of a conventional music detector) for setting the sbrPatchingMode[ch] parameter in response to detecting whether the bitstream indicates music (and asserting the setting parameter to level 203); a transient detector for setting the sbrOversamplingFlag[ch] parameter in response to detecting the presence or absence of transients in the audio content indicated by the bitstream (and asserting the setting parameter to level 203); and / or a spacing detector for setting the sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch] parameters in response to detecting spacing in the audio content indicated by the bitstream (and asserting the setting parameters to level 203). Other aspects of the invention are audio bitstream decoding methods performed by any embodiment of the inventive decoder described in this and previous paragraphs.

[0154] Aspects of the invention include encoding or decoding methods of any embodiment of the inventive APU, system, or apparatus configured (e.g., programmed) to perform. Other aspects of the invention include systems or apparatuses configured (e.g., programmed) to perform any embodiment of the inventive methods and computer-readable media (e.g., optical discs) storing (e.g., in a non-transitory manner) code for implementing any embodiment of the inventive methods or steps thereof. For example, the inventive system may be or include a programmable general-purpose processor, digital signal processor, or microprocessor that is programmed using software or firmware and / or otherwise configured to perform various operations on data (including embodiments of the inventive methods or steps thereof). This general-purpose processor may be or include a computer system that includes input devices, memory, and processing circuitry programmed (and / or otherwise configured) to perform embodiments of the inventive methods (or steps thereof) in response to data asserted thereto.

[0155] Embodiments of the present invention may be implemented in hardware, firmware, or software, or a combination of both (e.g., as a programmable logic array). Unless otherwise stated, the algorithms or processes included as part of the present invention are not inherently associated with any particular computer or other device. Specifically, various general-purpose machines may be used with programs written in accordance with the teachings herein, or more specialized devices (e.g., integrated circuits) may be more readily available for constructing to perform the desired method steps. Thus, the present invention may be implemented in one or more computer programs that execute on one or more programmable computer systems (e.g., ...). Figure 1 Components, or Figure 2 The encoder 100 (or its components), or Figure 3 decoder 200 (or its components), or Figure 4 The decoder 210 (or its components) or Figure 5 The decoder 400 (or any embodiment thereof) is described, and each of the one or more programmable computer systems includes at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known manner.

[0156] Each program can be implemented in any desired computer language (including machine, assembly, or high-level programming, logic, or object-oriented programming languages) to communicate with the computer system. In either case, the language can be a compiled or interpreted language.

[0157] For example, when implemented by a sequence of computer software instructions, the various functions and steps of embodiments of the present invention can be implemented by a multi-threaded sequence of software instructions that run in suitable digital signal processing hardware, in which case the various means, steps and functions of the embodiments can correspond to parts of the software instructions.

[0158] Each of these computer programs is preferably stored or downloaded to a storage medium or device (e.g., solid-state memory or media or magnetic or optical media) readable by a general-purpose or special-purpose programmable computer to configure and operate the computer when the storage medium or device is read by a computer system to execute the program described herein. The system of the present invention can also be implemented as a computer-readable storage medium configured (i.e., storing) a computer program, wherein such a storage medium causes the computer system to operate in a specific and predefined manner to perform the functions described herein.

[0159] Many embodiments of the invention have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the invention. Many modifications and variations of the invention can be made in light of the foregoing teachings. For example, to facilitate efficient implementation, phase shifting can be combined with complex QMF analysis and synthesis filter banks. The analysis filter bank is responsible for filtering the time-domain low-frequency signal generated by the core decoder into multiple sub-bands (e.g., QMF sub-bands). The synthesis filter bank is responsible for combining the regenerated high-frequency band generated by the selected HFR technique (as indicated by the received sbrPatchingMode parameter) with the decoded low-frequency band to produce a wideband output audio signal. However, a given filter bank implementation operating in a certain sampling rate mode (e.g., normal dual-rate operation or downsampled SBR mode) should not have a phase shift dependent on the bitstream. The QMF bank used in SBR is a theoretical complex exponential extension of the cosine modulation filter bank. It can be shown that when using complex exponential modulation to extend the cosine modulation filter bank, the frequency overlap cancellation constraint becomes obsolete. Therefore, for the SBR QMF bank, the analysis filter h k (n) and synthesis filter f k (n) Both can be defined by the following equation:

[0160]

[0161] Where p0(n) is a real-valued symmetric or asymmetric prototype filter (usually a low-pass prototype filter), M represents the number of channels, and N is the order of the prototype filter. The number of channels in the analysis filter bank can differ from the number of channels in the synthesis filter bank. For example, the analysis filter bank can have 32 channels and the synthesis filter bank can have 64 channels. When operating the synthesis filter bank in downsampling mode, the synthesis filter bank can have only 32 channels. Since the subband samples from the filter bank are complex values, an additive feasible channel-dependent phase shift step can be added to the analysis filter bank. These additional phase shifts need to be compensated before the synthesis filter bank. Although the phase shift term can theoretically have any value without disrupting the operation of the QMF analysis / synthesis chain, it can also be constrained to certain values ​​for consistency verification. The SBR signal is affected by the choice of phase factor, while the low-pass signal from the core decoder is not. The audio quality of the output signal is unaffected.

[0162] The coefficients p0(n) of the prototype filter can be defined as a length L of 640, as shown in Table 4 below.

[0163] Table 4

[0164]

[0165]

[0166]

[0167]

[0168]

[0169]

[0170]

[0171]

[0172] The prototype filter p0(n) can also be derived from Table 4 through one or more mathematical operations, such as rounding, subsampling, interpolation, and sampling.

[0173] Although the tuning of SBR-related control information generally does not depend on the transpose details (as previously discussed), in some embodiments, certain elements of the control data can be cascaded in the eSBR extension container (bs_extension_id == EXTENSION_ID_ESBR) to improve the quality of the regenerated signal. Some cascaded elements may include noise floor data (e.g., noise floor scaling factor and parameters indicating the direction (frequency or time direction) of differential encoding for each noise floor), inverse filtering data (e.g., parameters indicating inverse filtering modes selected from no inverse filtering, low inverse filtering, moderate inverse filtering, and strong inverse filtering), and missing harmonic data (e.g., parameters indicating whether a sine wave should be added to a specific frequency band of the regenerated high-frequency band). All these elements depend on the synthetic simulation of the decoder's transpose performed in the encoder and can therefore improve the quality of the regenerated signal after appropriate tuning according to the selected transpose.

[0174] Specifically, in some embodiments, missing harmonic and inverse filter control data (along with other bitstream parameters in Table 3) are transmitted in the eSBR extension container and tuned according to the eSBR's harmonic transposer. The additional bit rate required to transmit these two types of metadata from the eSBR's harmonic converter is relatively low. Therefore, sending the tuning missing harmonic and / or inverse filter control data in the eSBR extension container will improve the quality of the audio produced by the transposer while only slightly affecting the bit rate. To ensure backward compatibility with legacy decoders, the parameters tuned for the SBR's spectral shift operation can also be sent as part of the SBR control data in the bitstream using implicit or explicit transmission.

[0175] The complexity of the decoder with SBR enhancement described in this application must be limited to avoid significantly increasing the overall computational complexity of the implementation. Preferably, when using the eSBR tool, the PCU (MOP) of the SBR object type is equal to or less than 4.5, and the RCU of the SBR object type is equal to or less than 3. Approximate processing power is given in processor complexity units (PCU) (specified by an integer number of MOPS). Approximate RAM usage is given in RAM complexity units (RCU) (specified by an integer number of kWords (1000 words)). The number of RCUs does not include work buffers that can be shared between different objects and / or channels. Furthermore, the PCU is proportional to the sampling frequency. PCU values ​​are given in MOPS (millions of operations per second) per channel, and RCU values ​​are given in kilowords per channel.

[0176] Special attention needs to be paid to compressed data, such as HE-AAC encoded audio that can be decoded by different decoder configurations. In this case, decoding can be performed in a backward-compatible manner (AAC only) as well as in an enhanced manner (AAC+SBR). If the compressed data allows for both backward-compatible and enhanced decoding, and if the decoder operates in enhanced mode to use a post-processor that inserts some additional delay (e.g., an SBR post-processor in HE-AAC), then it must be ensured that this additional time delay caused by the backward-compatible mode is accounted for when presenting the combining unit, as described by the corresponding value n. To ensure proper handling of the combining timestamp (so that the audio is synchronized with other media), when the decoder operating mode includes the SBR enhancement described in this application (including eSBR), the additional delay introduced by post-processing, given the number of samples (per audio channel) at the output sampling rate, is 3010. Therefore, for the audio combining unit, when the decoder operating mode includes the SBR enhancement described in this application, the combining time is applied to the 3011th audio sample within the combining unit.

[0177] SBR enhancement should be enabled to improve the subjective quality of audio content with harmonic frequency structures and strong tonal characteristics, especially at low bit rates. The values ​​of the corresponding bitstream elements (i.e., esbr_data()) controlling these tools can be determined in the encoder by applying a signal dependency classification mechanism.

[0178] Generally, the use of the harmonic patching method (SBRPatchingMode == 0) is preferred for encoding music signals at very low bit rates, where the audio bandwidth of the core codec is significantly limited. This is especially true when these signals contain significant harmonic structures. Conversely, the use of the conventional SBR patching method is preferred for speech and mixed signals because it provides better preservation of the temporal structure of speech.

[0179] To improve the performance of the MPEG-4 SBR transposer, a preprocessing step (bs_sbr_preprocessing == 1) can be initiated, which avoids introducing spectral discontinuities of the signal into the subsequent envelope adjuster. The operation of this tool is beneficial for signal types where the coarse spectral envelope of the low-frequency band signal used for high-frequency reconstruction reveals large levels of variation.

[0180] To improve the transient response of harmonic SBR patching (sbrPatchingMode == 0), adaptive frequency domain oversampling (sbrOversamplingFlag == 1) can be applied. Since adaptive frequency domain oversampling increases the computational complexity of the transposer, but only benefits frames containing transients, the use of this tool is controlled by the bitstream element, which is transmitted once per frame and per independent SBR channel.

[0181] Typical bitrate settings for HE-AACv2 with SBR enhancement (i.e., a harmonic transposer with eSBR enabled) correspond to 20kbps to 32kbps for stereo audio content at sampling rates of 44.1kHz or 48kHz. The relative subjective quality gain of SBR enhancement increases towards the lower bitrate boundary, and a properly configured encoder allows this range to be extended even lower bitrates. The bitrates provided above are recommendations only and may be applicable to specific service requirements.

[0182] Decoders operating in the recommended enhanced SBR mode typically need to be able to switch between traditional SBR patching and enhanced SBR patching. Therefore, a delay as long as the duration of a core audio frame can be introduced, depending on the decoder settings. Generally, the delays of traditional SBR patching and enhanced SBR patching will be similar.

[0183] It should be understood that the invention may be practiced in ways other than those specifically described herein, within the scope of the appended claims. Any element symbols contained in the following claims are for illustrative purposes only and should in no way be used to interpret or limit the claims.

[0184] Various aspects of the invention can be understood from the following enumerated examples and embodiments (EEE):

[0185] EEE 1. A method for performing high-frequency reconstruction of an audio signal, the method comprising:

[0186] Receive an encoded audio bitstream, the encoded audio bitstream comprising audio data representing the low-frequency band portion of the audio signal and high-frequency reconstructed metadata;

[0187] Decode the audio data to generate a decoded low-frequency audio signal;

[0188] The high-frequency reconstruction metadata is extracted from the encoded audio bitstream, the high-frequency reconstruction metadata containing operational parameters of the high-frequency reconstruction process, the operational parameters including a patch mode parameter located in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates a spectral shift and a second value of the patch mode parameter indicates a harmonic transpose through frequency extension by a phase vocoder;

[0189] The decoded low-frequency audio signal is filtered to generate a filtered low-frequency audio signal;

[0190] The high-frequency band portion of the audio signal is regenerated using the filtered low-frequency audio signal and the high-frequency reconstruction metadata, wherein if the patching mode parameter is the first value, the regeneration includes spectral shifting, and if the patching mode parameter is the second value, the regeneration includes harmonic transposition via frequency stretching by a phase vocoder; and

[0191] The filtered low-frequency audio signal is combined with the regenerated high-frequency portion to form a broadband audio signal.

[0192] The filtering, regeneration, and combination are performed as a post-processing operation with a delay of 3010 samples or less per audio channel, and the spectral shift includes maintaining the ratio between the tone component and the noise-like component through adaptive inverse filtering.

[0193] EEE 2. The method according to EEE 1, wherein the encoded audio bitstream further includes a padding element having an identifier indicating the start of the padding element and padding data following the identifier, wherein the padding data includes the backward compatible extension container.

[0194] EEE 3. The method according to EEE 2, wherein the identifier is a 3-bit unsigned integer that transmits the most significant bit first and has a value of 0×6.

[0195] EEE 4. The method according to EEE 2 or EEE 3, wherein the padding data comprises an extended payload, the extended payload comprises spectrum band copy extended data, and the extended payload is identified by a 4-bit unsigned integer having a value of "1101" or "1110" transmitted first, and optionally,

[0196] The spectrum band replication extension data mentioned above includes:

[0197] Optional spectrum band copy header,

[0198] The spectrum band copy data, which is located after the header, and

[0199] A spectrum band copy extension element, which is located after the spectrum band copy data, wherein the marker is contained in the spectrum band copy extension element.

[0200] EEE 5. The method according to any one of EEE 1 to 4, wherein the high-frequency reconstruction metadata includes an envelope scaling factor, a background noise scaling factor, time / frequency grid information, or a parameter indicating a cross frequency.

[0201] EEE 6. The method according to any one of EEE 1 to 5, wherein the backward compatible extension container further includes a flag indicating whether additional preprocessing is used to avoid shape discontinuities in the spectral envelope of the high-frequency band portion when the patch mode parameter is equal to the first value, wherein a first value of the flag enables the additional preprocessing and a second value of the flag disables the additional preprocessing.

[0202] EEE 7. The method according to EEE 6, wherein the additional preprocessing includes using linear prediction filter coefficients to calculate a pregain curve.

[0203] EEE 8. The method according to any one of EEE 1 to 5, wherein the backward compatible extension container further includes a flag indicating whether signal adaptive frequency domain oversampling is applied when the patch mode parameter is equal to the second value, wherein a first value of the flag enables the signal adaptive frequency domain oversampling and a second value of the flag disables the signal adaptive frequency domain oversampling.

[0204] EEE 9. The method according to EEE 8, wherein the signal adaptive frequency domain oversampling is applied only to frames containing transients.

[0205] EEE 10. The method as described in any of the preceding EEEs, wherein the harmonic transpose via phase vocoder frequency extension is performed with an estimated complexity equal to or less than 4.5 million operations per second and 3,000 words of memory.

[0206] EEE 11. A non-transitory computer-readable medium containing instructions that, when executed by a processor, perform the method according to any one of EEE 1 to 10.

[0207] EEE 12. A computer program product having instructions which, when executed by a computing device or system, cause the computing device or system to perform the method according to any one of EEE 1 to 10.

[0208] EEE 13. An audio processing unit for performing high-frequency reconstruction of an audio signal, the audio processing unit comprising:

[0209] An input interface for receiving an encoded audio bitstream, the encoded audio bitstream comprising audio data representing the low-frequency band portion of the audio signal and high-frequency reconstructed metadata;

[0210] A core audio decoder, used to decode the audio data to produce a decoded low-frequency band audio signal;

[0211] A deformatter for extracting the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata containing operational parameters for the high-frequency reconstruction process, the operational parameters including a patch mode parameter located in a backward-compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates a spectral shift and a second value of the patch mode parameter indicates a harmonic transpose through frequency stretching by a phase vocoder;

[0212] An analysis filter bank is used to filter the decoded low-frequency band audio signal to produce a filtered low-frequency band audio signal;

[0213] A high-frequency regenerator for reconstructing the high-frequency portion of the audio signal using the filtered low-frequency audio signal and the high-frequency reconstruction metadata, wherein if the patching mode parameter is the first value, the reconstruction includes spectral shifting, and if the patching mode parameter is the second value, the reconstruction includes harmonic transposition via frequency extension by a phase vocoder; and

[0214] A synthesis filter bank is used to combine the filtered low-frequency audio signal with the regenerated high-frequency portion to form a broadband audio signal.

[0215] The analysis filter bank, high-frequency regenerator, and synthesis filter bank are executed in a postprocessor with a delay of 3010 samples or less per audio channel, and the spectral shift includes maintaining the ratio between the tone component and the noise-like component through adaptive inverse filtering.

[0216] EEE 14. The audio processing unit according to EEE 13, wherein the harmonic transpose via phase vocoder frequency extension is performed with an estimated complexity equal to or less than 4.5 million operations per second and 3,000 words of memory.

Claims

1. A method for performing high-frequency reconstruction of an audio signal, the method comprising: Receive an encoded audio bitstream, the encoded audio bitstream comprising audio data representing a low-frequency band portion of the audio signal and high-frequency reconstructed metadata, wherein the encoded audio bitstream further comprises a padding element having an identifier indicating the start of the padding element and padding data following the identifier, wherein the identifier is a 3-bit unsigned integer having a value of 0×6 and transmitting the most significant bit first. Decode the audio data to generate a decoded low-frequency audio signal; The high-frequency reconstruction metadata is extracted from the encoded audio bitstream, the high-frequency reconstruction metadata includes operating parameters of the high-frequency reconstruction process, the operating parameters include a patch mode parameter located in a backward compatible extension container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates a spectral shift and a second value of the patch mode parameter indicates a harmonic transpose through frequency extension by a phase vocoder, wherein the padding data includes the backward compatible extension container; The decoded low-frequency audio signal is filtered to generate a filtered low-frequency audio signal; The high-frequency band portion of the audio signal is regenerated using the filtered low-frequency audio signal and the high-frequency reconstruction metadata, wherein if the patching mode parameter is the first value, the regeneration includes spectral shifting, and if the patching mode parameter is the second value, the regeneration includes harmonic transposition via frequency stretching by a phase vocoder. The filtering and regeneration are performed as a post-processing operation with a delay of 3010 samples per audio channel, and the spectral shift includes maintaining the ratio between the tone component and the noise-like component through adaptive inverse filtering.

2. The method of claim 1, wherein the backward compatible extension container further includes a flag indicating whether additional preprocessing is used to avoid shape discontinuities in the spectral envelope of the high-frequency band portion when the patch mode parameter is equal to the first value, wherein a first value of the flag enables the additional preprocessing and a second value of the flag disables the additional preprocessing.

3. The method of claim 2, wherein the additional preprocessing comprises using linear prediction filter coefficients to calculate a pre-gain curve.

4. The method of claim 1, wherein the backward compatible extension container further includes a flag indicating whether to apply signal adaptive frequency domain oversampling when the patch mode parameter is equal to the second value, wherein a first value of the flag enables the signal adaptive frequency domain oversampling and a second value of the flag disables the signal adaptive frequency domain oversampling.

5. The method of claim 4, wherein the signal adaptive frequency domain oversampling is applied only to frames containing transients.

6. The method of claim 1, wherein the harmonic transpose via phase vocoder frequency stretching is performed with an estimated complexity of 4.5 million operations per second or less and 3,000 words or less.

7. An audio processing unit for performing high-frequency reconstruction of an audio signal, the audio processing unit comprising: An input interface is provided for receiving an encoded audio bitstream, the encoded audio bitstream comprising audio data representing a low-frequency band portion of the audio signal and high-frequency reconstructed metadata, wherein the encoded audio bitstream further comprises a padding element having an identifier indicating the start of the padding element and padding data following the identifier, wherein the identifier is a 3-bit unsigned integer having a value of 0×6 and transmitting the most significant bit first. A core audio decoder, used to decode the audio data to produce a decoded low-frequency band audio signal; A deformatter for extracting the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata containing operational parameters for the high-frequency reconstruction process, the operational parameters including a patch mode parameter located in a backward-compatible extended container of the encoded audio bitstream, wherein a first value of the patch mode parameter indicates a spectral shift and a second value of the patch mode parameter indicates a harmonic transpose through frequency extension by a phase vocoder, wherein the padding data includes the backward-compatible extended container; An analysis filter bank is used to filter the decoded low-frequency band audio signal to produce a filtered low-frequency band audio signal; as well as A high-frequency regenerator for reconstructing the high-frequency portion of the audio signal using the filtered low-frequency audio signal and the high-frequency reconstruction metadata, wherein if the patching mode parameter is the first value, the reconstruction includes spectral shifting, and if the patching mode parameter is the second value, the reconstruction includes harmonic transposition via frequency extension by a phase vocoder; and The analysis filter bank and high-frequency regenerator are executed in a postprocessor with a delay of 3010 samples per audio channel, and the spectral shift includes maintaining the ratio between the tone component and the noise-like component through adaptive inverse filtering.

Citation Information

Patent Citations

  • Audio Signal Synthesizer and Audio Signal Encoder

    US20110173006A1

  • Decoding audio bitstreams with enhanced spectral band replication metadata in at least one fill element

    US20180025737A1