Audio bit stream decoding method using enhanced spectrum band replicated metadata in at least one filling element
Enhanced spectral band replication metadata in MPEG-4 AAC bitstreams addresses spectral patching limitations by integrating eSBR tools, ensuring efficient and high-quality audio decoding with minimal bitrate and complexity.
Patent Information
- Application Number
- JP2025116276
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2015-03-16
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-03
AI Technical Summary
Spectral patching in MPEG-4 AAC standard is not ideal for certain audio types, particularly music content with low crossover frequencies, necessitating improved spectral band replication techniques.
Incorporation of enhanced spectral band replication (eSBR) metadata in MPEG-4 AAC bitstreams, including eSBR tools like harmonic transposition, QMF patching, and subband inter-sample Temporal Envelope Shaping, within filler elements to enhance high-frequency regeneration.
Enables efficient decoding of audio content with improved high-frequency reproduction, maintaining audio quality while minimizing bitrate impact and decoder complexity, supporting both legacy and eSBR-capable decoders.
Smart Images

Figure 2025146854000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to audio signal processing. Some embodiments relate to encoding and decoding audio bitstreams (e.g., bitstreams having MPEG-4 AAC format). Other embodiments relate to decoding such bitstreams by legacy decoders that are not configured to perform eSBR processing and that ignore such metadata, or to decoding audio bitstreams that do not include such metadata, including by generating eSBR control data in response to the bitstream. [Background technology]
[0002] A typical audio bitstream includes both audio data (e.g., encoded audio data) that describes one or more channels of audio content and metadata that describes at least one characteristic of the audio data or audio content. One well-known format for generating encoded audio bitstreams is the MPEG-4 Advanced Audio Coding (AAC) format, described in MPEG standard ISO / IEC 14496-3:2009. In the MPEG-4 standard, AAC stands for "advanced audio coding," and HE-AAC stands for "high-efficiency advanced audio coding."
[0003] The MPEG-4 AAC standard defines several audio profiles, which determine which objects and coding tools are present in a compliant encoder or decoder. Three of these audio profiles are (1) the AAC profile, (2) the HE-AAC profile, and (3) the HE-AAC v2 profile. The AAC profile includes the AAC low complexity (or "AAC-LC") object type. The AAC-LC object type corresponds to the MPEG-2 AAC low complexity profile with some adjustments and does not include the spectral band replication ("SBR") or parametric stereo ("PS") object types. The HE-AAC profile is a superset of the AAC profile and additionally includes the SBR object type. The HE-AAC v2 profile is a superset of the HE-AAC profile and additionally includes the PS object type.
[0004] The SBR object type includes a spectral band replication tool, an important coding tool that significantly improves the compression efficiency of perceptual audio codecs. SBR reconstructs the high-frequency components of an audio signal at the receiver (e.g., at the decoder). Therefore, the encoder only needs to encode and transmit low-frequency components, allowing much higher audio quality at low data rates. SBR is based on replicating, from available bandwidth-limited signal and control data obtained from the encoder, a sequence of harmonics that was previously truncated to reduce the data rate. The ratio between tone-like and noise-like components is maintained by adaptive inverse filtering and the optional addition of noise and sinusoids. In the MPEG-4 AAC standard, the SBR tool performs spectral patching, in which several adjacent quadrature mirror filter (QMF) subbands are copied from the transmitted low-frequency portion of the audio signal to the high-frequency portion of the audio signal generated at the decoder. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] MPEG standard ISO / IEC14496-3:2009 Summary of the Invention [Problem to be solved by the invention]
[0006] Spectral patching may not be ideal for certain audio types, such as music content, which has relatively low crossover frequencies. Therefore, techniques to improve spectral band replication are needed. [Means for solving the problem]
[0007] A first class of embodiments relates to an audio processing unit including a memory, a bitstream payload deformatter, and a decode subsystem. The memory is configured to store at least one block of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream). The bitstream payload deformatter is configured to demultiplex the encoded audio block. The decode subsystem is configured to decode audio content of the encoded audio block. The encoded audio block includes a fill element. The fill element has an identifier indicating the beginning of the fill element and fill data following the identifier. The fill data includes at least one flag identifying whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of the encoded audio block.
[0008] A second class of embodiments relates to a method for decoding an encoded audio bitstream, the method including receiving at least one block of the encoded audio bitstream, demultiplexing at least some portions of the at least one block of the encoded audio bitstream, and decoding at least some portions of the at least one block of the encoded audio bitstream. The at least one block of the encoded audio bitstream includes a fill element. The fill element has an identifier indicating the beginning of the fill element and fill data following the identifier. The fill data includes at least one flag identifying whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of the at least one block of the encoded audio bitstream.
[0009] Another class of embodiments relates to encoding and transcoding audio bitstreams that include metadata that identifies whether enhanced spectral band replication (eSBR) processing should be performed. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a block diagram of an embodiment of a system that may be configured to perform an embodiment of the method of the present invention. [Figure 2] FIG. 2 is a block diagram of an encoder that is an embodiment of the audio processing unit of the present invention. [Figure 3] 1 is a block diagram of a system including a decoder, which is an embodiment of an audio processing unit of the present invention, and optionally a post-processor coupled thereto; [Figure 4] FIG. 2 is a block diagram of a decoder that is an embodiment of the audio processing unit of the present invention. [Figure 5] FIG. 2 is a block diagram of a decoder, which is another embodiment of the audio processing unit of the present invention. [Figure 6] FIG. 2 is a block diagram of another embodiment of an audio processing unit of the present invention. [Figure 7] FIG. 1 illustrates a block of an MPEG-4 AAC bitstream containing divided segments. DETAILED DESCRIPTION OF THE INVENTION
[0011] Throughout this disclosure, including the claims, the expression performing an operation "on" a signal or data (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used broadly to refer to performing the operation either directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performing the operation).
[0012] Throughout this disclosure, including the claims, the term "audio processing unit" is used broadly to refer to a system, device, or apparatus configured to process audio data. Examples of audio processing units include, but are not limited to, encoders (e.g., transcoders), decoders, codecs, pre-processing systems, post-processing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools). Virtually every consumer electronic device, such as mobile phones, televisions, laptops, and tablet computers, includes an audio processing unit.
[0013] Throughout this disclosure, including the claims, the terms "couple" or "coupled" are used broadly to mean direct or indirect connections. Thus, when a first device couples to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections. Furthermore, components that are integrated within or together with other components are also coupled to each other.
[0014] Detailed Description of the Embodiments of the Present Invention The MPEG-4 AAC Standard contemplates that an encoded MPEG-4 AAC bitstream may contain metadata indicating each type of SBR processing, if any, to be applied by a decoder to decode the audio content of the bitstream and / or indicating at least one characteristic or parameter of at least one SBR tool to control such SBR processing and / or to be used to decode the audio content of the bitstream. The term "SBR metadata" is used herein to refer to this type of metadata described or referred to in the MPEG-4 AAC Standard.
[0015] The top level of an MPEG-4 AAC bitstream is a sequence of data blocks ("raw_data_block" elements), each of which is a segment of data (referred to herein as a "block") containing audio data and related information and / or other data (typically spanning a time period of 1024 or 960 samples). Here, we use the term "block" to refer to a segment of an MPEG-4 AAC bitstream containing audio data (and corresponding metadata and optionally other related data) that is determined or indicated by one (but not more than one) "raw_data_block" element.
[0016] Each block of an MPEG-4 AAC bitstream can contain several syntax elements (each of which is embodied as a segment of data in the bitstream). Seven types of such syntax elements are defined in the MPEG-4 AAC standard. Each syntax element is identified by a different value of the data element "id_syn_ele". Examples of syntax elements include "single_channel_element()", "channel_pair_element()", and "fill_element()". A single channel element is a container that contains audio data for a single audio channel (i.e., a monophonic audio signal). A channel pair element contains audio data for two audio channels (i.e., a stereo audio signal).
[0017] A fill element is an information container that contains an identifier (e.g., the value of the element "id_syn_ele" above) and subsequent data called "fill data". Fill elements have historically been used to adjust the instantaneous bit rate of a bitstream to be transmitted over a constant rate channel. By adding an appropriate amount of fill data to each block, a constant data rate can be achieved.
[0018] According to embodiments of the present invention, filler data may include one or more extension payloads that extend the types of data (e.g., metadata) that can be transmitted in a bitstream. A decoder that receives a bitstream with filler data that includes new types of data may optionally be used by a device (e.g., a decoder) that receives the bitstream to extend the capabilities of the device. Thus, as will be appreciated by those skilled in the art, filler elements are a special type of data structure that differs from data structures typically used to transmit audio data (e.g., audio payloads that include channel data).
[0019] In some embodiments of the invention, the identifier used to identify a filler element may consist of a three-bit unsigned integer transmitted most significant bit first ("uimsbf") with a value of 0x6. Several instances of the same type of syntax element (e.g., several filler elements) may occur in one block.
[0020] Another standard for encoding audio bitstreams is the MPEG Unified Speech and Audio Coding (USAC) standard (ISO / IEC 23003-3:2012). The MPEG USAC standard describes encoding and decoding audio content using a spectral band replication process (including the SBR process described in the MPEG-4 AAC standard, as well as other enhanced forms of spectral band replication). This process applies an extended and enhanced version of the set of SBR tools described in the MPEG-4 AAC standard (sometimes referred to in this document as "enhanced SBR tools" or "eSBR tools"). Thus, eSBR (as defined in the USAC standard) is an improvement over SBR (as defined in the MPEG-4 AAC standard).
[0021] In this document, the term "enhanced SBR processing" (or "eSBR processing") is used to refer to spectral band replication processing that uses at least one eSBR tool not described or referenced in MPEG-4 AAC (e.g., at least one eSBR tool described or referenced in the MPEG USAC standard). Examples of such eSBR tools are harmonic transposition, QMF patching, additional preprocessing or "pre-flattening," and subband inter-sample Temporal Envelope Shaping or "inter-TES."
[0022] Bitstreams produced in accordance with the MPEG USAC standard (sometimes referred to herein as USAC bitstreams) contain encoded audio content and typically include metadata indicating respective types of spectral band replication processes that should be applied by a decoder to decode the audio content of the USAC bitstream and / or metadata indicating at least one characteristic or parameter of at least one SBR tool and / or eSBR tool that should be used to control such spectral band replication processes and / or decode the audio content of the USAC bitstream.
[0023] As used herein, the term "enhanced SBR metadata" (or "eSBR metadata") refers to metadata not described or referenced in the MPEG-4 AAC Standard that indicates a respective type of spectral band replication process to be applied by a decoder to decode audio content in an encoded audio bitstream (e.g., a USAC bitstream) and / or that controls such spectral band replication process and / or that indicates at least one characteristic or parameter of at least one SBR tool and / or eSBR tool to be used to decode such audio content. An example of eSBR metadata is metadata (for indicating or controlling a spectral band replication process) that is described or referenced in the MPEG USAC Standard but not in the MPEG-4 AAC Standard. Thus, eSBR metadata herein refers to metadata that is not SBR metadata, and SBR metadata herein refers to metadata that is not eSBR metadata.
[0024] A USAC bitstream may include both SBR and eSBR metadata. More specifically, a USAC bitstream may include eSBR metadata that controls the decoder's performance of eSBR processing and SBR metadata that controls the decoder's performance of SBR processing. According to an exemplary embodiment of the present invention, eSBR metadata (e.g., eSBR-specific configuration data) is included (in accordance with the present invention) in an MPEG-4 AAC bitstream (e.g., in an sbr_extension() container at the end of the SBR payload).
[0025] During decoding of a bitstream encoded using an eSBR tool set (including at least one eSBR tool), the decoder performs eSBR processing to regenerate the high-frequency band of the audio signal based on replicating the harmonic sequences truncated during encoding. Such eSBR processing typically involves adjusting the spectral envelope of the generated high-frequency band, applying inverse filtering, and adding noise and sinusoidal components to recreate the spectral characteristics of the original audio signal.
[0026] According to an exemplary embodiment of the present invention, eSBR metadata (e.g., a small number of control bits that are eSBR metadata) is included in one or more metadata segments of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream). The encoded audio bitstream also contains encoded audio data in other segments (audio data segments). Typically, at least one such metadata segment in each block of the bitstream is (or contains) a filler element (including an identifier indicating the beginning of the filler element), and the eSBR metadata is included in the filler element after the identifier.
[0027] 1 is a block diagram of an exemplary audio processing chain (audio data processing system), one or more of whose elements may be configured in accordance with embodiments of the present invention. The system includes the following elements coupled together as shown: an encoder 1, a delivery subsystem 2, a decoder 3, and a post-processing unit 4. Variations on the illustrated system may omit one or more of the elements or include additional audio data processing units.
[0028] In some implementations, encoder 1 (which optionally includes a preprocessing unit) is configured to accept as input PCM (time-domain) samples containing audio content and to output an encoded audio bitstream (having a format conforming to the MPEG-4 AAC standard) representing the audio content. The data in the bitstream representing the audio content is sometimes referred to herein as "audio data" or "encoded audio data." When the encoder is configured according to exemplary embodiments of the present invention, the audio bitstream output from the encoder includes eSBR metadata (and typically other metadata) in addition to the audio data.
[0029] One or more encoded audio bitstreams output from Encoder 1 may be presented to Encoded Audio Delivery Subsystem 2. Subsystem 2 is configured to store and / or deliver each encoded bitstream output from Encoder 1. The encoded audio bitstreams output from Encoder 1 may be stored by Subsystem 2 (e.g., in the form of a DVD or Blu-ray disc), or may be transmitted by Subsystem 2 (which may implement a transmission link or network), or may be both stored and transmitted by Subsystem 2.
[0030] Decoder 3 is configured to decode the encoded MPEG-4 AAC audio bitstream (generated by Encoder 1) received via Subsystem 2. In some embodiments, decoder 3 is configured to extract eSBR metadata from each block of the bitstream and decode the bitstream (including by performing eSBR processing using the extracted eSBR metadata) to generate decoded audio data (e.g., a stream of decoded PCM audio samples). In some embodiments, decoder 3 is configured to extract SBR metadata from the bitstream (but ignore the eSBR metadata included in the bitstream) and decode the bitstream (including by performing SBR processing using the extracted SBR metadata) to generate decoded audio data (e.g., a stream of decoded PCM audio samples). Typically, decoder 3 includes a buffer that stores (e.g., non-temporarily) segments of the encoded audio bitstream received from Subsystem 2.
[0031] 1 is configured to accept a stream of decoded audio data (e.g., decoded PCM audio samples) from decoder 3 and perform post-processing thereon. The post-processing unit may be configured to render the post-processed audio content (or decoded audio received from decoder 3) for playback over one or more speakers.
[0032] FIG. 2 is a block diagram of an encoder (100), an embodiment of an audio processing unit of the present invention. Any of the components or elements of encoder 100 may be implemented in hardware, software, or a combination of hardware and software, as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit). Encoder 100 includes encoder 105, stuffer / formatter stage 107, metadata generation stage 106, and buffer memory 109, connected as shown. Typically, encoder 100 also includes other processing elements (not shown). Encoder 100 is configured to convert an input audio bitstream into an encoded output MPEG-4 AAC bitstream.
[0033] Metadata generator 106 is coupled and configured to generate (and / or pass through to stage 107) metadata (including eSBR metadata and SBR metadata) to be included by stage 107 in the encoded bitstream to be output from encoder 100.
[0034] Encoder 105 is coupled to and configured to encode input audio data (e.g., by performing compression on it) and to submit the resulting encoded audio to stage 107 for inclusion in an encoded bitstream to be output from stage 107.
[0035] Stage 107 is configured to multiplex the encoded audio from encoder 105 and the metadata (including eSBR metadata and SBR metadata) from generator 106 to generate an encoded bitstream to be output from stage 107. Preferably, the encoded bitstream has a format defined by one of the embodiments of the present invention.
[0036] The buffer memory 109 is configured to store (e.g., in a non-transient manner) at least one block of the encoded audio bitstream output from the stage 107. The sequence of blocks of the encoded audio bitstream is then presented from the buffer memory 109 to a delivery system as output from the encoder 100.
[0037] FIG. 3 is a block diagram of a system including a decoder (200), an embodiment of an audio processing unit of the present invention, and optionally a post-processor (300) coupled thereto. Any of the components or elements of decoder 200 may be implemented in hardware, software, or a combination of hardware and software, as one or more processes and / or one or more circuits (e.g., an ASIC, FPGA, or other integrated circuit). Decoder 200 includes, connected as shown, a buffer memory 201, a bitstream payload deformatter (parser) 205, an audio decode subsystem 202 (sometimes referred to as the "core" decode stage or "core" decode subsystem), an eSBR processing stage 203, and a control bit generation stage 204. Typically, decoder 200 also includes other processing elements (not shown).
[0038] Buffer memory (buffer) 201 stores (e.g., non-temporarily) at least one block of an encoded MPEG-4 AAC audio bitstream received by decoder 200. In operation of decoder 200, a sequence of blocks of the bitstream is presented from buffer 201 to deformatter 205.
[0039] In a variation of the FIG. 3 embodiment (or the FIG. 4 embodiment described below), an APU that is not a decoder (e.g., APU 500 of FIG. 6) includes a buffer memory (e.g., the same buffer memory as buffer 201) that stores (e.g., in a non-temporary manner) at least one block of an encoded audio bitstream (e.g., an MPEG-4 AAC audio bitstream) of the same type as that received by buffer 201 of FIG. 3 or FIG. 4 (i.e., an encoded audio bitstream that includes eSBR metadata).
[0040] 3, deformatter 205 is coupled and configured to demultiplex each block of the bitstream, extract therefrom SBR metadata (including quantized envelope data) and eSBR metadata (and typically other metadata as well), and provide at least the eSBR metadata and the SBR metadata to eSBR processing stage 203, and typically also provide other extracted metadata to decode subsystem 202 (and optionally to control bit generator 204). Deformatter 205 is also coupled and configured to extract audio data from each block of the bitstream and provide the extracted audio data to decode subsystem (decode stage) 202.
[0041] 3 also optionally includes a post-processor 300. Post-processor 300 includes a buffer memory (buffer) 301 and other processing elements (not shown), including at least one processing element coupled to buffer 301. Buffer 301 stores (e.g., non-temporarily) at least one block (or frame) of decoded audio data received by post-processor 300 from decoder 200. The processing element of post-processor 300 is coupled and configured to receive the sequence of decoded audio blocks (or frames) output from buffer 301 and adaptively process them using metadata output from decoding subsystem 202 (and / or deformatter 205) and / or control bits output from stage 204 of decoder 200.
[0042] Audio decoding subsystem 202 of decoder 200 is configured to decode the audio data extracted by parser 205 (such decoding may be referred to as a "core" decoding operation) to generate decoded audio data, and to submit the decoded audio data to eSBR processing stage 203. The decoding is performed in the frequency domain and typically includes inverse quantization followed by spectral processing. Typically, the final stage of processing in subsystem 202 applies a frequency-domain to time-domain transform to the decoded frequency-domain audio data, so that the output of the subsystem is time-domain decoded audio data. Stage 203 is configured to apply the SBR and eSBR tools indicated by the SBR and eSBR metadata (extracted by parser 205) to the decoded audio data (i.e., perform SBR and eSBR processing on the output of decode subsystem 202 using the SBR and eSBR metadata) to generate fully decoded audio data that is output from decoder 200 (e.g., to post-processor 300). Typically, decoder 200 includes a memory (accessible by subsystems 202 and stage 203) that stores the deformatted audio data and metadata output from deformatter 205, and stage 203 is configured to access the audio data and metadata (including the SBR and eSBR metadata) as needed during SBR and eSBR processing. The SBR and eSBR processing in stage 203 may be considered post-processing on the output of core decode subsystem 202. Optionally, decoder 200 also includes a final upmix subsystem, which may apply parametric stereo ("PS") tools defined in the MPEG-4 AAC standard using PS metadata extracted by deformatter 205 and / or control bits generated in subsystem 204.The upmix subsystem is coupled and configured to perform an upmix on the output of stage 203 to generate fully decoded, upmixed audio that is output from decoder 200. Alternatively, post-processor 300 is configured to perform an upmix on the output of decoder 200 (e.g., using PS metadata extracted by deformatter 205 and / or control bits generated in subsystem 204).
[0043] In response to the metadata extracted by the deformatter 205, the control bit generator 204 may generate control data. The control data may be used within the decoder 200 (e.g., in a final upmix subsystem) and / or may be presented as an output of the decoder 200 (e.g., to a post-processor 300 for use in post-processing). In response to the metadata extracted from the input bitstream (and optionally also in response to the control data), stage 204 may generate (present to post-processor 300) a control bit indicating that the decoded audio data output from the eSBR processing stage 203 should undergo a particular type of post-processing. In some implementations, the decoder 200 is configured to present the metadata extracted from the input bitstream by the deformatter 205 to the post-processor 300, and the post-processor 300 is configured to perform post-processing on the decoded audio data output from the decoder 200 using the metadata.
[0044] FIG. 4 is a block diagram of an audio processing unit ("APU") 210, another embodiment of an audio processing unit of the present invention. The APU 210 is a legacy decoder not configured to perform eSBR processing. Any of the components or elements of the APU 210 may be implemented in hardware, software, or a combination of hardware and software, as one or more processes and / or one or more circuits (e.g., an ASIC, FPGA, or other integrated circuit). The APU 210 includes, connected as shown, a buffer memory 201, a bitstream payload deformatter (parser) 215, an audio decode subsystem 202 (sometimes referred to as a "core" decode stage or "core" decode subsystem), and an SBR processing stage 213. Typically, the APU 210 also includes other processing elements (not shown).
[0045] Elements 201 and 202 of APU 210 are identical to the like-numbered elements of decoder 200 (of FIG. 3), and their description above will not be repeated. In operation of APU 210, a sequence of blocks of an encoded audio bitstream (MPEG-4 AAC bitstream) received by APU 210 is presented from buffer 201 to deformatter 215.
[0046] Deformatter 215 is coupled to and configured to demultiplex each block of the bitstream and extract therefrom SBR metadata (including quantized envelope data), typically along with other metadata, but ignoring eSBR that may be included in the bitstream according to any embodiment of the present invention. Deformatter 215 is configured to provide at least the SBR metadata to SBR processing stage 213. Deformatter 215 is also coupled to and configured to extract audio data from each block of the bitstream and provide the extracted audio data to decoding subsystem (decode stage) 202.
[0047] The audio decode subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the deformatter 215 (such decoding may be referred to as a "core" decoding operation) to generate decoded audio data and present the decoded audio data to the SBR processing stage 213. The decoding is performed in the frequency domain. Typically, the final stage of processing in the subsystem 202 applies a frequency-domain to time-domain transform to the decoded frequency-domain audio data, so that the output of the subsystem is time-domain decoded audio data. The stage 213 is configured to apply the SBR tools (but not the eSBR tools) indicated by the SBR metadata (extracted by the deformatter 215) to the decoded audio data (i.e., perform SBR processing on the output of the decode subsystem 202 using the SBR metadata) to generate fully decoded audio data that is output from the APU 210 (e.g., to the post-processor 300). Typically, APU 210 includes a memory (accessible by subsystem 202 and stage 213) that stores the deformatted audio data and metadata output from deformatter 215, with stage 213 configured to access the audio data and metadata (including SBR metadata) as needed during SBR processing. The SBR processing in stage 213 may be considered post-processing on the output of core decode subsystem 202. Optionally, APU 210 also includes a final upmix subsystem (which may apply parametric stereo ("PS") tools defined in the MPEG-4 AAC standard using the PS metadata extracted by deformatter 215). The upmix subsystem is coupled and configured to perform an upmix on the output of stage 213 to generate fully decoded, upmixed audio that is output from APU 210.Alternatively, a post-processor may be configured to perform an upmix on the output of the APU 210 (eg, using PS metadata extracted by the deformatter 215 and / or control bits generated in the APU 210).
[0048] Various implementations of the encoder 100, decoder 200 and APU 210 are configured to perform different embodiments of the method of the present invention.
[0049] According to some embodiments, eSBR metadata (e.g., a small number of control bits that are the eSBR metadata) is included in the encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) so that legacy decoders (not configured to parse the eSBR metadata or use any eSBR tools related to the eSBR metadata) ignore the eSBR metadata but can still decode the bitstream to the extent possible without using the eSBR metadata or any eSBR tools related to the eSBR metadata, typically without any significant penalty in decoded audio quality. However, eSBR decoders configured to parse the bitstream to identify the eSBR metadata and use at least one eSBR tool in response to the eSBR metadata still benefit from the use of at least one such eSBR tool. Thus, embodiments of the present invention provide a means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner.
[0050] Typically, the eSBR metadata in the bitstream indicates one or more of the following eSBR tools (as described in the MPEG USAC standard, and which may or may not have been applied by an encoder in generating the bitstream) (e.g., indicates at least one characteristic or parameter of one or more of the following eSBR tools): · Harmonic conversion; QMF patching additional preprocessing (pre-flattening); and Subband Inter-sample Temporal Envelope Shaping or "Inter-TES". For example, eSBR metadata included in the bitstream may indicate values for the following parameters (described in the MPEG USAC standard and this disclosure): harmonicSBR[ch], sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch], bs_interTes, bs_temp_shape[ch][env], bs_inter_temp_shape_mode[ch][env], and bs_sbr_preprocessing.
[0051] Here, the notation X[ch], where X is some parameter, denotes that the parameter relates to a certain channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, we sometimes abbreviate the notation [ch] and assume that the associated parameter relates to a certain channel of audio content.
[0052] Here, the notation X[ch][env], where X is some parameter, denotes that the parameter relates to the SBR envelope ("env") of a certain channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, we sometimes abbreviate the expressions [env] and [ch], and assume that the associated parameter relates to the SBR envelope of a certain channel of audio content.
[0053] As mentioned above, MPEG USAC contemplates that USAC bitstreams contain eSBR metadata that controls the decoder's execution of eSBR processing. The eSBR metadata includes the following one-bit metadata parameters: harmonicSBR; bs_interTES; and bs_pvc.
[0054] The parameter harmonicSBR indicates the use of harmonic patching (harmonic transposition) for SBR. Specifically, harmonicSBR=0 indicates non-harmonic spectral patching as described in Section 4.6.18.6.3 of the MPEG-4 AAC Standard; harmonicSBR=1 indicates harmonic SBR patching (of the type used in eSBR as described in Sections 7.5.3 or 7.5.4 of the MPEG USAC Standard). Harmonic SBR patching is not used with non-eSBR spectral band replication (i.e., non-eSBR SBR). Throughout this disclosure, the basic form of spectral band replication is referred to as spectral patching, and the enhanced form of spectral band replication is referred to as harmonic transposition.
[0055] The value of the parameter bs_interTES indicates the use of the InterTES tool of eSBR.
[0056] The value of the parameter bs_pvc indicates the use of the PVC tool of eSBR.
[0057] During decoding of an encoded bitstream, the performance of harmonic transformations during the eSBR processing stage of the decoding (for each channel "ch" of the audio content represented by the bitstream) is controlled by the following eSBR metadata parameters: sbrPatchingMode[ch]; sbrOversamplingFlag[ch]; sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch].
[0058] The value of sbrPatchingMode[ch] indicates the type of transposer used in eSBR: sbrPatchingMode[ch]=1 indicates non-harmonic patching as described in section 4.6.18.6.3 of the MPEG-4 AAC standard; sbrPatchingMode[ch]=0 indicates harmonic SBR patching as described in sections 7.5.3 or 7.5.4 of the MPEG USAC standard.
[0059] The value of sbrOversamplingFlag[ch] indicates the use of signal-adaptive frequency-domain oversampling in eSBR in combination with DFT-based harmonic SBR patching as described in Section 7.5.3 of the MPEG USAC Standard. This flag controls the size of the DFT used in the commutator. 1 indicates signal-adaptive frequency-domain oversampling enabled as described in Section 7.5.3.1 of the MPEG USAC Standard; 0 indicates signal-adaptive frequency-domain oversampling disabled as described in Section 7.5.3.1 of the MPEG USAC Standard.
[0060] The value of sbrPitchInBinsFlag[ch] controls the interpretation of the sbrPitchInBins[ch] parameter: 1 indicates that the value in sbrPitchInBins[ch] is valid and greater than 0; 0 indicates that the value of sbrPitchInBins[ch] is set to 0.
[0061] The value of sbrPitchInBins[ch] controls the addition of cross-product terms in the SBR harmonic converter. The value sbrPitchInBins[ch] is an integer value in the range [0,127] and represents the distance measured in frequency bins for a 1536-line DFT operating on the sampling frequency of the core encoder.
[0062] If an MPEG-4 AAC bitstream exhibits a disjointed SBR channel pair (rather than a single SBR channel), the bitstream exhibits two instances of the above syntax (for harmonic or non-harmonic conversion): one instance for each channel of sbr_channel_pair_element().
[0063] Harmonic transposition in eSBR tools typically improves the quality of the decoded music signal at relatively low crossover frequencies. Non-harmonic transposition (i.e., legacy spectral patching) typically improves the speech signal. Therefore, the starting point in determining which type of transposition is preferred for encoding specific audio content is to select a transposition method depending on the speech / music detection. Here, harmonic transposition is used for music content and spectral patching is used for speech content.
[0064] The implementation of pre-flattening during eSBR processing is controlled by the value of a single-bit eSBR metadata parameter known as bs_sbr_preprocessing, in the sense that pre-flattening is either performed or not performed depending on the value of this single bit. When the SBR QMF patching algorithm described in section 4.6.18.6.3 of the MPEG-4 AAC standard is used, a pre-flattening stage may be performed (as indicated by the bs_sbr_preprocessing parameter) to avoid discontinuities in the shape of the high-frequency signal's spectral envelope from being input to the subsequent envelope adjuster (which performs another stage of the eSBR processing). Pre-flattening typically improves the performance of the subsequent envelope adjustment stage, resulting in a more stable perceived high-frequency signal.
[0065] The performance of inter-subband sample Temporal Envelope Shaping (the "inter-TES" tool) during eSBR processing in the decoder is controlled by the following eSBR metadata parameters for each SBR envelope ("env") for each channel ("ch") of the audio content of the USAC bitstream being decoded: bs_temp_shape[ch][env] and bs_inter_temp_shape_mode[ch][env].
[0066] The InterTES tool processes QMF subband samples after the envelope adjuster. This processing stage shapes the temporal envelope of the higher frequency bands with a finer temporal granularity than that of the envelope adjuster. InterTES shapes the temporal envelope between QMF subband samples by applying a gain factor to each QMF subband sample in the SBR envelope.
[0067] The parameter bs_temp_shape[ch][env] is a flag that signals the use of Inter TES. The parameter bs_inter_temp_shape_mode[ch][env] indicates the value of the parameter γ in Inter TES (as defined in the MPEG USAC standard).
[0068] The overall bitrate requirement for including eSBR metadata indicating the above-mentioned eSBR tools (harmonic conversion, pre-flattening, and InterTES) in an MPEG-4 AAC bitstream is expected to be on the order of a few hundred bits per second, since, according to some embodiments of the present invention, only the differential control data required to perform the eSBR processing is transmitted. This information is included in a backwards-compatible manner (as explained below), so that legacy decoders can ignore it. Therefore, the negative bitrate impact associated with including eSBR metadata is negligible for several reasons, including the following: The bitrate penalty (due to the inclusion of eSBR metadata) is a very small percentage of the total bitrate, since only the differential control data required to perform eSBR processing is transmitted (not a simulcast of SBR control data); Tuning of control information related to SBR is typically independent of conversion details; and · The InterTES tool (used during eSBR processing) performs single-ended post-processing of the converted signal.
[0069] Thus, embodiments of the present invention provide a means for efficiently transmitting enhanced Spectral Band Replication (eSBR) control data or metadata in a backward-compatible manner. This efficient transmission of eSBR control data reduces memory requirements in decoders, encoders, and transcoders employing aspects of the present invention, without any appreciable adverse impact on bitrate. Furthermore, the complexity and processing requirements associated with implementing eSBR in accordance with embodiments of the present invention are also reduced, as the SBR data only needs to be processed once and does not need to be simulcast, as would be the case if eSBR were treated as an entirely separate object type in MPEG-4 AAC rather than being integrated into the MPEG-4 AAC codec in a backward-compatible manner.
[0070] The elements of a block (raw_data_block) of an MPEG-4 AAC bitstream into which eSBR metadata may be included according to some embodiments of the present invention will now be described with reference to Figure 7. Figure 7 is a diagram of a block (raw_data_block) of an MPEG-4 AAC bitstream showing some of its segments.
[0071] A block of an MPEG-4 AAC bitstream may contain at least one single_channel_element() (e.g., the single channel element shown in Figure 7) and / or at least one channel_pair_element() (not specifically shown in Figure 7, but which may be present) that contain audio data for an audio program. A block may also contain several fill_elements (e.g., fill element 1 and / or fill element 2 in Figure 7) that contain program-related data (e.g., metadata). Each single_channel_element() contains an identifier (e.g., "ID1" in Figure 7) that indicates the start of the single channel element and may contain audio data representing different channels of a multi-channel audio program. Each channel_pair_element contains an identifier (not shown in Figure 7) that indicates the start of a channel pair element and may contain audio data representing two channels of the program.
[0072] A fill_element (referred to herein as a fill element) in an MPEG-4 AAC bitstream contains an identifier (e.g., "ID2" in Figure 7) that indicates the beginning of the fill element, followed by fill data. The identifier ID2 may consist of a three-bit unsigned integer ("uimsbf") transmitted most significant bit first with a value of 0x6. The fill data may contain an extension_payload() element (sometimes referred to herein as an extension payload), the syntax of which is shown in Table 4.57 of the MPEG-4 AAC standard. Several types of extension payload exist, and are identified through the extension_type parameter, which is a four-bit unsigned integer ("uimsbf") transmitted most significant bit first.
[0073] The filler data (e.g., its extension payload) may include a header or identifier (e.g., "Header 1" in FIG. 7) that indicates a segment of the filler data that represents an SBR object (i.e., the header initializes an "SBR object" type, referred to as sbr_extension_data() in the MPEG-4 AAC standard). For example, a Spectral Band Replication (SBR) extension payload is identified with a value of "1101" or "1110" for the extension_type field in the header, where the identifier "1101" identifies an extension payload with SBR data and "1110" identifies an extension payload with SBR data with a cyclic redundancy check (CRC) to verify the correctness of the SBR data.
[0074] When a header (e.g., the extension_type field) initializes an SBR object type, the header is followed by SBR metadata (sometimes referred to herein as "spectral band replication data" and referred to as sbr_data() in the MPEG-4 AAC standard), which may be followed by at least one spectral band replication extension element (e.g., the "SBR extension element" of filler element 1 in Figure 7). Such a spectral band replication extension element (a segment of a bitstream) is referred to as an sbr_extension() container in the MPEG-4 AAC standard. A spectral band replication extension element optionally includes a header (e.g., the "SBR extension header" of filler element 1 in Figure 7).
[0075] The MPEG-4 AAC standard contemplates that a spectral band replication extension element can contain PS (parametric stereo) data for the audio data of a program. The MPEG-4 AAC standard contemplates that a filler element (e.g., its extension payload) contains a spectral band replication data bs_extension_id parameter when the header of the filler element (e.g., its extension payload) initializes an SBR object type (such as "Header 1" in Figure 7) and the filler element's spectral band replication extension element contains PS data. A value of this parameter (i.e., bs_extension_id=2) indicates that PS data is included in the filler element's spectral band replication extension element.
[0076] According to some embodiments of the invention, eSBR metadata (e.g., a flag indicating whether enhanced spectral band replication (eSBR) processing is to be performed on the audio content of that block) is included in the spectral band replication extension element of a filler element. For example, such a flag is included in filler element 1 of FIG. 7, where the flag appears after the header of the "SBR extension element" of filler element 1 ("SBR extension header" of filler element 1). Optionally, such a flag and additional eSBR metadata are included in the spectral band replication extension element after the header of the spectral band replication extension element (e.g., after the SBR extension header in the SBR extension element of filler element 1 in FIG. 7). According to some embodiments of the invention, a filler element containing eSBR metadata also includes a bs_extension_id parameter. A value of that parameter (e.g., bs_extension_id=3) indicates that the filler element contains eSBR metadata and that eSBR processing should be performed on the audio content of that block.
[0077] According to some embodiments of the present invention, eSBR metadata is included in a filler element of an MPEG-4 AAC bitstream other than the filler element's spectral band replication extension element (SBR extension element) (e.g., filler element 2 in FIG. 7). This is because a filler element containing an extension_payload() with SBR data or SBR data with a CRC does not contain any other extension payload of any other extension type. Therefore, in embodiments in which the eSBR metadata is stored in its own extension payload, a separate filler element is used to store the eSBR metadata. Such a filler element includes an identifier (e.g., "ID2" in FIG. 7) that indicates the beginning of the filler element, followed by the filler data. The filler data can include an extension_payload() element (sometimes referred to herein as the extension payload), the syntax of which is shown in Table 4.57 of the MPEG-4 AAC Standard. The filler data (e.g., its extension payload) may include a header (e.g., "Header 2" of filler element 2 in FIG. 7) that indicates an eSBR object (i.e., the header initializes an enhanced spectral band replication (eSBR) object type), and the filler data (e.g., its extension payload) includes eSBR metadata after the header. For example, filler element 2 in FIG. 7 includes such a header ("Header 2") that also includes eSBR metadata after the header (i.e., a "flag" in filler element 2 that indicates whether enhanced spectral band replication (eSBR) processing is to be performed on the audio content of the block). Optionally, additional eSBR metadata may also be included in the filler data of filler element 2 in FIG. 7 after Header 2. In the embodiment described in this paragraph, the header (e.g., Header 2 in FIG. 7) has a distinguishing value that indicates an eSBR extension payload (thus, the extension_type field of the header indicates that the filler data includes eSBR metadata) rather than one of the usual values specified in Table 4.57 of the MPEG-4 AAC Standard.
[0078] In a first class of embodiments, the present invention provides an audio processing unit (e.g., a decoder) comprising: a memory (e.g., buffer 201 of FIG. 3 or FIG. 4) configured to store at least one block of an encoded audio bitstream (e.g., at least one block of an MPEG-4 AAC bitstream); a bitstream payload deformatter (e.g., element 205 of FIG. 3 or element 215 of FIG. 4 ) coupled to the memory and configured to demultiplex at least some of the blocks of the bitstream; a decoding subsystem (e.g., elements 202 and 203 of FIG. 3 or elements 202 and 213 of FIG. 4) coupled and configured to decode at least a portion of the audio content of the block of the bitstream, the block comprising: It includes a filler element, an identifier indicating the beginning of the filler element (for example, the id_syn_ele identifier having the value 0x6 in Table 4.85 of the MPEG-4 AAC standard), and filler data following the identifier, the filler data being: at least one flag that identifies whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of the block (e.g., using spectral band replication data and eSBR metadata included in the block); An audio processing unit.
[0079] The flags are eSBR metadata, and an example of such a flag is the sbrPatchingMode flag. Another example of such a flag is the harmonicSBR flag. Both of these flags indicate whether basic or enhanced spectral band replication should be performed on the audio data of the block. Basic spectral band replication is spectral patching, and enhanced spectral band replication is harmonic translation.
[0080] In some embodiments, the filler data also includes additional eSBR metadata (ie, eSBR metadata other than the flag).
[0081] The memory may be a buffer memory (eg, an implementation of buffer 201 of FIG. 4) that stores (eg, in a non-transient manner) the at least one block of an encoded audio bitstream.
[0082] The complexity of the execution of eSBR processing (using eSBR harmonic conversion, pre-flattening, and InterTES tools) by an eSBR decoder during decoding of an MPEG-4 AAC bitstream containing eSBR metadata (the eSBR metadata indicates these eSBR tools) is estimated to be (for a typical decoding using the parameters shown) as follows: ●Harmonic conversion (16kbps, 14400 / 28800Hz) DFT-based: 3.68 WMOPS (weighted million operations per second); ○WMF base: 0.98WMOPS; ●QMF patching preprocessing (pre-flattening): 0.1WMOPS; ●Subband / Inter-sample Temporal Envelope Shaping (Inter-TES): at most 0.16 WMOPS For transient components, it has been found that DFT-based transforms typically perform better than QMF-based transforms.
[0083] According to some embodiments of the present invention, a filler element (of an encoded audio bitstream) containing eSBR metadata also includes a parameter (e.g., a bs_extension_id parameter) with a value (e.g., bs_extension_id=3) that signals that eSBR metadata is included in the filler element and that eSBR processing should be performed on the audio content of that block, and / or a parameter (e.g., the same bs_extension_id parameter) with a value (e.g., bs_extension_id=2) that signals that the filler element's sbr_extension() container contains PS data. For example, as shown in Table 1 below, such a parameter with the value bs_extension_id=2 may signal that the filler element's sbr_extension() container contains PS data, and such a parameter with the value bs_extension_id=3 may signal that the filler element's sbr_extension() container contains eSBR metadata.
[0084] [Table 1] According to some embodiments of the present invention, the syntax of each spectrum band duplication extension element containing eSBR metadata and / or PS data is as shown in Table 2 below (where sbr_extension() represents a container that is a spectrum band duplication extension element, bs_extension_id is as described in Table 1 above, ps_data represents PS data, and esbr_data represents eSBR metadata).
[0085] [Table 2] In one exemplary embodiment, the esbr_data() referenced in Table 2 above indicates values for the following metadata parameters: 1. The above one-bit metadata parameters harmonicSBR; bs_interTES; and bs_sbr_preprocessing; 2. For each channel ("ch") of the audio content of the encoded bitstream to be decoded, each of the above parameters: sbrPatchingMode[ch]; sbrOversamplingFlag[ch]; sbrPitchInBinsFlag[ch]; and sbrPitchInBins[ch]; and 3. For each SBR envelope ("env") of each channel ("ch") of the audio content of the encoded bitstream to be decoded, specify the above parameters: bs_temp_shape[ch][env]; and bs_inter_temp_shape_mode[ch][env], respectively.
[0086] For example, in some embodiments, esbr_data() may have the syntax shown in Table 3 to indicate these metadata parameters.
[0087] [Table 3] The above syntax allows for the efficient implementation of enhanced forms of spectral band replication, such as harmonic transposition, as extensions to legacy decoders. Specifically, the eSBR data in Table 3 includes only parameters required to perform the enhanced form of spectral band replication that are not already supported in the bitstream or directly derivable from parameters already supported in the bitstream. All other parameters and processing data required to perform the enhanced form of spectral band replication are extracted from existing parameters in already defined positions in the bitstream.
[0088] For example, an MPEG-4 HE-AAC or HE-AAC-v2 compliant decoder may be extended to include an enhanced form of spectral band replication, such as harmonic transformation, in addition to the basic form of spectral band replication already supported by the decoder. In the context of an MPEG-4 HE-AAC or HE-AAC-v2 compliant decoder, this basic form of spectral band replication is the QMF spectral patching SBR tool defined in clause 4.6.18 of the MPEG-4 AAC standard.
[0089] When performing the improved form of spectral band replication, the enhanced HE-AAC decoder may reuse many of the bitstream parameters already included in the SBR extension payload of the bitstream. Specific parameters that may be reused include, for example, various parameters that determine the master frequency band table. These parameters include bs_start_freq (a parameter that determines the start of the master frequency table parameter), bs_stop_freq (a parameter that determines the end of the master frequency table parameter), bs_freq_scale (a parameter that determines the number of frequency bands per octave), and bs_alter_scale (a parameter that alters the frequency band scale). Parameters that may be reused also include a parameter that determines the noise band table (bs_noise_bands) and a limiter band table parameter (bs_limiter_bands). Thus, in various embodiments, at least some of the equivalent parameters specified in the USAC standard are omitted from the bitstream, thereby reducing control overhead in the bitstream. Typically, if a parameter specified in the AAC standard has an equivalent parameter specified in the USAC standard, the equivalent parameter specified in the USAC standard will have the same name as the parameter specified in the AAC standard, e.g., envelope scale factor E. OrigMappedHowever, the equivalent parameters specified in the USAC standard typically have different values that are "tuned" for the enhanced SBR process defined in the USAC standard rather than for the SBR process defined in the AAC standard.
[0090] In addition to many of the parameters mentioned above, other data elements may also be reused by an enhanced HE-AAC decoder when performing enhanced spectral band replication in accordance with embodiments of the present invention. For example, envelope data and noise floor data may be extracted from the bs_data_env and bs_noise_env data and used during enhanced spectral band replication.
[0091] Essentially, these embodiments leverage configuration parameters and envelope data already supported by legacy HE-AAC or HE-AAC v2 decoders in the SBR extension payload to enable enhanced spectral band replication with as little additional transmission data as possible. Thus, enhanced decoders supporting enhanced spectral band replication can be generated in a highly efficient manner by relying on already defined bitstream elements (e.g., those in the SBR extension payload) and adding only the parameters needed to support the enhanced spectral band replication (in the filler element extension payload). This data reduction feature, combined with placing newly added parameters in reserved data fields such as extension containers, substantially reduces the barrier to creating decoders that support enhanced spectral band replication by ensuring that the bitstream is backward compatible with legacy decoders that do not support enhanced spectral band replication.
[0092] In Table 3, the numbers in the center column indicate the number of bits of the corresponding parameter in the left column.
[0093] In some embodiments, the invention is a method that includes encoding audio data to generate an encoded bitstream (e.g., an MPEG-4 AAC bitstream) by including eSBR metadata in at least one segment of at least one block of the encoded bitstream and audio data in at least one other segment of the block. In an exemplary embodiment, the method includes multiplexing the audio data with the eSBR metadata in each block of the encoded bitstream. In exemplary decoding of the encoded bitstream in an eSBR decoder, the decoder extracts the eSBR metadata from the bitstream (including by parsing and demultiplexing the eSBR metadata and audio data) and uses the eSBR metadata to process the audio data to generate a stream of decoded audio data.
[0094] Another aspect of the invention is an eSBR decoder configured to perform eSBR processing (e.g., using at least one of the eSBR tools known as harmonic transformation, pre-flattening, or InterTES) during decoding of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that does not contain eSBR metadata. An example of such a decoder is described with reference to FIG. 5.
[0095] The eSBR decoder (400) of Figure 5 includes, connected as shown, buffer memory 201 (which is identical to memory 201 of Figures 3 and 4), bitstream payload deformatter 215 (which is identical to deformatter 215 of Figure 4), audio decode subsystem 202 (sometimes referred to as the "core" decode stage or "core" decode subsystem, which is identical to core decode subsystem 202 of Figure 3), eSBR control data generation subsystem 401, and eSBR processing stage 203 (which is identical to stage 203 of Figure 3). Typically, decoder 400 also includes other processing elements (not shown).
[0096] In operation of the decoder 400 , a sequence of blocks of an encoded audio bitstream (MPEG-4 AAC bitstream) received by the decoder 400 is presented from the buffer 201 to the deformatter 215 .
[0097] Deformatter 215 is coupled and configured to demultiplex each block of the bitstream and extract therefrom SBR metadata (including quantized envelope data), and typically other metadata as well. Deformatter 215 is configured to provide at least said SBR metadata to eSBR processing stage 203. Deformatter 215 is also coupled and configured to extract audio data from each block of the bitstream and provide the extracted audio data to decoding subsystem (decoding stage) 202.
[0098] The audio decoding subsystem 202 of the decoder 400 is configured to decode the audio data extracted by the deformatter 215 (such decoding may be referred to as a "core" decoding operation) to generate decoded audio data and present the decoded audio data to the eSBR processing stage 203. The decoding is performed in the frequency domain. Typically, the final stage of processing in subsystem 202 applies a frequency-domain to time-domain transform to the decoded frequency-domain audio data, so that the output of the subsystem is time-domain decoded audio data. Stage 203 is configured to apply the SBR tools (and eSBR tools) indicated by the SBR metadata (extracted by the deformatter 215) and the eSBR metadata generated in subsystem 401 to the decoded audio data (i.e., perform SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata) to generate fully decoded audio data that is output from the decoder 400. Typically, decoder 400 includes a memory (accessible by subsystem 202 and stage 203) that stores deformatted audio data and metadata output from deformatter 215 (and optionally subsystem 401), with stage 203 configured to access the audio data and metadata as needed during SBR and eSBR processing. The SBR processing in stage 203 may be considered post-processing on the output of core decode subsystem 202. Optionally, decoder 400 also includes a final upmix subsystem (which may apply parametric stereo ("PS") tools defined in the MPEG-4 AAC standard using the PS metadata extracted by deformatter 215). The upmix subsystem is coupled and configured to perform an upmix on the output of stage 203 to generate fully decoded, upmixed audio, which is output from APU 210.
[0099] 5 is coupled and configured to detect at least one attribute of the encoded audio bitstream to be decoded and to generate eSBR control data (which may be or include eSBR metadata of any type included in the encoded audio bitstream, according to other embodiments of the present invention) in response to at least one result of the detection step. The eSBR control data is submitted to stage 203 to trigger and / or control the application of individual eSBR tools or combinations of eSBR tools upon detection of a particular attribute (or combination of attributes) of the bitstream. For example, to control the execution of eSBR processing using harmonic transformation, some embodiments of the control data generation subsystem 401 will include: a music detector (e.g., a simplified version of a conventional music detector) for setting the sbrPatchingMode[ch] parameter in response to detecting whether the bitstream indicates music (and providing the set parameter to stage 203); a transient detector for setting the sbrOversamplingFlag[ch] parameter in response to detecting the presence or absence of transient components in the audio content indicated by the bitstream (and providing the set parameter to stage 203); and / or a pitch detector for setting the sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch] parameters in response to detecting the pitch of the audio content indicated by the bitstream (and providing the set parameter to stage 203). Another aspect of the present invention is an audio bitstream decoding method performed by any of the embodiments of the inventive decoder described in this and the previous paragraphs.
[0100] Aspects of the present invention include encoding or decoding methods of the type that any embodiment of an APU, system, or device of the present invention is configured (e.g., programmed) to perform. Other aspects of the present invention include systems or devices configured (e.g., programmed) to perform any embodiment of the method of the present invention, as well as computer-readable media (e.g., disks) storing (e.g., non-transitory) code for implementing any embodiment of the method of the present invention or steps thereof. For example, a system of the present invention may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed and / or otherwise configured with software or firmware to perform any of a variety of operations on data, including embodiments of the method of the present invention, or steps thereof. Such a general-purpose processor may be or include a computer system, including input devices, memory, and processing circuitry, programmed (and / or otherwise configured) to perform embodiments of the method of the present invention (or steps thereof) in response to data presented to it.
[0101] Embodiments of the present invention may be implemented in hardware, firmware, or software, or a combination of both (e.g., as a programmable logic array). Unless otherwise specified, the algorithms or processes included as part of the present invention are not inherently related to any particular computer or other apparatus. In particular, various general-purpose machines may be used with programs written in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus (e.g., integrated circuits) to perform the required method steps. Thus, the present invention may be implemented in one or more computer programs running on one or more programmable computer systems (e.g., an implementation of any of the elements of FIG. 1 or the encoder 100 (or an element thereof) of FIG. 2 or the decoder 200 (or an element thereof) of FIG. 3 or the decoder 210 (or an element thereof) of FIG. 4 or the decoder 400 (or an element thereof) of FIG. 5). Each computer system has at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices, in known fashion.
[0102] Each such program may be implemented in any desired computer language to communicate with a computer system, including machine language, assembly language, or a high-level procedural, logical, or object-oriented programming language, and in any case, the language may be a compiled or interpreted language.
[0103] For example, when implemented by computer software instruction sequences, various functions and steps of embodiments of the present invention may be implemented by multi-threaded software instruction sequences running on suitable digital signal processing hardware, in which case various units, steps and functions of the embodiments may correspond to portions of the software instructions.
[0104] Each such computer system is preferably stored on or downloaded to a general-purpose or special-purpose programmable computer-readable storage medium or device (e.g., semiconductor memory or media, or magnetic or optical media) that, when read by the computer system, configures and operates the computer to perform the procedures described herein. The system of the present invention may be implemented as a computer-readable storage medium configured with (i.e., having stored thereon) a computer program, where the configured storage medium causes the computer system to operate in a specific, predefined manner to perform the functions described herein.
[0105] Several embodiments of the present invention have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the invention. Numerous modifications and variations of the present invention are possible in light of the above teachings. It is understood that, within the scope of the appended claims, the invention may be practiced other than as specifically described herein. Any reference signs included in the claims are for illustrative purposes only and should not be used to interpret or limit the claims in any manner.
[0106] Several aspects will be described. [Aspect 1] a buffer configured to store at least one block of the encoded audio bitstream; a bitstream payload deformatter coupled to the buffer and configured to demultiplex at least a portion of the at least one block of the encoded audio bitstream; a decoding subsystem coupled to the bitstream payload deformatter and configured to decode at least a portion of the at least one block of the encoded audio bitstream, wherein the at least one block of the encoded audio bitstream comprises: The filler element includes an identifier indicating the beginning of the filler element and filler data following the identifier, the filler data comprising: and at least one flag identifying whether an enhanced spectral band replication process should be performed on the audio content of the at least one block of the encoded audio bitstream. Audio processing unit. [Aspect 2] 2. The audio processing unit of claim 1, wherein the filler data further includes enhanced spectral band replication metadata. Aspect 3 3. The audio processing unit of aspect 2, wherein the enhanced spectral band replication metadata does not include one or more parameters used for both spectral patching and harmonic transformation. Aspect 4 4. The audio processing unit of aspect 2 or 3, wherein the enhanced spectral band replication metadata does not include parameters for selecting between harmonic transformation and spectral patching. Aspect 5 5. The audio processing unit of claim 2, wherein the enhanced spectral band replication metadata includes at least one of: i) a parameter indicating whether to perform pre-flattening; ii) a parameter indicating whether to perform subband sample-to-sample temporal envelope shaping; and iii) a parameter indicating whether to perform signal-adaptive frequency-domain oversampling. Aspect 6 6. The audio processing unit of any one of aspects 2 to 5, wherein the enhanced spectral band replication metadata is metadata configured to enable at least one eSBR tool that is described or mentioned in the MPEG USAC standard and not described or mentioned in the MPEG-4 AAC standard. Aspect 7 7. The audio processing unit of any one of aspects 1 to 6, wherein the at least one block of the encoded audio bitstream includes spectral band replication metadata. Aspect 8 An audio processing unit according to aspect 7 when citing aspect 2, wherein the enhanced spectral band replication metadata does not include parameters equivalent to parameters of the spectral band replication metadata. Aspect 9 9. The audio processing unit of claim 7 or 8, wherein the spectral band replication metadata is metadata configured to enable at least one SBR tool described or referenced in the MPEG-4 AAC standard. Aspect 10 10. The audio processing unit of any one of aspects 7 to 9, wherein the spectral band replication metadata includes one or more parameters used for both spectral patching and harmonic transformation. Aspect 11 11. The audio processing unit of any one of aspects 1 to 10, wherein the enhanced spectral band replication processing includes harmonic transformation but does not include spectral patching. Aspect 12 12. The audio processing unit of any one of aspects 1 to 11, wherein one value of the at least one flag indicates that the enhanced spectral band replication process should be performed on the audio content of the at least one block of the encoded audio bitstream, and another value of the at least one flag indicates that a basic spectral band replication process should be performed on the audio content of the at least one block of the encoded audio bitstream. Aspect 13 13. The audio processing unit of claim 12, wherein the basic spectral band replication process includes spectral patching but does not include harmonic transformation. Aspect 14 14. The audio processing unit according to claim 12 or 13, wherein the basic spectral band replication process is a spectral band replication process using spectral patching as described in the MPEG-4 AAC standard. Aspect 15 15. The audio processing unit of any one of aspects 1 to 14, wherein the enhanced spectral band replication process is a spectral band replication process that uses at least one eSBR tool that is described or mentioned in the MPEG USAC standard and not described or mentioned in the MPEG-4 AAC standard. Aspect 16 16. The audio processing unit of any one of aspects 1 to 15, wherein the audio processing unit is an audio decoder and the identifier is a three-bit unsigned integer having a value of 0x6, with the most significant bit transmitted first. Aspect 17 the filler data comprises an extended payload, the extended payload comprising spectral band replication extension data, the extended payload being identified using a four-bit unsigned integer transmitted most significant bit first, having a value of "1101" or "1110", and optionally The spectral band replication extension data comprises: an optional spectrum band replication header; spectral band replication data following said header; and a spectral band replication extension element following the spectral band replication data, wherein the flag is included in the spectral band replication extension element; 17. The audio processing unit of any one of aspects 1 to 16. Aspect 18 18. An audio processing unit as described in any one of aspects 1 to 17, wherein the at least one block of the encoded audio bitstream includes a first filler element and a second filler element, the first filler element including spectral band replication data, and the second filler element including the flag but not spectral band replication data. Aspect 19 19. The audio processing unit of any one of aspects 1 to 18, further comprising an enhanced spectral band replication processing subsystem configured to perform enhanced spectral band replication processing using or in response to the at least one flag. Aspect 20 A method of decoding an encoded audio bitstream, comprising: receiving at least one block of an encoded audio bitstream; demultiplexing at least a portion of the at least one block of the encoded audio bitstream; and decoding at least a portion of the at least one block of the encoded audio bitstream; wherein said at least one block of said encoded audio bitstream: The filler element includes an identifier indicating the beginning of the filler element and filler data following the identifier, the filler data comprising: and at least one flag identifying whether an enhanced spectral band replication process should be performed on the audio content of the at least one block of the encoded audio bitstream. method. Aspect 21 21. The method of embodiment 20, wherein the identifier is a three-bit unsigned integer transmitted most significant bit first, having a value of 0x6. Aspect 22 the filler data comprises an extended payload, the extended payload comprising spectral band replication extension data, the extended payload being identified using a four-bit unsigned integer transmitted most significant bit first, having a value of "1101" or "1110", and optionally The spectral band replication extension data comprises: an optional spectrum band replication header; spectral band replication data following said header; and a spectral band replication extension element following the spectral band replication data, wherein the flag is included in the spectral band replication extension element; 22. The method of embodiment 20 or 21. Aspect 23 23. The method of any one of aspects 20 to 22, wherein the enhanced spectral band replication process is a harmonic transformation, and one value of the at least one flag indicates that the enhanced spectral band replication process should be performed on the audio content of the at least one block of the encoded audio bitstream, and another value of the at least one flag indicates that spectral patching should be performed on the audio content of the at least one block of the encoded audio bitstream but that the harmonic transformation should not be performed. Aspect 24 the spectral band replication extension element includes enhanced spectral band replication metadata other than the flag, the enhanced spectral band replication metadata including a parameter indicating whether to perform pre-flattening, or the spectral band replication extension element includes enhanced spectral band replication metadata other than the flag, and the enhanced spectral band replication metadata includes a parameter indicating whether to perform subband sample-to-sample temporal envelope shaping. 24. The method of embodiment 22 or 23. Aspect 25 25. The method of any one of aspects 20-24, further comprising: performing an enhanced spectral band replication process using the at least one flag, the enhanced spectral band replication comprising harmonic translation. Aspect 26 The method of any one of aspects 20 to 25 or the audio processing unit of any one of aspects 1 to 19, wherein the encoded audio bitstream is an MPEG-4 AAC bitstream.
Claims
1. a bitstream payload deformatter configured to demultiplex blocks of the encoded audio bitstream; a decoding subsystem coupled to the bitstream payload deformatter and configured to decode at least a portion of the blocks of the encoded audio bitstream, wherein the blocks of the encoded audio bitstream comprise: a filler element having an identifier indicating the beginning of the filler element and filler data following the identifier, the filler data comprising an extended payload, the extended payload comprising spectral band replication extension data, the extended payload being identified using a four-bit unsigned integer transmitted most significant bit first having a value of "1101" or "1110", and the spectral band replication extension data being: including a spectral band replication header, spectral band replication data; the extended payload includes a spectrum band duplication extension element following the spectrum duplication data, and the flag is included in the spectrum band duplication extension element; The filling data further comprises: at least one flag identifying whether an enhanced spectral band replication process should be performed on the audio content of the block of the encoded audio bitstream; and enhanced spectral band replication metadata that does not include one or more parameters used for both spectral patching and harmonic transformation, wherein the enhanced spectral band replication metadata is metadata configured to enable at least one eSBR tool described or mentioned in the MPEG USAC standard and not described or mentioned in the MPEG-4 AAC standard. Audio processing unit.
2. 2. The audio processing unit of claim 1, wherein the encoded audio bitstream is an MPEG-4 AAC bitstream.
3. 2. The audio processing unit of claim 1, wherein the identifier is a 3-bit unsigned integer transmitted most significant bit first and has a value of 0x6.
4. 1. A method for decoding an encoded audio bitstream, the method comprising: demultiplexing blocks of the encoded audio bitstream; and decoding at least a portion of the blocks of the encoded audio bitstream; The blocks of the encoded audio bitstream: a filler element having an identifier indicating the beginning of the filler element and filler data following the identifier, the filler data comprising an extended payload, the extended payload comprising spectral band replication extension data, the extended payload being identified using a four-bit unsigned integer transmitted most significant bit first having a value of "1101" or "1110", and the spectral band replication extension data being: including a spectral band replication header, spectral band replication data; the extension payload further includes a spectrum band duplication extension element, and the flag is included in the spectrum band duplication extension element; The filling data further comprises: a flag identifying whether an enhanced spectral band replication process should be performed on the audio content of the block of the encoded audio bitstream; and enhanced spectral band replication metadata that does not include one or more parameters used for both spectral patching and harmonic transformation, wherein the enhanced spectral band replication metadata is metadata configured to enable at least one eSBR tool described or mentioned in the MPEG USAC standard and not described or mentioned in the MPEG-4 AAC standard. method.
5. 5. The method of claim 4, wherein the identifier is a 3-bit unsigned integer transmitted most significant bit first and has a value of 0x6.
6. The method of claim 5 , wherein the encoded audio bitstream is an MPEG-4 AAC bitstream.
Citation Information
Patent Citations
IEC14496-3