Decoding of an audio bit stream using enhanced spectral band replication metadata within at least one fill element
By integrating eSBR metadata in audio bitstreams, the method addresses suboptimal spectral band replication for music content, improving audio quality and reducing decoder complexity with minimal bitrate increase.
Patent Information
- Application Number
- JP2025011797
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2015-03-16
- Filing Date
- 2025-01-28
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2036-03-10
AI Technical Summary
Spectral band replication techniques in audio encoding, such as MPEG-4 AAC, are not optimal for certain types of audio content, particularly music with low crossover frequencies, necessitating improved methods for enhanced spectral band replication.
Incorporation of enhanced spectral band replication (eSBR) metadata in audio bitstreams, including flags and control data within fill elements, to guide decoders in performing eSBR processing, allowing for improved high-frequency regeneration using tools like harmonic transposition and inverse filtering.
Enables efficient and backward-compatible eSBR processing, enhancing audio quality at low data rates with minimal bitrate impact, reducing complexity and memory requirements in decoders.
Smart Images

Figure 0007712502000004 
Figure 0007712502000005 
Figure 0007712502000006
Abstract
Description
Technical Field
[0001] The present invention relates to audio signal processing. Some embodiments relate to the encoding and decoding of audio bitstreams (e.g., bitstreams having the MPEG-4 AAC format). Other embodiments relate to the decoding of such bitstreams by legacy decoders configured not to perform eSBR processing and to ignore such metadata, or to the decoding of audio bitstreams that do not include such metadata, which includes generating eSBR control data in response to the bitstream.
Background Art
[0002] A typical audio bitstream includes both audio data (e.g., encoded audio data) indicating one or more channels of audio content and metadata indicating at least one characteristic of the audio data or audio content. One well-known format for generating an encoded audio bitstream is the MPEG-4 Advanced Audio Coding (AAC) format described in the MPEG standard ISO / IEC 14496-3:2009. In the MPEG-4 standard, AAC represents "advanced audio coding" and HE-AAC represents "high-efficiency advanced audio coding".
[0003] The MPEG-4 AAC standard defines several audio profiles that determine which objects and coding tools are present in an encoder or decoder compliant with those audio profiles. Three of these audio profiles are: (1) the AAC profile, (2) the HE-AAC profile, and (3) the HE-AAC v2 profile. The AAC profile includes the AAC low complexity (or "AAC-LC") object type. The AAC-LC object type corresponds to the MPEG-2 AAC low complexity profile with some adjustments and does not include the spectral band replication ("SBR") object type nor the parametric stereo ("PS") object type. The HE-AAC profile is a superset of the AAC profile and additionally includes the SBR object type. The HE-AAC v2 profile is a superset of the HE-AAC profile and additionally includes the PS object type.
[0004] The SBR object type includes a spectral band replication tool. This is an important encoding tool that significantly improves the compression efficiency of perceptual audio coders. SBR reconstructs the high-frequency components of the audio signal on the receiver side (e.g., in the decoder). Therefore, the encoder only needs to encode and transmit the low-frequency components, allowing much higher audio quality at low data rates. SBR is based on replicating a sequence of harmonics that were previously truncated to reduce the data rate, from the available bandwidth-limited signal and control data obtained from the encoder. The ratio between the tone-like components and the noise-like components is maintained by adaptive inverse filtering and optional addition of noise and sine waves. In the MPEG-4 AAC standard, the SBR tool performs spectral patching where several adjacent quadrature mirror filter (QMF) subbands are copied from the transmitted low-frequency part of the audio signal to the high-frequency part of the audio signal generated in the decoder.
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] Spectral patching may not be ideal for certain types of audio, such as music content with a relatively low crossover frequency. Therefore, techniques for improving spectral band replication are needed.
Means for Solving the Problems
[0007] A first class of embodiments relates to an audio processing unit including a memory, a bitstream payload formatter, and a decode subsystem. The memory is configured to store at least one block of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream). The bitstream payload formatter is configured to demultiplex the encoded audio block. The decode subsystem is configured to decode the audio content of the encoded audio block. The encoded audio block includes a fill element. The fill element has an identifier indicating the start of the fill element and fill data following the identifier. The fill data includes at least one flag for identifying whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of the encoded audio block.
[0008] A second class of embodiments relates to a method for decoding an encoded audio bitstream. The method includes receiving at least one block of the encoded audio bitstream, demultiplexing at least some portions of the at least one block of the encoded audio bitstream, and decoding at least some portions of the at least one block of the encoded audio bitstream. The at least one block of the encoded audio bitstream includes a fill element. The fill element has an identifier indicating the start of the fill element and fill data following the identifier. The fill data includes at least one flag for identifying whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of the at least one block of the encoded audio bitstream.
[0009] Other embodiments relate to encoding and transcoding an audio bitstream that includes metadata identifying whether an enhanced spectral band replication (eSBR) process should be performed. BRIEF DESCRIPTION OF THE DRAWINGS
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
[0011] Throughout the present disclosure including the claims, the expression of performing an operation "on" a signal or data (such as filtering, scaling, converting, or applying a gain to the signal or data) is used in a broad sense to represent performing the operation directly on the signal or data, or on a processed version of the signal or data (such as a version of the signal that has undergone preliminary filtering or preprocessing prior to performing the operation).
[0012] Throughout the present disclosure including the claims, the expression "audio processing unit" is used in a broad sense to represent a system, device, or apparatus configured to process audio data. Examples of audio processing units include, but are not limited to, encoders (such as transcoders), decoders, codecs, preprocessing systems, postprocessing systems, and bitstream processing systems (sometimes referred to as bitstream processing tools). Virtually any consumer electronic device, such as a mobile phone, television, laptop, and tablet computer, includes an audio processing unit.
[0013] Throughout the present disclosure including the claims, the terms "coupling" or "coupled" are used in a broad sense to mean a direct or indirect connection. Thus, when a first device is coupled to a second device, the connection may be through a direct connection or through an indirect connection via other devices and connections. Further, components integrated within or together with other components are also coupled to each other.
[0014] 〈Detailed Description of Embodiments of the Present Invention〉 The MPEG-4 AAC standard contemplates that an encoded MPEG-4 AAC bitstream contains metadata indicating each type of SBR processing (if any) to be applied by a decoder to decode the audio content of the bitstream and / or controlling such SBR processing and / or indicating at least one characteristic or parameter of at least one SBR tool to be used to decode the audio content of the bitstream. Here, the expression "SBR metadata" is used to represent this type of metadata described or referred to in the MPEG-4 AAC standard.
[0015] The top level of an MPEG-4 AAC bitstream is a sequence of data blocks ("raw_data_block" elements), where each data block is a segment of data (referred to herein as a "block") containing audio data (typically over a time period of 1024 or 960 samples) and related information and / or other data. Here, the term "block" is used to represent a segment of an MPEG-4 AAC bitstream containing audio data (and corresponding metadata and optionally other related data) that determines or indicates one (and no more than one) "raw_data_block" element.
[0016] Each block of an MPEG-4 AAC bitstream can contain several syntax elements (each of which is also embodied as a segment of data in the bitstream). Seven types of such syntax elements are defined in the MPEG-4 AAC standard. Each syntax element is identified by a different value of the data element "id_syn_ele". Examples of syntax elements include "single_channel_element()", "channel_pair_element()", and "fill_element()". A single channel element is a container that includes audio data for a single audio channel (monophonic audio signal). A channel pair element includes audio data for two audio channels (i.e., a stereo audio signal).
[0017] A fill element is a container of information that includes an identifier (e.g., the value of the above element "id_syn_ele") and the subsequent data called "fill data". Fill elements have historically been used to adjust the instantaneous bitrate of a bitstream that is to be transmitted through a constant rate channel. By adding an appropriate amount of fill data to each block, a constant data rate can be achieved.
[0018] According to embodiments of the present invention, the fill data can include one or more extended payloads that extend the types of data (e.g., metadata) that can be transmitted in the bitstream. A decoder that receives a bitstream with fill data containing new types of data can optionally be used by the device (e.g., decoder) that receives the bitstream to extend the functionality of the device. Thus, as will be understood by those skilled in the art, a fill element is a special type of data structure that is different from the data structures typically used to transmit audio data (e.g., an audio payload including channel data).
[0019] In some embodiments of the present invention, the identifier used to identify a stuffing element may consist of a three-bit unsigned integer transmitted most significant bit first (「uimsbf」) having a value of 0x6. In one block, several instances of the same type of syntax element (e.g., several stuffing elements) may occur.
[0020] Another standard for encoding an audio bit stream is the MPEG Unified Speech and Audio Coding (USAC) standard (ISO / IEC 23003-3:2012). The MPEG USAC standard describes encoding and decoding audio content using spectral band replication processing (including the SBR processing described in the MPEG-4 AAC standard and other improved forms of spectral band replication processing). This process applies an extended and improved version of the spectral band replication tool (sometimes referred to herein as the "enhanced SBR tool" or "eSBR tool") of the set of SBR tools described in the MPEG-4 AAC standard. Thus, eSBR (defined in the USAC standard) is an improvement over SBR (defined in the MPEG-4 AAC standard).
[0021] In this document, the expression "enhanced SBR processing" (or "eSBR processing") is used to represent spectral band replication processing that uses at least one eSBR tool not described or referred to in MPEG-4 AAC (for example, at least one eSBR tool described or referred to in the MPEG USAC standard). Examples of such eSBR tools are harmonic transposition, QMF patching additional preprocessing or "pre-flattening", and subband sample-to-sample temporal envelope shaping or "inter-TES".
[0022] A bitstream generated according to the MPEG USAC standard (sometimes referred to as a USAC bitstream in this document) contains encoded audio content and typically includes metadata indicating the respective type of spectral band replication processing to be applied by a decoder to decode the audio content of the USAC bitstream and / or metadata indicating at least one characteristic or parameter of at least one SBR tool and / or eSBR tool to be used to control such spectral band replication processing and / or to decode the audio content of the USAC bitstream.
[0023] Here, the expression "enhanced SBR metadata" (or "eSBR metadata") is used to represent metadata that indicates each type of spectral band replication process to be applied by a decoder to decode the audio content of an encoded audio bitstream (e.g., a USAC bitstream) and / or controls such spectral band replication process and / or indicates at least one characteristic or parameter of at least one SBR tool and / or eSBR tool to be used to decode such audio content, which is not described or mentioned in the MPEG-4 AAC standard. An example of eSBR metadata is metadata that is described or mentioned in the MPEG USAC standard but not described or mentioned in the MPEG-4 AAC standard (for indicating or controlling spectral band replication process). Thus, the eSBR metadata in this document represents metadata that is not SBR metadata, and the SBR metadata in this document represents metadata that is not eSBR metadata.
[0024] The USAC bitstream may contain both SBR metadata and eSBR metadata. More specifically, the USAC bitstream may contain eSBR metadata that controls the execution of eSBR processing by a decoder and SBR metadata that controls the execution of SBR processing by a decoder. According to an exemplary embodiment of the present invention, eSBR metadata (e.g., configuration setting data specific to eSBR) is included in the MPEG-4 AAC bitstream (in accordance with the present invention) (e.g., in the sbr_extension() container at the end of the SBR payload).
[0025] During the decoding of an encoded bitstream using a set of eSBR tools (including at least one eSBR tool), the execution of eSBR processing by a decoder regenerates the high-frequency band of an audio signal based on the replication of a sequence of harmonics interrupted during encoding. Such eSBR processing typically adjusts the spectral envelope of the generated high-frequency band, applies inverse filtering, and adds noise and sine wave components to reproduce the spectral characteristics of the original audio signal.
[0026] According to an exemplary embodiment of the present invention, eSBR metadata (for example, a few control bits that are eSBR metadata) is included in one or more of the metadata segments of an encoded audio bitstream (for example, an MPEG-4 AAC bitstream). The encoded audio bitstream also includes encoded audio data in other segments (audio data segments). Typically, at least one such metadata segment of each block of the bitstream is a stuffing element (including an identifier indicating the start of the stuffing element) (or includes a stuffing element), and the eSBR metadata is included in the stuffing element after the identifier.
[0027] FIG. 1 is a block diagram of an exemplary audio processing chain (audio data processing system), and one or more of the elements of the system may be configured according to an embodiment of the present invention. The system includes the following elements coupled together as shown: an encoder 1, a delivery subsystem 2, a decoder 3, and a post-processing unit 4. In a variation of the illustrated system, one or more of the elements are omitted, or additional audio data processing units are included.
[0028] In some implementations, the encoder 1 (which optionally includes a preprocessing unit) is configured to receive as input PCM (time domain) samples including audio content and output an encoded audio bitstream (in a format compliant with the MPEG-4 AAC standard) representing the audio content. The data of the bitstream representing the audio content is sometimes referred to herein as "audio data" or "encoded audio data". When the encoder is configured according to a typical embodiment of the present invention, the audio bitstream output from the encoder includes eSBR metadata (typically other metadata as well) in addition to the audio data.
[0029] One or more encoded audio bitstreams output from the encoder 1 may be presented to the encoded audio delivery subsystem 2. The subsystem 2 is configured to store and / or deliver each encoded bitstream output from the encoder 1. The encoded audio bitstream output from the encoder 1 may be stored by the subsystem 2 (e.g., in the form of a DVD or Blu-ray disc), or may be transmitted by the subsystem 2 (which may implement a transmission link or network), or may be stored and transmitted by the subsystem 2.
[0030] Decoder 3 is configured to decode an encoded MPEG-4 AAC audio bitstream received (generated by Encoder 1) via Subsystem 2. In some embodiments, Decoder 3 extracts eSBR metadata from each block of the bitstream and decodes the bitstream (including by performing eSBR processing using the extracted eSBR metadata) to generate decoded audio data (e.g., a stream of decoded PCM audio samples). In some embodiments, Decoder 3 extracts SBR metadata from the bitstream (but ignores any eSBR metadata included in the bitstream) and decodes the bitstream (including by performing SBR processing using the extracted SBR metadata) to generate decoded audio data (e.g., a stream of decoded PCM audio samples). Typically, Decoder 3 includes a buffer for storing (e.g., in a non-transitory manner) segments of the encoded audio bitstream received from Subsystem 2.
[0031] The post-processing unit 4 of FIG. 1 is configured to receive a stream of decoded audio data (e.g., decoded PCM audio samples) from Decoder 3 and perform post-processing thereon. The post-processing unit may be configured to render the post-processed audio content (or the decoded audio received from Decoder 3) for playback by one or more speakers.
[0032] Figure 2 is a block diagram of an encoder (100) which is an embodiment of the audio processing unit of the present invention. Any of the components or elements of encoder 100 may be implemented in hardware, software, or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit). Encoder 100 has an encoder 105, a stuffing / formatter stage 107, a metadata generation stage 106, and a buffer memory 109 connected as shown in the figure. Typically, encoder 100 also includes other processing elements (not shown). Encoder 100 is configured to convert an input audio bitstream into an encoded output MPEG-4 AAC bitstream.
[0033] The metadata generator 106 is coupled and configured to generate (and / or pass through to stage 107) the metadata (including eSBR metadata and SBR metadata) to be included by stage 107 in the encoded bitstream to be output from encoder 100.
[0034] The encoder 105 is coupled and configured to encode the input audio data (e.g., by performing compression thereon) and present the resulting encoded audio to stage 107 for inclusion in the encoded bitstream to be output from stage 107.
[0035] Stage 107 is configured to multiplex the encoded audio from encoder 105 and the metadata (including eSBR metadata and SBR metadata) from generator 106 to generate the encoded bitstream to be output from stage 107. Preferably, the encoded bitstream has a format defined by one of the embodiments of the present invention.
[0036] Buffer memory 109 is configured to store (e.g., in a non-transitory manner) at least one block of the encoded audio bitstream output from stage 107. Thereafter, a sequence of blocks of the encoded audio bitstream is presented from buffer memory 109 to the delivery system as the output from encoder 100.
[0037] FIG. 3 is a block diagram of a system including a decoder (200), which is an embodiment of the audio processing unit of the present invention, and optionally a post-processor (300) coupled thereto. Any of the components or elements of decoder 200 may be implemented in hardware, software, or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit). Decoder 200 has a buffer memory 201, a bitstream payload demultiplexer (parser) 205, an audio decode subsystem 202 (sometimes referred to as the "core" decode stage or "core" decode subsystem), an eSBR processing stage 203, and a control bit generation stage 204, connected as shown in the figure. Typically, decoder 200 also includes other processing elements (not shown).
[0038] Buffer memory (buffer) 201 stores at least one block of the encoded MPEG-4 AAC audio bitstream received by decoder 200 (e.g., in a non-transitory manner). In the operation of decoder 200, a sequence of blocks of the bitstream is presented from buffer 201 to demultiplexer 205.
[0039] In a variation of the embodiment of FIG. 3 (or the embodiment of FIG. 4 described later), an APU that is not a decoder (e.g., the APU 500 of FIG. 6) includes a buffer memory (e.g., the same buffer memory as buffer 201) that stores (e.g., in a non-transitory manner) at least one block of an encoded audio bitstream of the same type as that received by buffer 201 of FIG. 3 or FIG. 4 (e.g., an MPEG-4 AAC audio bitstream), i.e., an encoded audio bitstream that includes eSBR metadata.
[0040] Referring again to FIG. 3, the demultiplexer 205 is coupled and configured to demultiplex each block of the bitstream and then extract SBR metadata (including quantized envelope data) and eSBR metadata (typically also other metadata), and present at least the eSBR metadata and the SBR metadata to the eSBR processing stage 203, and typically also present the other extracted metadata to the decode subsystem 202 (optionally also to the control bit generator 204). The demultiplexer 205 is also coupled and configured to extract audio data from each block of the bitstream and present the extracted audio data to the decode subsystem (decode stage) 202.
[0041] The system of FIG. 3 optionally also includes a post-processor 300. The post-processor 300 includes a buffer memory (buffer) 301 and other processing elements (not shown) including at least one processing element coupled to the buffer 301. The buffer 301 stores (e.g., in a non-transitory manner) at least one block (or frame) of the decoded audio data received by the post-processor 300 from the decoder 200. The processing elements of the post-processor 300 are coupled and configured to adaptively process a sequence of blocks (or frames) of the decoded audio output from the buffer 301, using the metadata output from the decode subsystem 202 and / or the format stripper 205 and / or the control bits output from stage 204 of the decoder 200.
[0042] The audio decoding subsystem 202 of decoder 200 decodes the audio data extracted by parser 205 (such decoding may be referred to as "core" decoding operation), generates the decoded audio data, and is configured to present the decoded audio data to eSBR processing stage 203. The decoding is performed in the frequency domain and typically includes inverse quantization followed by spectral processing. Typically, the final stage of processing in subsystem 202 applies a conversion from the frequency domain to the time domain to the decoded frequency domain audio data, so the output of the subsystem is the decoded audio data in the time domain. Stage 203 is configured to apply the SBR tools and eSBR tools indicated by the SBR metadata and eSBR metadata (extracted by parser 205) to the decoded audio data (i.e., perform SBR and eSBR processing on the output of decoding subsystem 202 using the SBR and eSBR metadata) to generate the fully decoded audio data output from decoder 200 (e.g., to post-processor 300). Typically, decoder 200 includes a memory (accessible by subsystem 202 and stage 203) that stores the demultiplexed audio data and metadata output from demultiplexer 205, and stage 203 is configured to access the audio data and metadata (including SBR metadata and eSBR metadata) as needed during SBR and eSBR processing. The SBR processing and eSBR processing in stage 203 may be considered as post-processing on the output of core decoding subsystem 202. Optionally, decoder 200 also includes a final upmixing subsystem (which can apply the parametric stereo ("PS") tools defined in the MPEG-4 AAC standard using the PS metadata extracted by demultiplexer 205 and / or the control bits generated in subsystem 204).The upmix subsystem is coupled and configured to perform an upmix on the output of stage 203 to produce fully decoded, upmixed audio output from decoder 200. Alternatively, post-processor 300 may be configured to perform an upmix on the output of decoder 200 (e.g., using PS metadata extracted by demultiplexer 205 and / or control bits generated in subsystem 204).
[0043] In response to metadata extracted by demultiplexer 205, control bit generator 204 may generate control data. The control data may be used within decoder 200 (e.g., in a final upmix subsystem), and / or presented as an output of decoder 200 (e.g., to post-processor 300 for use in post-processing). In response to metadata extracted from the input bitstream (optionally also in response to control data), stage 204 may generate control bits (presented to post-processor 300) indicating that the decoded audio data output from eSBR processing stage 203 should undergo a particular type of post-processing. In some implementations, decoder 200 is configured to present metadata extracted by demultiplexer 205 from the input bitstream to post-processor 300, and post-processor 300 is configured to perform post-processing on the decoded audio data output from decoder 200 using the metadata.
[0044] Figure 4 is a block diagram of an audio processing unit ("APU") (210) which is another embodiment of the audio processing unit of the present invention. APU 210 is a legacy decoder not configured to perform eSBR processing. Any of the components or elements of APU 210 may be implemented in hardware, software, or a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA or other integrated circuit). APU 210 has a buffer memory 201, a bitstream payload formatter (parser) 215, an audio decode subsystem 202 (sometimes referred to as the "core" decode stage or "core" decode subsystem), and an SBR processing stage 213, connected as shown in the figure. Typically, APU 210 also includes other processing elements (not shown).
[0045] Elements 201 and 202 of APU 210 are identical to the identically numbered elements of decoder 200 (of FIG. 3), and the above description thereof will not be repeated. In the operation of APU 210, a sequence of blocks of an encoded audio bitstream (MPEG-4 AAC bitstream) received by APU 210 is presented from buffer 201 to formatter 215.
[0046] Formatter 215 is configured to demultiplex each block of the bitstream and then extract SBR metadata (including quantized envelope data), typically other metadata as well, but ignore eSBR which may be included in the bitstream according to any embodiment of the present invention, and is coupled and configured to present at least said SBR metadata to SBR processing stage 213. Formatter 215 is also coupled and configured to extract audio data from each block of the bitstream and present the extracted audio data to decode subsystem (decode stage) 202.
[0047] The audio decoding subsystem 202 of decoder 200 decodes the audio data extracted by the demultiplexer 215 (such decoding may be referred to as "core" decoding operation), generates decoded audio data, and is configured to present the decoded audio data to the SBR processing stage 213. The decoding is performed in the frequency domain. Typically, the final stage of processing in subsystem 202 applies an inverse transform from the frequency domain to the time domain to the decoded frequency domain audio data, so the output of the subsystem is time domain decoded audio data. Stage 213 is configured to apply the SBR tools indicated by the SBR metadata (extracted by the demultiplexer 215) to the decoded audio data (but not the eSBR tools) (i.e., perform SBR processing on the output of the decoding subsystem 202 using the SBR metadata) to generate fully decoded audio data output from the APU 210 (e.g., to the post-processor 300). Typically, the APU 210 includes a memory (accessible by the subsystem 202 and stage 213) that stores the demultiplexed audio data and metadata output from the demultiplexer 215, and stage 213 is configured to access the audio data and metadata (including SBR metadata) as needed during SBR processing. The SBR processing in stage 213 may be considered a post-processing of the output of the core decoding subsystem 202. Optionally, the APU 210 also includes a final upmixing subsystem (which may apply parametric stereo ("PS") tools defined in the MPEG-4 AAC standard using the PS metadata extracted by the demultiplexer 215). The upmixing subsystem is coupled and configured to perform upmixing on the output of stage 213 to generate fully decoded, upmixed audio output from the APU 210.Alternatively, the post-processor is configured to perform upmixing on the output of the APU 210 (e.g., using the PS metadata extracted by the formatter 215 and / or the control bits generated in the APU 210).
[0048] Various implementations of the encoder 100, the decoder 200, and the APU 210 are configured to perform different embodiments of the method of the present invention.
[0049] According to some embodiments, a legacy decoder (configured not to parse eSBR metadata or use any eSBR tools related to eSBR metadata) ignores the eSBR metadata, but still decodes the bitstream as much as possible without using the eSBR metadata or any eSBR tools related to the eSBR metadata, typically without any significant penalty in the decoded audio quality. The eSBR metadata (e.g., a few control bits that are the eSBR metadata) is included in the encoded audio bitstream (e.g., an MPEG-4 AAC bitstream). However, an eSBR decoder configured to parse the bitstream to identify the eSBR metadata and use at least one eSBR tool in response to the eSBR metadata enjoys the benefits of using at least one such eSBR tool. Accordingly, embodiments of the present invention provide a means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner.
[0050] Typically, the eSBR metadata in the bitstream indicates one or more of the following eSBR tools (e.g., indicates at least one characteristic or parameter of one or more of the following eSBR tools), which are described in the MPEG USAC standard and may or may not be applied by the encoder during generation of the bitstream): · Harmonic conversion; ·Pre-processing for additional QMF patching (pre-flattening); and ·Temporal Envelope Shaping between sub-band samples or "inter TES". For example, the eSBR metadata included in the bitstream may indicate the values of the parameters: harmonicSBR[ch], sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch], bs_interTes, bs_temp_shape[ch][env], bs_inter_temp_shape_mode[ch][env] and bs_sbr_preprocessing (described in the MPEG USAC standard and this disclosure).
[0051] Here, assuming X is some parameter, the notation X[ch] indicates that the parameter relates to a channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, sometimes the expression [ch] is omitted, assuming that the relevant parameter relates to a channel of the audio content.
[0052] Here, assuming X is some parameter, the notation X[ch][env] indicates that the parameter relates to the SBR envelope ("env") of a channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, sometimes the expressions [env] and [ch] are omitted, assuming that the relevant parameter relates to the SBR envelope of a channel of the audio content.
[0053] As described above, MPEG USAC contemplates that the USAC bitstream includes eSBR metadata that controls the execution of eSBR processing by the decoder. The eSBR metadata includes the following one-bit metadata parameters: harmonicSBR; bs_interTES; and bs_pvc.
[0054] The parameter harmonicSBR indicates the use of harmonic patching (harmonic transposition) for SBR. Specifically, harmonicSBR = 0 indicates non-harmonic spectral patching as described in section 4.6.18.6.3 of the MPEG-4 AAC standard; harmonicSBR = 1 indicates harmonic SBR patching (of the type used in eSBR as described in sections 7.5.3 or 7.5.4 of the MPEG USAC standard). Harmonic SBR patching is not used for non-eSBR spectral band replication (i.e., SBR that is not eSBR). Throughout this disclosure, the basic form of spectral band replication is referred to as spectral patching, and the improved form of spectral band replication is referred to as harmonic transposition.
[0055] The value of the parameter bs_interTES indicates the use of the interTES tool for eSBR.
[0056] The value of the parameter bs_pvc indicates the use of the PVC tool for eSBR.
[0057] During the decoding of the encoded bitstream, the execution of harmonic transposition during the eSBR processing stage of decoding (for each channel "ch" of the audio content indicated by the bitstream) is controlled by the following eSBR metadata parameters: sbrPatchingMode[ch]; sbrOversamplingFlag[ch]; sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch].
[0058] The value of sbrPatchingMode[ch] indicates the type of transposer used in eSBR. sbrPatchingMode[ch]=1 indicates non-harmonic patching as described in section 4.6.18.6.3 of the MPEG-4 AAC standard; sbrPatchingMode[ch]=0 indicates harmonic SBR patching as described in sections 7.5.3 or 7.5.4 of the MPEG USAC standard.
[0059] The value of sbrOversamplingFlag[ch] indicates the use of signal-adaptive frequency-domain oversampling in eSBR combined with DFT-based harmonic SBR patching as described in section 7.5.3 of the MPEG USAC standard. This flag controls the size of the DFT used in the transposer. 1 indicates signal-adaptive frequency-domain oversampling enabled as described in section 7.5.3.1 of the MPEG USAC standard; 0 indicates signal-adaptive frequency-domain oversampling disabled as described in section 7.5.3.1 of the MPEG USAC standard.
[0060] The value of sbrPitchInBinsFlag[ch] controls the interpretation of the sbrPitchInBins[ch] parameter. 1 indicates that the value in sbrPitchInBins[ch] is valid and greater than 0; 0 indicates that the value of sbrPitchInBins[ch] is set to 0.
[0061] The value of sbrPitchInBins[ch] controls the addition of the cross-product terms in the SBR harmonic transposer. The value sbrPitchInBins[ch] is an integer value in the range [0,127] and represents the distance measured in the frequency bins of a 1536-line DFT operating on the sampling frequency of the core coder.
[0062] When an MPEG-4 AAC bitstream indicates SBR channel pairs where the channels are not combined (rather than a single SBR channel), the bitstream indicates two instances of the above syntax (for harmonic or non-harmonic conversion). One instance for each channel of the sbr_channel_pair_element().
[0063] Harmonic conversion of the eSBR tool typically improves the quality of the decoded music signal at relatively low crossover frequencies. Non-harmonic conversion (i.e., legacy spectral patching) typically improves the speech signal. Thus, the starting point in determining which type of conversion is preferred for encoding a particular audio content is to select the conversion method depending on speech / music detection. Here, harmonic conversion is used for music content and spectral patching is used for speech content.
[0064] The execution of pre-emphasis during eSBR processing is controlled by the value of a one-bit eSBR metadata parameter known as bs_sbr_preprocessing. It is in the sense that pre-emphasis is either executed or not depending on this single-bit value. When the SBR QMF patching algorithm described in section 4.6.18.6.3 of the MPEG-4 AAC standard is used, a pre-emphasis stage may be executed (when indicated by the bs_sbr_preprocessing parameter) to avoid discontinuities in the shape of the spectral envelope of the high-frequency signal from being input to a subsequent envelope adjuster (which performs another stage of the eSBR processing). Pre-emphasis typically improves the operation of the subsequent envelope adjustment stage and, as a result, the perceived high-frequency signal becomes more stable.
[0065] The execution of inter-subband sample Temporal Envelope Shaping (the “inter-TES” tool) during eSBR processing in a decoder is controlled by the following eSBR metadata parameters for each SBR envelope (“env”) of each channel (“ch”) of the audio content of the decoded USAC bitstream: bs_temp_shape[ch][env] and bs_inter_temp_shape_mode[ch][env].
[0066] The inter-TES tool processes the QMF subband samples after the envelope adjuster. This processing stage shapes the temporal envelope of the higher frequency bands with a finer temporal granularity than the temporal granularity of the envelope adjuster. The inter-TES shapes the temporal envelope among the QMF subband samples by applying a gain factor to each QMF subband sample in the SBR envelope.
[0067] The parameter bs_temp_shape[ch][env] is a flag indicating the use of inter-TES. The parameter bs_inter_temp_shape_mode[ch][env] indicates the value of the parameter γ in inter-TES (as defined in the MPEG USAC standard).
[0068] The overall bitrate requirement for including eSBR metadata indicating the above-described eSBR tools (harmonic conversion, pre-emphasis flattening, and inter-TES) in an MPEG-4 AAC bitstream is expected to be on the order of hundreds of bits per second. This is because, according to some embodiments of the present invention, only the differential control data required to perform the eSBR processing is transmitted. This information is included in a backward-compatible manner (as will be described later) so that legacy decoders can ignore this information. Therefore, the adverse impact on the bitrate associated with including eSBR metadata can be ignored for several reasons including the following: ·The bitrate penalty (due to including eSBR metadata) is a very small percentage of the total bitrate since only the differential control data needed to perform eSBR processing is transmitted (not simulcast of SBR control data); ·Tuning of control information related to SBR typically does not depend on the details of the conversion; and ·The inter-TES tool (used during eSBR processing) performs single-ended post-processing of the converted signal.
[0069] Thus, embodiments of the present invention provide means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner. This efficient transmission of eSBR control data reduces the memory requirements in decoders, encoders, and transcoders that use aspects of the present invention without a clear adverse impact on the bitrate. Further, the complexity and processing requirements associated with performing eSBR according to embodiments of the present invention are also reduced. This is because the SBR data only needs to be processed once and does not need to be simulcast as it would if eSBR were integrated into the MPEG-4 AAC codec in a backward-compatible manner but rather is treated as a completely separate object type in MPEG-4 AAC.
[0070] Next, referring to FIG. 7, elements of a block (raw_data_block) of an MPEG-4 AAC bitstream that includes eSBR metadata according to some embodiments of the present invention are described. FIG. 7 is a diagram of a block (raw_data_block) of an MPEG-4 AAC bitstream, showing some of its segments.
[0071] A block of an MPEG-4 AAC bitstream may include at least one single_channel_element() (e.g., the single channel element shown in FIG. 7) and / or at least one channel_pair_element() (not specifically shown in FIG. 7 but may exist) that contain audio data for an audio program. The block may also include several fill_element (e.g., fill element 1 and / or fill element 2 in FIG. 7) that contain data related to the program (e.g., metadata). Each single_channel_element() includes an identifier (e.g., "ID1" in FIG. 7) indicating the start of the single channel element and can contain audio data representing different channels of a multi-channel audio program. Each channel_pair_element includes an identifier (not shown in FIG. 7) indicating the start of the channel pair element and can contain audio data representing two channels of the program.
[0072] A fill_element (referred to as a fill element in this document) of an MPEG-4 AAC bitstream includes an identifier (e.g., "ID2" in FIG. 7) indicating the start of the fill element, followed by fill data. The identifier ID2 may consist of a three-bit unsigned integer (a "uimsbf") with the value of 0x6, where the most significant bit is transmitted first. The fill data can include an extension_payload() element (sometimes referred to as an extended payload in this document). Its syntax is shown in Table 4.57 of the MPEG-4 AAC standard. There are several types of extended payloads, identified through the extension_type parameter. This parameter is a four-bit unsigned integer (a "uimsbf") where the most significant bit is transmitted first.
[0073] The filler data (e.g., its extended payload) can include a header or identifier (e.g., "Header 1" in FIG. 7) indicating a segment of the filler data that represents an SBR object (i.e., the header initializes an "SBR object" type called sbr_extension_data() in the MPEG-4 AAC standard). For example, a spectral band replication (SBR) extended payload is identified by having a value of "1101" or "1110" for the extension_type field in the header, where the identifier "1101" identifies an extended payload using SBR data, and "1110" identifies an extended payload using SBR data with a cyclic redundancy check (CRC) for verifying the correctness of the SBR data.
[0074] When the header (e.g., the extension_type field) initializes an SBR object type, SBR metadata (sometimes referred to herein as "spectral band replication data" and called sbr_data() in the MPEG-4 AAC standard) follows the header, and at least one spectral band replication extension element (e.g., the "SBR extension element" of filler element 1 in FIG. 7) can follow the SBR metadata. Such a spectral band replication extension element (a segment of the bitstream) is called an sbr_extension() container in the MPEG-4 AAC standard. The spectral band replication extension element optionally includes a header (e.g., the "SBR extension header" of filler element 1 in FIG. 7).
[0075] The MPEG-4 AAC standard contemplates that the spectral band replication extension element can include PS (parametric stereo) data for the audio data of the program. The MPEG-4 AAC standard contemplates that when the header of the stuffing element (e.g., of its extended payload) initializes the SBR object type (like "header 1" in FIG. 7) and the spectral band replication extension element of the stuffing element includes PS data, the stuffing element (e.g., its extended payload) includes the spectral band replication data bs_extension_id parameter. The value of this parameter (i.e., bs_extension_id = 2) indicates that PS data is included in the spectral band replication extension element of the stuffing element.
[0076] According to some embodiments of the present invention, eSBR metadata (e.g., a flag indicating whether enhanced spectral band replication (eSBR) processing is to be performed on the audio content of that block) is included in the spectral band replication extension element of the stuffing element. For example, such a flag is included in stuffing element 1 of FIG. 7, and the flag appears after the header of the "SBR extension element" of stuffing element 1 (the "SBR extension header" of stuffing element 1). Optionally, such a flag and additional eSBR metadata are included in the spectral band replication extension element after the header of the spectral band replication extension element (e.g., in the SBR extension element of stuffing element 1 in FIG. 7, after the SBR extension header). According to some embodiments of the present invention, a stuffing element containing eSBR metadata also includes the bs_extension_id parameter. The value of this parameter (e.g., bs_extension_id = 3) indicates that the stuffing element includes eSBR metadata and that eSBR processing should be performed on the audio content of that block.
[0077] According to some embodiments of the present invention, eSBR metadata is included in padding elements of the MPEG-4 AAC bitstream other than the spectral band replication extension element (SBR extension element) of the padding element (e.g., padding element 2 in FIG. 7). This is because padding elements containing SBR data or extension_payload() with SBR data having CRC do not contain any other extension payloads of any other extended type. Therefore, in embodiments where eSBR metadata is stored in its own extension payload, a separate padding element is used to store the eSBR metadata. Such a padding element includes an identifier indicating the start of the padding element (e.g., "ID2" in FIG. 7), followed by padding data. The padding data can include an extension_payload() element (sometimes referred to as an extension payload in this document). Its syntax is shown in Table 4.57 of the MPEG-4 AAC standard. The padding data (e.g., its extension payload) can include a header indicating an eSBR object (e.g., "header 2" of padding element 2 in FIG. 7) (i.e., the header initializes the enhanced spectral band replication (eSBR) object type), and the padding data (e.g., its extension payload) includes eSBR metadata after the header. For example, padding element 2 in FIG. 7 includes such a header ("header 2"), and after the header, it also includes eSBR metadata (i.e., a "flag" within padding element 2 indicating whether enhanced spectral band replication (eSBR) processing is performed on the audio content of that block). Optionally, additional eSBR metadata is also included in the padding data of padding element 2 in FIG. 7 after header 2. In the embodiments described in this paragraph, the header (e.g., header 2 in FIG. 7) has an identification information value indicating an eSBR extension payload rather than one of the normal values specified in Table 4.57 of the MPEG-4 AAC standard (thus, the extension_type field of the header indicates that the padding data contains eSBR metadata).
[0078] In a first class of embodiments, the present invention is an audio processing unit (e.g., a decoder) that: A memory (e.g., buffer 201 of FIG. 3 or FIG. 4) configured to store at least one block of an encoded audio bitstream (e.g., at least one block of an MPEG-4 AAC bitstream); A bitstream payload format demultiplexer (e.g., element 205 of FIG. 3 or element 215 of FIG. 4) coupled to the memory and configured to demultiplex at least a portion of the block of the bitstream; A decode subsystem (e.g., elements 202 and 203 of FIG. 3 or elements 202 and 213 of FIG. 4) coupled and configured to decode at least one portion of the audio content of the block of the bitstream, wherein the block Includes padding elements, an identifier indicating the start of the padding elements (e.g., an id_syn_ele identifier having a value of 0x6 in Table 4.85 of the MPEG-4 AAC standard), and padding data after the identifier, and the padding data: Includes at least one flag for identifying whether an enhanced spectral band replication (eSBR) process should be performed on the audio content of the block (e.g., using the spectral band replication data and eSBR metadata included in the block); Is an audio processing unit.
[0079] The flag is eSBR metadata, and examples of the flag are the sbrPatchingMode flag. Another example of the flag is the harmonicSBR flag. All of these flags indicate whether basic form spectral band replication or enhanced form spectral replication should be performed on the audio data of the block. Basic form spectral replication is spectral patching, and enhanced form spectral band replication is harmonic conversion.
[0080] In some embodiments, the padding data also includes additional eSBR metadata (i.e., eSBR metadata other than the flag).
[0081] The memory may be a buffer memory (e.g., the implementation of buffer 201 in FIG. 4) that stores (e.g., in a non-transitory manner) the at least one block of the encoded audio bitstream.
[0082] The complexity of the execution of eSBR processing (using eSBR harmonic conversion, pre-emphasis flattening, and inter-TES tools) by an eSBR decoder during the decoding of an MPEG-4 AAC bitstream containing eSBR metadata (where the eSBR metadata indicates these eSBR tools) is estimated to be as follows for a typical decoding using the parameters shown: ● Harmonic conversion (16 kbps, 14400 / 28800 Hz) ○ DFT-based: 3.68 WMOPS (weighted million operations per second); ○ WMF-based: 0.98 WMOPS; ● QMF patching pre-processing (pre-emphasis flattening): 0.1 WMOPS; ● Sub-band sample-to-sample temporal envelope shaping (inter-TES): at most 0.16 WMOPS For transients, it has been found that DFT-based conversion typically exhibits better performance than QMF-based conversion.
[0083] According to some embodiments of the present invention, a stuffing element (of an encoded audio bitstream) containing eSBR metadata includes a parameter (e.g., the bs_extension_id parameter) having a value (e.g., bs_extension_id = 3) indicating that the eSBR metadata is included in the stuffing element and that eSBR processing should be performed on the audio content of the block, and / or a parameter (e.g., the same bs_extension_id parameter) having a value (e.g., bs_extension_id = 2) indicating that the sbr_extension() container of the stuffing element contains PS data. For example, as shown in Table 1 below, such a parameter having the value bs_extension_id = 2 may indicate that the sbr_extension() container of the stuffing element contains PS data, and such a parameter having the value bs_extension_id = 3 may indicate that the sbr_extension() container of the stuffing element contains eSBR metadata.
[0084]
Table 1
[0085]
Table 2
[0086] For example, in some embodiments, esbr_data() may have the syntax shown in Table 3 to indicate these metadata parameters.
[0087]
Table 3
[0088] For example, an MPEG-4 HE-AAC or HE-AAC-v2 compliant decoder may be extended to include an improved form of spectral band replication such as harmonic conversion. This improved form of spectral band replication is in addition to the basic form of spectral band replication already supported by the decoder. In the context of an MPEG-4 HE-AAC or HE-AAC-v2 compliant decoder, this basic form of spectral band replication is the QMF spectral patching SBR tool defined in section 4.6.18 of the MPEG-4 AAC standard.
[0089] When performing the improved form of spectral band replication, the extended HE-AAC decoder may reuse many of the bitstream parameters already included in the SBR extension payload of the bitstream. Specific parameters that may be reused include, for example, various parameters that determine the master frequency band table. These parameters include bs_start_freq (a parameter that determines the start of the master frequency table parameter), bs_stop_freq (a parameter that determines the end of the master frequency table), bs_freq_scale (a parameter that determines the number of frequency bands per octave), and bs_alter_scale (a parameter that changes the scale of the frequency bands). The parameters that may be reused also include the parameter (bs_noise_bands) that determines the noise band table and the limiter band table parameter (bs_limiter_bands). Thus, in various embodiments, at least some of the equivalent parameters specified in the USAC standard are omitted from the bitstream, thereby reducing the control overhead in the bitstream. Typically, when a parameter specified in the AAC standard has an equivalent parameter specified in the USAC standard, the equivalent parameter specified in the USAC standard has the same name as the parameter specified in the AAC standard. For example, the envelope scale factor E OrigMappedHowever, the equivalent parameters specified in the USAC standard typically have different values that are "tuned" for the enhanced SBR processing defined in the USAC standard rather than for the SBR processing defined in the AAC standard.
[0090] In addition to the many parameters described above, other data elements may also be reused by the enhanced HE-AAC decoder when performing an improved form of spectral band replication according to embodiments of the present invention. For example, envelope data and noise floor data may be extracted from the bs_data_env and bs_noise_env data and used during the improved form of spectral band replication.
[0091] In essence, these embodiments utilize the configuration parameters and envelope data already supported by a legacy HE-AAC or HE-AAC v2 decoder in the SBR extension payload to enable an improved form of spectral band replication with as little additional transmission data as possible. Thus, an enhanced decoder that supports an improved form of spectral band replication can be generated in a very efficient manner by relying on already defined bitstream elements (such as those within the SBR extension payload) and adding only the parameters required to support the improved form of spectral band replication (within the filler element extension payload). This data reduction feature, combined with placing the newly added parameters in a reserved data field such as an extension container, substantially reduces the barrier to creating a decoder that supports an improved form of spectral band replication by ensuring that the bitstream is backward compatible with legacy decoders that do not support an improved form of spectral band replication.
[0092] In Table 3, the numbers in the central column indicate the number of bits of the corresponding parameter in the left column.
[0093] In some embodiments, the present invention is a method that includes encoding audio data to generate an encoded bitstream (e.g., an MPEG-4 AAC bitstream). The generation includes including eSBR metadata in at least one segment of at least one block of the encoded bitstream and including the audio data in at least one other segment of the at least one block. In a typical embodiment, the method includes multiplexing the audio data with the eSBR metadata in each block of the encoded bitstream. In a typical decoding of the encoded bitstream in an eSBR decoder, the decoder extracts the eSBR metadata from the bitstream (which includes parsing and demultiplexing the eSBR metadata and the audio data), and uses the eSBR metadata to process the audio data to generate a stream of decoded audio data.
[0094] Another aspect of the present invention is an eSBR decoder configured to perform eSBR processing (e.g., using at least one of harmonic conversion, pre-emphasis flattening, or an eSBR tool known as Inter-TES) during the decoding of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that does not include eSBR metadata. An example of such a decoder is described with reference to FIG. 5.
[0095] The eSBR decoder (400) of FIG. 5 includes a buffer memory 201 (which is the same as the memory 201 of FIGS. 3 and 4), a bitstream payload formatter 215 (which is the same as the formatter 215 of FIG. 4), an audio decoding subsystem 202 (sometimes referred to as the "core" decoding stage or "core" decoding subsystem, which is the same as the core decoding subsystem 202 of FIG. 3), an eSBR control data generation subsystem 401, and an eSBR processing stage 203 (which is the same as the stage 203 of FIG. 3), connected as shown in the figure. Typically, the decoder 400 also includes other processing elements (not shown).
[0096] In the operation of the decoder 400, a sequence of blocks of the encoded audio bitstream (MPEG-4 AAC bitstream) received by the decoder 400 is presented from the buffer 201 to the demultiplexer 215.
[0097] The demultiplexer 215 is combined and configured to demultiplex each block of the bitstream and then extract SBR metadata (including quantized envelope data), typically other metadata as well. The demultiplexer 215 is configured to present at least the SBR metadata to the eSBR processing stage 203. The demultiplexer 215 is also combined and configured to extract audio data from each block of the bitstream and present the extracted audio data to the decode subsystem (decode stage) 202.
[0098] The audio decoding subsystem 202 of the decoder 400 decodes the audio data extracted by the demultiplexer 215 (such decoding may be referred to as the "core" decoding operation), generates the decoded audio data, and is configured to present the decoded audio data to the eSBR processing stage 203. The decoding is performed in the frequency domain. Typically, the final stage of processing in the subsystem 202 applies an inverse Fourier transform to the decoded frequency domain audio data, so that the output of the subsystem is the decoded audio data in the time domain. The stage 203 is configured to apply the SBR (and eSBR) tools indicated by the SBR metadata (extracted by the demultiplexer 215) and the eSBR metadata generated in the subsystem 401 to the decoded audio data (i.e., perform SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata) to generate the fully decoded audio data output from the decoder 400. Typically, the decoder 400 includes a memory (accessible by the subsystem 202 and the stage 203) that stores the demultiplexed audio data and metadata output from the demultiplexer 215 (and optionally the subsystem 401), and the stage 203 is configured to access the audio data and metadata as needed during SBR and eSBR processing. The SBR processing in the stage 203 may be considered a post-processing of the output of the core decoding subsystem 202. Optionally, the decoder 400 also includes a final upmixing subsystem (which can apply parametric stereo ("PS") tools defined in the MPEG-4 AAC standard using the PS metadata extracted by the demultiplexer 215). The upmixing subsystem is coupled and configured to perform an upmixing on the output of the stage 203 to generate the fully decoded, upmixed audio output from the APU 210.
[0099] The control data generation subsystem 401 of FIG. 5 detects at least one attribute of an encoded audio bitstream to be decoded and is coupled and configured to generate eSBR control data (which, according to other embodiments of the present invention, may be any type of eSBR metadata of the type included in the encoded audio bitstream or may include it) in response to at least one result of the detection stage. The eSBR control data is presented at stage 203 to trigger and / or control the application of individual eSBR tools or combinations of eSBR tools when a particular attribute (or combination of attributes) of the bitstream is detected. For example, to control the execution of eSBR processing using harmonic conversion, some embodiments of the control data generation subsystem 401: a music detector (e.g., a simplified version of a normal music detector) for setting the sbrPatchingMode[ch] parameter in response to detecting whether the bitstream indicates music or not (and presenting the set parameter at stage 203); a transient detector for setting the sbrOversamplingFlag[ch] parameter in response to detecting the presence or absence of transient components in the audio content indicated by the bitstream (and presenting the set parameter at stage 203); and / or a pitch detector for setting the sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch] parameters in response to detecting the pitch of the audio content indicated by the bitstream (and presenting the set parameters at stage 203). Another aspect of the present invention is an audio bitstream decoding method performed by any embodiment of the decoder of the present invention described in this paragraph and the previous paragraph.
[0100] Aspects of the invention include an encoding or decoding method of a type configured (e.g., programmed) to be executed by an embodiment of an APU, system, or device of the invention. Other aspects of the invention include a system or device configured (e.g., programmed) to execute an embodiment of a method of the invention and a computer-readable medium (e.g., a disk) storing code (e.g., in a non-transitory manner) for implementing an embodiment or stage of a method of the invention. For example, a system of the invention can be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed (and / or otherwise configured) with software or firmware to perform any of a variety of operations including embodiments or stages of a method of the invention on data. Such a general-purpose processor can be or include a computer system including an input device, memory, and processing circuitry programmed (and / or otherwise configured) to perform an embodiment (or stage) of a method of the invention in response to data presented thereto, or alternatively can include it.
[0101] Embodiments of the invention may be implemented in hardware, firmware, software, or a combination of both (e.g., as a programmable logic array). Unless otherwise specified, algorithms or processes included as part of the present invention are not inherently related to any particular computer or other device. In particular, various general-purpose machines may be used in conjunction with programs written in accordance with the teachings herein, or it may be convenient to construct more specialized devices (e.g., integrated circuits) to perform the required method steps. Thus, the present invention may be implemented in one or more computer programs executed on one or more programmable computer systems (e.g., any implementation of the elements of FIG. 1 or the encoder 100 (or certain elements thereof) of FIG. 2 or the decoder 200 (or certain elements thereof) of FIG. 3 or the decoder 210 (or certain elements thereof) of FIG. 4 or the decoder 400 (or certain elements thereof) of FIG. 5). Each computer system has at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known manner.
[0102] Each such program may be implemented in any desired computer language (including machine language, assembly, or high-level procedural, logical, or object-oriented programming languages) to interface with the computer system. In any case, the language may be a compiled or interpreted language.
[0103] For example, when implemented by a computer software instruction sequence, the various functions and steps of the embodiments of the present invention may be implemented by a multi-threaded software instruction sequence running on suitable digital signal processing hardware. In that case, the various devices, steps and functions of the embodiments may correspond to portions of the software instructions.
[0104] Each such computer system is preferably stored in or downloaded from a storage medium or device readable by a general-purpose or special-purpose programmable computer (such as a semiconductor memory or media or magnetic or optical media). This is for configuring and operating the computer to execute the procedures described herein when the storage medium or device is read by the computer system. The system of the present invention may be implemented as a computer-readable storage medium configured with a computer program (i.e., storing a computer program). Here, the storage medium configured in this way causes the computer system to operate in a specific predefined manner so as to execute the functions described herein.
[0105] Some embodiments of the present invention have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the present invention. Numerous modifications and variations of the present invention are possible in light of the above teachings. It is understood that within the scope of the appended claims, the present invention may be practiced in ways other than specifically described herein. Even if there are reference signs included in the claims, they are for illustrative purposes only and should not be used to interpret or limit the claims in any way.
[0106] Some aspects will be described. 〔Aspect 1〕 A buffer configured to store at least one block of an encoded audio bitstream; A bitstream payload format descrambler coupled to the buffer and configured to descramble at least a portion of at least one block of the encoded audio bitstream; An audio processing unit having a decoding subsystem coupled to the bitstream payload format descrambler and configured to decode at least a portion of at least one block of the encoded audio bitstream, wherein the at least one block of the encoded audio bitstream: Includes padding elements, each padding element having an identifier indicating the start of the padding element and padding data following the identifier, and the padding data: Includes at least one flag for identifying whether an enhanced spectral band replication process should be performed on the audio content of the at least one block of the encoded audio bitstream, An audio processing unit. [Aspect 2] The audio processing unit according to Aspect 1, wherein the padding data further includes enhanced spectral band replication metadata. [Aspect 3] The audio processing unit according to Aspect 2, wherein the enhanced spectral band replication metadata does not include one or more parameters used for both spectral patching and harmonic conversion. [Aspect 4] The audio processing unit according to Aspect 2 or 3, wherein the enhanced spectral band replication metadata does not include a parameter for selecting between harmonic conversion and spectral patching. [Aspect 5] The audio processing unit according to any one of Aspects 2 to 4, wherein the enhanced spectral band replication metadata includes at least one of: i) a parameter indicating whether to perform pre-emphasis flattening; ii) a parameter indicating whether to perform inter-subband sample temporal envelope shaping; and iii) a parameter indicating whether to perform signal adaptive frequency domain oversampling. 〔Aspect 6〕 The up-spectral band replication metadata is metadata configured to enable at least one eSBR tool described or referred to in the MPEG USAC standard and not described or referred to in the MPEG-4 AAC standard, and is the audio processing unit according to any one of Aspects 2 to 5. 〔Aspect 7〕 The audio processing unit according to any one of Aspects 1 to 6, wherein at least one block of the encoded audio bitstream includes spectral band replication metadata. 〔Aspect 8〕 The audio processing unit according to Aspect 7 when citing Aspect 2, wherein the up-spectral band replication metadata does not include parameters equivalent to the parameters of the spectral band replication metadata. 〔Aspect 9〕 The audio processing unit according to Aspect 7 or 8, wherein the spectral band replication metadata is metadata configured to enable at least one SBR tool described or referred to in the MPEG-4 AAC standard. 〔Aspect 10〕 The audio processing unit according to any one of Aspects 7 to 9, wherein the spectral band replication metadata includes one or more parameters used for both spectral patching and harmonic conversion. 〔Aspect 11〕 The audio processing unit according to any one of Aspects 1 to 10, wherein the up-spectral band replication process includes harmonic conversion but does not include spectral patching. 〔Aspect 12〕 The value with the at least one flag indicates that the enhanced spectral band replication process should be performed on the audio content of the at least one block of the encoded audio bitstream, and another value of the at least one flag indicates that the basic spectral band replication process should be performed on the audio content of the at least one block of the encoded audio bitstream. The audio processing unit according to any one of Aspects 1 to 11. [Aspect 13] The basic spectral band replication process includes spectral patching but does not include harmonic conversion. The audio processing unit according to Aspect 12. [Aspect 14] The basic spectral band replication process is a spectral band replication process using spectral patching described in the MPEG-4 AAC standard. The audio processing unit according to Aspect 12 or 13. [Aspect 15] The enhanced spectral band replication process is a spectral band replication process using at least one eSBR tool described or referred to in the MPEG USAC standard and not described or referred to in the MPEG-4 AAC standard. The audio processing unit according to any one of Aspects 1 to 14. [Aspect 16] The audio processing unit is an audio decoder, and the identifier is a three-bit unsigned integer with a value of 0x6 and the most significant bit is transmitted first. The audio processing unit according to any one of Aspects 1 to 15. [Aspect 17] The padding data includes an extended payload, the extended payload includes spectral band replication extension data, and the extended payload is identified using a four-bit unsigned integer with a value of "1101" or "1110" and the most significant bit is transmitted first. Optionally, The spectral band replication extension data is: An optional spectral band replication header, The spectral band replication data after the header and including a spectral band replication extension element after the spectral band replication data, the flag being included in the spectral band replication extension element, An audio processing unit according to any one of Aspects 1 to 16. [Aspect 18] The at least one block of the encoded audio bitstream includes a first padding element and a second padding element, the first padding element including spectral band replication data, the second padding element including the flag but not including spectral band replication data, an audio processing unit according to any one of Aspects 1 to 17. [Aspect 19] Further comprising an enhanced spectral band replication processing subsystem configured to perform enhanced spectral band replication processing using or in response to the at least one flag, an audio processing unit according to any one of Aspects 1 to 18. [Aspect 20] A method for decoding an encoded audio bitstream, comprising: Receiving at least one block of the encoded audio bitstream; Demultiplexing at least a part of the at least one block of the encoded audio bitstream; Decoding at least a part of the at least one block of the encoded audio bitstream, The at least one block of the encoded audio bitstream being: Including a padding element having an identifier indicating the start of the padding element and padding data after the identifier, the padding data being: Including at least one flag for identifying whether enhanced spectral band replication processing should be performed on the audio content of the at least one block of the encoded audio bitstream. Method [Aspect 21] The method according to aspect 20, wherein the identifier is a three-bit unsigned integer having a value of 0x6 and with the most significant bit transmitted first without a code. [Aspect 22] The filling data includes an extended payload, the extended payload includes spectrum band replication extension data, the extended payload is identified using a four-bit unsigned integer having a value of "1101" or "1110" and with the most significant bit transmitted first, and optionally, The spectrum band replication extension data is: An optional spectrum band replication header, Spectrum band replication data after the header and Spectrum band replication extension elements after the spectrum band replication data, and the flag is included in the spectrum band replication extension elements, The method according to aspect 20 or 21. [Aspect 23] The method according to any one of aspects 20 to 22, wherein the up-spectral band replication process is harmonic conversion, and a certain value of the at least one flag indicates that the up-spectral band replication process should be performed on the audio content of the at least one block of the encoded audio bitstream, and another value of the at least one flag indicates that spectral patching should be performed on the audio content of the at least one block of the encoded audio bitstream but the harmonic conversion should not be performed. [Aspect 24] The spectrum band replication extension element includes up-spectral band replication metadata other than the flag, and the up-spectral band replication metadata includes a parameter indicating whether to perform pre-flattening, or The spectrum band replication extension element includes improvement spectrum band replication metadata other than the flag, and the improvement spectrum band replication metadata includes a parameter indicating whether to perform sub-band sample inter-time envelope shaping. The method according to aspect 22 or 23. [Aspect 25] The method according to any one of aspects 20 to 24, further comprising the step of performing an improved spectrum band replication process using the at least one flag, wherein the improved spectrum band replication includes harmonic conversion. [Aspect 26] The method according to any one of aspects 20 to 25 or the audio processing unit according to any one of aspects 1 to 19, wherein the encoded audio bitstream is an MPEG-4 AAC bitstream.
Claims
1. A bitstream payload format parser configured to demultiplex blocks of an encoded audio bitstream; An audio processing apparatus having a decoding subsystem coupled to the bitstream payload format parser and configured to decode at least a portion of the blocks of the encoded audio bitstream, wherein the blocks of the encoded audio bitstream: Contain padding elements, the padding elements having an identifier indicating the start of the padding element and padding data following the identifier, the identifier being a 3-bit unsigned integer with the most significant bit transmitted first and having a value of 0x6, and the padding data: At least one flag for identifying whether an enhanced spectral band replication process should be performed on the audio content of the blocks of the encoded audio bitstream; Enhanced spectral band replication metadata that does not include one or more parameters used for both spectral patching and harmonic conversion, the enhanced spectral band replication metadata being metadata configured to enable at least one eSBR tool described or referred to in the MPEG USAC standard and not described or referred to in the MPEG-4 AAC standard, The enhanced spectral band replication metadata includes a parameter indicating whether to perform signal-adaptive frequency domain oversampling, and the decoding subsystem is configured to perform signal-adaptive frequency domain oversampling when the parameter indicates that signal-adaptive frequency domain oversampling should be performed. An audio processing apparatus.
2. The padding data includes an extended payload, the extended payload includes spectral band replication extension data, the extended payload is identified using a four-bit unsigned integer with the most significant bit transmitted first and having a value of "1101" or "1110", and the spectral band replication extension data: A spectral band replication header, Spectral band replication data following the header and including a spectral band replication extension element after the spectral band replication data, wherein the flag is included in the spectral band replication extension element, The audio processing apparatus according to claim 1. **Claim 3** A method of decoding an encoded audio bitstream, executed by an audio processing apparatus, the method comprising: multiplexing and separating blocks of the encoded audio bitstream; decoding at least a part of the blocks of the encoded audio bitstream by a decoding subsystem of the audio processing apparatus, wherein the blocks of the encoded audio bitstream: include padding elements, each padding element having an identifier indicating the start of the padding element and padding data after the identifier, the identifier being a 3-bit unsigned integer with the most significant bit transmitted first and having a value of 0x6, and the padding data: a flag for identifying whether enhanced spectral band replication processing should be performed on the audio content of the blocks of the encoded audio bitstream; enhanced spectral band replication metadata that does not include one or more parameters used for both spectral patching and harmonic conversion, the enhanced spectral band replication metadata being metadata configured to enable at least one eSBR tool described or referred to in the MPEG USAC standard and not described or referred to in the MPEG-4 AAC standard, the enhanced spectral band replication metadata includes a parameter indicating whether to perform signal-adaptive frequency domain oversampling, and the decoding subsystem is further configured to perform signal-adaptive frequency domain oversampling when the parameter indicates that signal-adaptive frequency domain oversampling should be performed. Method. **Claim 4** wherein the padding data includes an extended payload, the extended payload includes spectral band replication extension data, the extended payload is identified using a 4-bit unsigned integer with the most significant bit transmitted first and having a value of "1101" or "1110", and the spectral band replication extension data: a spectral band replication header, the spectral band replication data after the header and a spectral band replication extension element after the spectral band replication data, the flag being included in the spectral band replication extension element, The method according to claim 3.
5. The method according to claim 3, wherein the encoded audio bitstream is an MPEG-4 AAC bitstream.
6. A non-transitory computer-readable medium having instructions that, when executed by a processor, cause the processor to perform the method according to claim 3.
Citation Information
Patent Citations
IEC14496-3
Decoder, encoder, encoding decoding system, decoding method, encoding method, decoding program and encoding program
JP2013125187A
Frame element positioning in a bitstream frame representing audio content
JP2014512020A
Bandwidth expansion parameter-generator, encoder, decoder, bandwidth expansion parameter-generating method, encoding method, and decoding method
WO2014115225A1