Method for high-frequency reconstruction of audio signals and audio processing unit
Enhanced spectral band replication (eSBR) processing addresses the limitations of MPEG-4 AAC by regenerating high-frequency bands using harmonic transposition and QMF patching, improving audio quality for music content while maintaining low data rates and compatibility.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- DOLBY INTERNATIONAL AB
- Filing Date
- 2024-11-13
- Publication Date
- 2026-04-27
AI Technical Summary
Existing audio encoding techniques, such as MPEG-4 AAC, struggle with spectral band reproduction for certain audio types, particularly music content, due to limitations in high-frequency reconstruction methods like spectral patching.
Implementing enhanced spectral band replication (eSBR) processing, which includes harmonic transposition and QMF patching, to regenerate high-frequency bands based on metadata within the audio bitstream, ensuring improved spectral characteristics and stability.
Enhances audio quality by accurately reconstructing high-frequency components, particularly for music content, while maintaining low data rates and being backward-compatible with existing standards.
Smart Images

Figure 0007852013000017 
Figure 0007852013000018 
Figure 0007852013000019
Abstract
Description
[Background technology]
[0001] Cross-reference of related applications This application claims priority based on the following application, which is incorporated herein: U.S. Provisional Application No. 62 / 622,205, filed on 26 January 2018.
[0002] Technical field The embodiments relate to audio signal processing, and more specifically to encoding, decoding, or transcoding of an audio bitstream with control data specifying that either a basic form of high-frequency reconstruction (HFR) or an enhanced form of HFR should be performed with respect to the audio data.
[0003] Background of the Invention A typical audio bitstream includes both audio data (e.g., encoded audio data) that represents one or more channels of audio content, and metadata that represents at least one feature of the audio data or audio content. One well-known format for generating encoded audio bitstreams is the MPEG-4 Advanced Audio Coding (AAC) format, which is described in the MPEG standard ISO / IEC 14496-3:2009. In the MPEG4 standard, AAC stands for "Advanced Audio Coding," and HE-AAC stands for "High Efficiency Advanced Audio Coding."
[0004] The MPEG-4 AAC standard defines several audio profiles that determine whether objects and encoding tools are present in the compliant encoder or decoder. Three of these audio profiles are (1) the AAC profile, (2) the HE-AAC profile, and (3) the HE-AAC v2 profile. The AAC profile includes the AAC Low Complexity (or "AAC-LC") object type. The AAC-LC object corresponds to the MPEG-2 AAC Low Complexity profile with some modifications and does not include the Spectral Band Reproduction ("SBR") object type or the Parametric Stereo ("PS") object type. The HE-AAC profile is a superset of the AAC profile and additionally includes the SBR object type. The HE-AACv2 profile is a superset of the HE-AAC profile and additionally includes the PS object type.
[0005] The SBR object type includes a spectral band replication tool, which is a crucial high-frequency reconstruction ("HFR") coding tool that significantly improves the compression efficiency of perceptual speech codecs. SBR reconstructs the high-frequency components of the audio signal at the receiver side (e.g., in the decoder). Therefore, the encoder only needs to encode and transmit the low-frequency components, enabling very high audio quality at low data rates. SBR is based on replicating a sequence of harmonics pre-transcribed to reduce the data rate from control data and available bandwidth-limiting signals obtained from the encoder. The ratio between the tonal and noise components is maintained by adaptive inverse filtering, in addition to the selective addition of noise and sine waves. In the MPEG-4 AAC standard, the SBR tool performs spectral patching (also called linear transformation or spectral transformation), in which a number of consecutive quadrature mirror filter (QMF) subbands are copied (or "patched") from the transmitted low-band portion of the audio signal to the high-band portion of the audio signal, generated by the decoder.
[0006] Spectral patching or linear transformation may not be ideal for certain audio types, such as music content, which involves relatively low crossover frequencies. Therefore, techniques are needed to improve spectral band reproduction. [Overview of the Initiative]
[0007] A method for decoding an encoded audio bitstream is disclosed relating to a first-class embodiment. The method includes receiving an encoded audio bitstream and decoding the audio data to generate a decoded lowband audio signal. The method further includes extracting high-frequency reconstruction metadata and filtering the decoded lowband audio signal through an analysis filter bank to generate a filtered lowband audio signal. The method further includes extracting flags indicating whether spectral transformation or harmonic transposition should be performed on the audio data and regenerating the highband portion of the audio signal using the high-frequency reconstruction metadata and the filtered lowband audio signal according to the flags. Finally, the method includes combining the filtered lowband audio signal and the regenerated highband portion to form a broadband audio signal.
[0008] A second-class embodiment relates to an audio decoder for decoding an encoded audio bitstream. The decoder includes an input interface for receiving an encoded audio bitstream (the encoded audio bitstream contains audio data representing the low-band portion of an audio signal) and a core decoder for decoding the audio data to produce a decoded low-band audio signal. The decoder also includes a demultiplexer for extracting high-frequency reconstruction metadata from the encoded audio bitstream (the high-frequency reconstruction metadata contains operating parameters for a high-frequency reconstruction process that linearly transforms a number of consecutive subbands from the low-band portion of the audio signal to the high-band portion of the audio signal) and an analysis filter bank for filtering the decoded low-band audio signal to produce a filtered low-band audio signal. The decoder further includes a demultiplexer for extracting flags from the encoded audio bitstream indicating whether a linear transformation or harmonic transposition should be performed on the audio data, and a high-frequency regenerator for regenerating the high-band portion of the audio signal using the high-frequency reconstruction metadata and the filtered low-band audio signal according to the flags. Finally, the decoder includes a synthesis filter bank for combining the filtered low-band audio signal with the regenerated high-band portion to form a wideband audio signal.
[0009] Other embodiments of the class relate to encoding and transcoding an audio bitstream that includes metadata identifying whether enhanced spectral band replication (eSBR) processing should be performed. [Brief explanation of the drawing]
[0010] [Figure 1] This is a block diagram of an embodiment of a system that may be configured to carry out embodiments of the method of the present invention. [Figure 2] A block diagram of an encoder, which is an embodiment of the audio processing unit of the present invention. [Figure 3] A block diagram of a system including a decoder, which is an embodiment of the audio processing unit of the present invention, and a post-processor optionally coupled thereto. [Figure 4] A block diagram of a decoder, which is an embodiment of the audio processing unit of the present invention. [Figure 5] A block diagram of a decoder, which is another embodiment of the audio processing unit of the present invention. [Figure 6] A block diagram of another embodiment of the audio processing unit of the present invention. [Figure 7] A diagram showing a block of an MPEG-4 AAC bitstream including segmented segments.
Embodiments for Carrying Out the Invention
[0011] Notation and Terminology Throughout this disclosure, including the claims, the expression "performing processing on" a signal or data (e.g., filtering, scaling, conversion, or application of gain to the signal or data) is used in a broad sense to indicate performing direct processing on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has been pre-filtered or pre-processed prior to the execution of that processing).
[0012] Throughout this disclosure, including the claims, the terms "audio processing unit" or "audio processor" are used broadly to refer to a system, device, or apparatus configured to process audio data. Examples of audio processing units include, but are not limited to, encoders, transcoders, decoders, codec, pre-processing systems, post-processing systems, and bitstream processing systems (often referred to as bitstream processing tools). Virtually all consumer electronics products, such as mobile phones, televisions, laptops, and tablet computers, incorporate an audio processing unit or audio processor.
[0013] Throughout this disclosure, including the claims, the terms "coupled" or "coupling" are used broadly to mean a direct or indirect connection. Thus, when a first device is coupled to a second device, that connection may be through a direct connection or through an indirect connection via other devices and connections. Further, components integrated within or with other components are also coupled to each other.
[0014] Detailed description of embodiments of the invention The MPEG-4 AAC standard is assumed to include metadata indicating each type of high-frequency reconstruction (HFR) processing to be applied (if to be applied) by a decoder to decode the audio content of a bitstream of an encoded MPEG-4 AAC bitstream, and / or controlling such HFR processing, and / or indicating at least one characteristic or parameter of at least one HFR tool to be used to decode the audio content of the bitstream. Here, we use the expression "SBR metadata" to denote this type of metadata described or referred to in the MPEG-4 AAC standard for use in spectral band replication ("SBR"). As will be understood by those skilled in the art, SBR is a form of HFR.
[0015] SBR is preferably used as a dual-rate system, where the underlying codec operates at half the original sampling rate, while the SBR operates at the original sampling rate. The SBR encoder operates in parallel with the underlying core codec, albeit at a higher sampling rate. While SBR is primarily post-processing in the decoder, key parameters are extracted by the encoder to compensate for the highest accuracy high-frequency reconstruction in the decoder. The encoder estimates the spectral envelope of the SBR range for a time and frequency range / resolution appropriate to the current input signal segment characteristics. The spectral envelope is estimated by complex QMF analysis and subsequent energy calculations. The time and frequency resolution of the spectral envelope can be selected with a high degree of freedom to ensure optimal time-frequency resolution for a given input segment. Envelope estimation must consider that transients located in the original region, primarily in the high-frequency region (e.g., a high-hat), are present in small quantities in the SBR high band generated before envelope adjustment, because the high band in the decoder is based on the low band, and its transients are judged to be much smaller compared to the high band. This aspect imposes different conditions on the time-frequency resolution of the spectral envelope data compared to normal spectral envelope estimation as used in other audio coding algorithms.
[0016] Apart from the spectral envelope, several additional parameters are extracted that represent the spectral characteristics of the input signal for different time and frequency domains. The encoder, naturally, has access not only to the original signal but also to information about how the SBR unit within the decoder generates the high bands. Therefore, the system can handle situations where the low bands constitute a strong harmonic sequence and the regenerated high bands mainly consist of random signal components, as well as situations where the high band region has no corresponding low band and the original high bands have strong tonal components. Furthermore, the SBR encoder operates in close relation to the underlying core codec to evaluate which frequency ranges should be covered by the SBR at a given time. In the case of stereo signals, the SBR data is efficiently encoded before transmission by utilizing not only the channel dependency of the control data but also entropy coding.
[0017] Control parameter extraction algorithms typically require careful tuning to the underlying codec, given a given bitrate and sampling rate. This is due to the fact that lower bitrates usually exhibit a larger SBR range compared to higher bitrates, and different sampling rates correspond to different temporal resolutions of SBR frames.
[0018] An SBR decoder typically comprises several different parts. These include a bitstream decoding module, a high-frequency reconstruction (HFR) module, an additional high-frequency component module, and an envelope tuning module. The system is based on a complex-valued QMF filter bank (for high-quality SBRs) or a real-valued QMF filter bank (for low-power SBRs). Embodiments of the present invention are applicable to both high-quality and low-power SBRs. In the bitstream extraction module, control data is read from the bitstream and decoded. A time-frequency grid is obtained for the current frame before reading envelope data from the bitstream. The underlying core decoder decodes the audio signal of the current frame (albeit at a lower sampling rate) to generate time-domain audio samples. The resulting audio data frame is used for high-frequency reconstruction by the HFR module. The decoded low-band signal is then analyzed using the QMF filter bank. Subsequently, high-frequency reconstruction and envelope tuning are performed on subband samples of the QMF filter bank. The high frequencies are reconstructed from the low bands in a flexible manner based on given control parameters. Furthermore, the reconstructed high bands are adaptively filtered on a subband channel basis according to control data to ensure appropriate spectral characteristics in a given time / frequency domain.
[0019] The top level of an MPEG-4 AAC bitstream is a sequence of data blocks ("raw_data_block" elements), each of which is a segment of data (hereinafter referred to as a "block") containing audio data (typically spanning a period of 1024 or 960 samples) and associated information and / or other data. Here, we use the term "block" to refer to a segment of an MPEG-4 AAC bitstream containing audio data (and corresponding metadata and optionally other associated data) that determines or indicates one (not more than one) "raw_data_block" element.
[0020] Each block of an MPEG-4 AAC bitstream can contain numerous syntactic elements (each of which also appears in the bitstream as a data segment). Seven types of such syntactic elements are defined in the MPEG-4 AAC standard. Each syntactic element is identified by a different value of the data element "id_syn_ele". Specific examples of syntactic elements include "single_channel_element()", "channel_pair_element()", and "fill_element()". A single-channel element is a container (mono audio signal) containing audio data for a single audio channel. A channel-pair element contains audio data for two audio channels (stereo audio signals).
[0021] A fill element is a container of information containing an identifier (for example, the value of the id_syn_ele element above) followed by data, which is referred to as "fill data". Historically, fill elements have been used to adjust the instantaneous bitrate of a bitstream transmitted over a channel of a constant rate. By adding an appropriate amount of fill data to each block, it is possible to achieve a constant data rate.
[0022] According to embodiments of the present invention, fill data may include one or more augmentation payloads that extend the types of data (e.g., metadata) that can be transmitted in the bitstream. A decoder that receives the bitstream with fill data containing new types of data may be optionally used by a device receiving the bitstream (e.g., a decoder) to extend the functionality of the device. Thus, as will be understood by those skilled in the art, fill elements are a special type of data structure, distinct from data structures typically used to transmit audio data (e.g., audio payloads containing channel data).
[0023] In some embodiments of the present invention, the identifier used to identify a fill element may consist of a three-bit unsigned integer transmitted most significant bit first (uimsbf) having the value 0x6. A single block may contain several instances of syntactic elements of the same type (e.g., multiple fill elements).
[0024] Another standard for encoding audio bitstreams is the MPEG (Unified Speech and Audio Coding: USAC) standard (ISO / IEC 23003-3:2012). The MPEG USAC standard describes the encoding and decoding of audio content using spectral band duplication processing (including SBR processing as described in the MPEG-4 AAC standard, and other enhanced forms of spectral band duplication processing). This processing applies extended and enhanced versions of spectral band duplication tools (often referred to here as “extended SBR tools” or “eSBR tools”) of the set of SBR tools described in the MPEG-4 AAC standard. Thus, eSBR (as defined in the USAC standard) is an improvement over SBR (as defined in the MPEG-4 AAC standard).
[0025] Here, the term "enhanced SBR processing" (or "eSBR processing") is used to describe spectral band replication processing using at least one eSBR tool not described or mentioned in the MPEG-4 AAC standard (for example, at least one eSBR tool described or mentioned in the MPEG USAC standard). Examples of such eSBR tools include harmonic transposition and QMF patch processing, or "pre-flattening".
[0026] An integer-order T harmonic transposer maps a sine wave of frequency ω to a sine wave of frequency Tω while preserving the signal duration. Typically, three orders T=2, 3, and 4 are used in sequence to generate each part of the desired output frequency range using the smallest possible transposition order. If an output above the 4th-order transposition range is required, it can be generated by frequency shifting. Where possible, a baseband-time domain sampled near critically is created for processing to minimize computational complexity.
[0027] The harmonic transposer may be either QMF or DFT based. When using a QMF-based harmonic transposer, bandwidth expansion of the core encoder time-domain signal is performed entirely within the QMF domain using a modified phase vocoder structure, followed by time extension after decimation for all QMF subbands. Transposition using several transposition factors (e.g., T=2, 3, 4) is performed in a common QMF analysis / synthesis transformation stage. Since the QMF-based harmonic transposer does not feature signal-adaptive frequency-domain oversampling, the corresponding flag in the bitstream (sbrOversamplingFlag[ch]) may be ignored.
[0028] When using a DFT-based harmonic transposer, the factor 3 and 4 transposers (third- and fourth-order transposers) are preferably incorporated into the factor 2 transposer (second-order transposer) by interpolation to reduce complexity. For each frame (corresponding to corecoder samples, coreCoderFrameLength), the nominal "full-size" transposition size of the transposer is first determined by the signal-adaptive frequency-domain oversampling flag (sbrOverSamplingFlag[ch]) in the bitstream.
[0029] When sbrPatchingMode==1, it indicates that linear transposition should be used to generate the high band, and an additional step may be introduced to avoid discontinuities in the shape of the spectral envelope of the high-frequency signal input to the subsequent envelope tuner. This improves the operation of the next envelope tuner stage and, as a result, produces a high-band signal that is perceived as more stable. The operation of the additional preprocessing is beneficial for signal types in which the coarse spectral envelope of the low-band signal used for high-frequency reconstruction exhibits large fluctuations in level. However, the value of the bitstream element may be determined within the encoder by applying any kind of signal-dependent classification. The additional preprocessing is preferably activated by a 1-bit bitstream element bs_sbr_preprocessing. When bs_sbr_preprocessing is set to 1, the additional processing is enabled. When bs_sbr_preprocessing is set to zero, the additional processing is disabled. The additional processing is applied to the low-band X for each patch. Low It is preferable to utilize the pre-gain curve used by the high-frequency generator to scale the signal. For example, the pre-gain curve may be calculated as follows:
number
number
number
number
[0030] A bitstream generated in accordance with the MPEG USAC standard (often referred to hereby as a “USAC bitstream”) includes encoded audio content and typically includes metadata indicating each type of spectral band replication process applied by a decoder to decode the audio content of the USAC bitstream, and / or metadata indicating at least one characteristic or parameter of at least one SBR tool and / or eSBR tool used to control such spectral band replication processes and / or decode the audio content of the USAC bitstream.
[0031] In this application, the term “enhanced SBR metadata” (or “eSBR metadata”) is used to indicate metadata that describes each type of spectral band duplication process applied by a decoder to decode the audio content of an encoded audio bitstream (e.g., a USAC bitstream), and / or controls such spectral band duplication processes, and / or controls such audio content, and / or controls such audio content, but is not described or mentioned in the MPEG4AAC standard. Specific examples of eSBR metadata include metadata described or mentioned in the MPEG USAC standard but not described or mentioned in the MPEG-4AAC standard (for describing or controlling spectral band duplication processes). Therefore, in this application, eSBR metadata refers to metadata that is not SBR metadata, and SBR metadata refers to metadata that is not eSBR metadata.
[0032] A USAC bitstream may contain both SBR metadata and eSBR metadata. More specifically, a USAC bitstream may contain eSBR metadata that controls the performance of eSBR processing by the decoder, and SBR metadata that controls the performance of SBR processing by the decoder. According to a typical embodiment of the present invention, the eSBR metadata (e.g., eSBR-specific configuration data) is contained in the MPEG-4 AAC bitstream (e.g., the sbr_extension() container at the end of the SBR payload) (according to the present invention).
[0033] During the decoding of an encoded bitstream using an eSBR toolset (including at least one eSBR tool), the decoder's execution of eSBR processing regenerates the high-frequency bands of the audio signal based on a replica of the harmonic sequence truncated during encoding. Such eSBR processing typically adjusts the spectral envelope of the generated high-frequency bands, applies inverse filtering, and adds noise and sinusoidal components to recreate the spectral characteristics of the original audio signal.
[0034] According to a typical embodiment of the present invention, the eSBR metadata is contained in one or more metadata segments of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) containing audio data encoded in other segments (audio data segments) (e.g., containing a small number of control bits which are eSBR metadata). Typically, at least one such metadata segment in each block of the bitstream is (or contains) a fill element (containing an identifier indicating the start of a fill element), and the eSBR metadata is contained in the fill element after the identifier.
[0035] Figure 1 is a block diagram of an exemplary audio processing chain (audio data processing system), in which one or more elements of the system can be configured according to embodiments of the present invention. The system includes elements that are coupled together, as shown as encoder 1, distribution subsystem 2, decoder 3, and post-processing unit 4. Variations of the illustrated system may omit one or more elements or include additional audio data processing units.
[0036] In some implementations, encoder 1 (with an optional preprocessing unit) is configured to accept a PCM (time-domain) sample containing audio content as input and to output an encoded audio bitstream representing the audio content (having a format compliant with the MPEG-4 AAC standard). The data of the bitstream representing the audio content is often referred to herein as “audio data” or “encoded audio data.” When the encoder is configured according to a typical embodiment of the present invention, the audio bitstream output from the encoder includes eSBR metadata (and typically other metadata) as well as audio data.
[0037] One or more encoded audio bitstreams output from encoder 1 may be asserted to encoded audio distribution subsystem 2. Subsystem 2 is configured to store and / or distribute each encoded bitstream output from encoder 1. The encoded audio bitstreams output from encoder 1 may be stored by subsystem 2 (for example, in the form of a DVD or Blu-ray disc), transmitted by subsystem 2 (capable of realizing a transmission link or network), or both stored and transmitted by subsystem 2.
[0038] Decoder 3 is configured to decode an encoded MPEG-4 AAC audio bitstream (generated by encoder 1) received via subsystem 2. In some embodiments, decoder 3 is configured to extract eSBR metadata from each block of the bitstream, decode the bitstream (including performing eSBR processing using the extracted eSBR metadata), and generate decoded audio data (e.g., a stream of decoded PCM audio samples). In some embodiments, decoder 3 is configured to extract SBR metadata from the bitstream (but ignoring the eSBR metadata contained in the bitstream), decode the bitstream (including performing SBR processing using the extracted SBR metadata), and generate decoded audio data (e.g., a stream of decoded PCM audio samples). Typically, decoder 3 includes a buffer that stores (e.g., non-temporarily) segments of the encoded audio bitstream received from subsystem 2.
[0039] The post-processing unit 4 in Figure 1 is configured to receive a stream of decoded audio data from the decoder 3 (e.g., decoded PCM audio samples) and perform post-processing on it. The post-processing unit may also be configured to render the post-processed audio content (or the decoded audio received from the decoder 3) for playback by one or more speakers.
[0040] Figure 2 is a block diagram of an encoder (100) which is an embodiment of the audio processing unit of the present invention. Any of the components or elements of the encoder 100 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware and software. The encoder 100 includes an encoder 105, a stuffer / formatter stage 107, a metadata generation stage 106, and a buffer memory 109 connected as shown. Typically, the encoder 100 also includes other processing elements (not shown). The encoder 100 is configured to convert an input audio bitstream into an encoded output MPEG-4 AAC bitstream.
[0041] The metadata generation unit 106 is configured and coupled to generate (and / or pass through) metadata (including eSBR metadata and SBR metadata) that should be included by the stage 107 in the encoded bitstream to be output from the encoder 100.
[0042] Encoder 105 is configured to encode the input audio data (for example, by performing compression therein) and assert and combine the resulting encoded audio with stage 107 so that it can be included in the encoded bitstream to be output from stage 107.
[0043] Stage 107 is configured to multiplex the encoded audio from encoder 105 and metadata (including eSBR metadata and SBR metadata) from generation unit 106 to generate an encoded bitstream output from stage 107, preferably having a format specified by any embodiment of the present invention.
[0044] The buffer memory 109 is configured to store (for example, in a non-transient manner) at least one block of the encoded audio bitstream output from the stage 107, and then a sequence of blocks of the encoded audio bitstream is asserted from the buffer memory 109 as output from the encoder 100 to the transmission system.
[0045] Figure 3 is a block diagram of a system including a decoder (200) and an optional post-processor (300) coupled thereto, which is an embodiment of the audio processing unit of the present invention. Any component or element of the decoder 200 and the post-processor 300 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware and software. The decoder 200 includes a buffer memory 201, a bitstream payload deformator 205, an audio decoding subsystem 202 (sometimes referred to as the “core” decoding stage or “core” decoding subsystem), an eSBR processing stage 203, and a control bit generation stage 204, connected as shown. Typically, the decoder 200 also includes other processing elements (not shown).
[0046] Buffer memory (buffer) 201 stores at least one block (for example, in a non-temporary format) of the encoded MPEG-4 AAC audio bitstream received by decoder 200. In the operation of decoder 200, a sequence of blocks of the bitstream is asserted from buffer 201 to deformatter 205.
[0047] In a modified embodiment of Figure 3 (or the embodiment of Figure 4 described later), an APU that is not a decoder (e.g., APU 500 in Figure 6) includes a buffer memory (e.g., the same buffer memory as buffer 201), which stores (e.g., in a non-temporary manner) at least one block of the same type of encoded audio bitstream (e.g., an MPEG-4 AAC audio bitstream) (i.e., an encoded audio bitstream including eSBR metadata) received by buffer 201 in Figure 3 or Figure 4.
[0048] Referring again to Figure 3, the deformator 205 is configured and coupled to demultiplex (or separate) each block of the bitstream, extract therefrom the SBR metadata and eSBR metadata (and typically other metadata) (including the quantized envelope data), assert at least the eSBR metadata and SBR metadata to the eSBR processing stage 203, and typically assert other extracted metadata to the decoding subsystem 202 (and optionally also to control the bit generation unit 204). The deformator 205 is also configured and coupled to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.
[0049] The system in Figure 3 also optionally includes a post-processor 300. The post-processor 300 includes a buffer memory (buffer) 301 and other processing elements (not shown) including at least one processing element coupled to buffer 301. Buffer 301 stores (in a non-temporary manner) at least one block (or frame) of decoded audio data received by the post-processor 300 from decoder 200. The processing elements of the post-processor 300 are configured and coupled to receive and adaptively process a sequence of blocks of the decoded audio output from buffer 301 using metadata output from the decoding subsystem 202 (and / or deformatter 205) and / or control bits output from stage 204 of decoder 200.
[0050] The audio decoding subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the parser 205 (such decoding may be referred to as the “core” decoding process) to produce decoded audio data and to assert the decoded audio data to the eSBR processing stage 203. Decoding is performed in the frequency domain and typically involves inverse quantization and subsequent spectral processing. Typically, the final stage of processing in subsystem 202 is to apply a frequency-domain to time-domain transformation to the decoded frequency-domain audio data, so that the subsystem output is time-domain decoded audio data. Stage 203 is configured to apply SBR and eSBR tools, indicated by eSBR metadata and the eSBR (extracted by the parser 205), to the decoded audio data (i.e., perform SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata) to produce fully decoded audio data output from the decoder 200 (e.g., to the post-processor 300). Typically, the decoder 200 includes memory (accessible by subsystem 202 and stage 203) for storing metadata and deformatted audio data from the deformatter 205, and stage 203 is configured to access the audio data and metadata (including SBR metadata and eSBR metadata) as needed during SBR and eSBR processing. The SBR and eSBR processing in stage 203 may be considered post-processing for the output of the core decoding subsystem 202.Optionally, the decoder 200 also includes a final upmixing subsystem (which can apply the parametric stereo (PS) tools defined in the MPEG-4 AAC standard, using PS metadata extracted by the deformator 205 and / or control bits generated in subsystem 204) that is configured to perform upmixing on the output of the decoder 203 and to produce fully decoded, upmixed audio output from the decoder 200. Alternatively, the post-processor 300 is configured to perform upmixing on the output of the decoder 200 (for example, using PS metadata extracted by the deformator 205 and / or control bits generated in subsystem 204).
[0051] In response to metadata extracted by the deformator 205, the control bit generator 204 can generate control data, which can be used within the decoder 200 (e.g., in the final upmixing subsystem) and / or asserted as an output of the decoder 200 (e.g., to the post-processor 300 for use in post-processing). In response to metadata extracted from the input bitstream (and optionally in response to control data), the stage 204 can generate (and assert to the post-processor 300) control bits indicating that the decoded audio data output from the eSBR processing stage 203 should undergo a particular type of post-processing. In some implementations, the decoder 200 is configured to assert the metadata extracted by the deformator 205 from the input bitstream to the post-processor 300, and the post-processor 300 is configured to perform post-processing on the decoded audio data output from the decoder 200 using the metadata.
[0052] Figure 4 is a block diagram of an audio processing unit ("APU") (210), which is another embodiment of the audio processing unit of the present invention. The APU 210 is a legacy decoder not configured to perform eSBR processing. Any component or element of the APU 210 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware. The APU 210 includes a buffer memory 201, a bitstream payload deformatter (parser) 215, an audio decoding subsystem 202 (often referred to as the "core" decoding stage or "core" decoding subsystem), and an SBR processing stage 213 in the configuration shown. Typically, the APU 210 includes other processing elements (not shown). The APU 210 may represent, for example, an audio encoder, decoder, or transcoder.
[0053] Elements 201 and 202 of APU210 are identical to the numbered elements of decoder 200 (in Figure 3), and their above description will not be repeated. In the operation of APU210, a sequence of blocks of the encoded audio bitstream (MPEG-4 AAC bitstream) received by APU210 is asserted from buffer 201 to deformatter 215.
[0054] The deformator 215 is configured to demultiplex each block of the bitstream and extract therefrom SBR metadata (including quantization envelope data) and typically other metadata, but to ignore and combine eSBR metadata that may be included in the bitstream according to some embodiment of the present invention. The deformator 215 is configured to assert at least the SBR metadata to the SBR processing stage 213. The deformator 215 is also configured to extract audio data from each block of the bitstream and to assert the extracted audio data to the decoding subsystem (decoding stage) 202.
[0055] The audio decoding subsystem 202 of the decoder 200 is configured to decode the audio data extracted by the deformatter 215 to produce decoded audio data (such decoding may be referred to as the “core” decoding process) and to assert the decoded audio data to the SBR processing stage 213. Decoding is performed in the frequency domain. Typically, the final stage of processing in subsystem 202 is to apply a frequency-domain to time-domain conversion to the decoded frequency-domain audio data, so that the output of the subsystem is the time-domain decoded audio data. Stage 213 is configured to apply an SBR tool (but not an eSBR tool) indicated by the SBR metadata (extracted by the deformatter 215) to the decoded audio data (i.e., to perform SBR processing on the output of the decoding subsystem 202 using the SBR metadata) to produce fully decoded audio data to be output from the APU 210 (e.g., to the post-processor 300). Typically, the APU210 includes memory (accessible by subsystems 202 and stage 213) for storing deformatted audio data and metadata output from the deformator 215, and stage 213 is configured to access the audio data and metadata (including SBR metadata) as needed during SBR processing. The SBR processing in stage 213 may be considered post-processing on the output of the core decoding subsystem 202. Optionally, the APU210 also includes a final upmixing subsystem configured and combined to perform upmixing on the output of stage 213 to produce fully decoded, upmixed audio output from the APU210 (this subsystem can apply the parametric stereo (PS) tool as defined in the MPEG-4 AAC standard, using the PS metadata extracted by the deformator 215).Alternatively, the post-processor is configured to perform upmixing on the output of the APU 210 (for example, using PS metadata extracted by the deformator 215 and / or control bits generated in the APU 210).
[0056] Various implementations of the encoder 100, decoder 200, and APU 210 are configured to perform various embodiments of the method of the present invention.
[0057] According to several embodiments, eSBR metadata is included in an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) (e.g., including a small number of control bits that are eSBR metadata), and as a result, legacy decoders (not configured to parse eSBR metadata or to use any eSBR tools that involve eSBR metadata) can ignore the eSBR metadata, but nevertheless, it is possible to decode the bitstream to the extent possible without using any eSBR tools or eSBR metadata that involve eSBR metadata, typically without any significant penalty to the decoded audio quality. However, an eSBR decoder configured to parse the bitstream to identify eSBR metadata and to use at least one eSBR tool depending on the eSBR metadata will enjoy the tonal patterns that utilize at least one such eSBR tool. Thus, embodiments of the present invention provide means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner.
[0058] Typically, eSBR metadata in a bitstream indicates one or more of the following eSBR tools (which are described in the MPEG USAC standard and may or may not be applied by the encoder during bitstream generation) (for example, indicating at least one characteristic or parameter of it): • Harmonic transposition; and QMF patching additional pre-processing (pre-flattening).
[0059] For example, eSBR metadata contained in a bitstream may include parameter values such as sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch], and bs_sbr_preprocessing (these are described in the MPEG USAC standard and this disclosure).
[0060] Here, if X is some parameter, the notation X[ch] indicates that the parameter relates to the channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, we often omit the notation [ch] and assume that the relevant parameter relates to the channel of the audio content.
[0061] Here, if X is some parameter, the notation X[ch][env] indicates that the parameter relates to the SBR envelope ("env") of the channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, we often omit the notations [env] and [ch] and assume that the relevant parameter relates to the SBR envelope of the channel of the audio content.
[0062] During the decoding of the encoded bitstream, the performance of harmonic transposition between the eSBR processing stages of decoding (each channel "ch" of the audio content represented by the bitstream) is controlled by the following eSBR metadata parameters: sbrPatchingMode[ch]; sbrOversamplingFlag[ch]; sbrPitchInBinsFlag[ch]; and sbrPitchInBins[ch].
[0063] The value "sbrPatchingMode[ch]" indicates the type of transposer used in the eSBR: sbrPatchingMode[ch]=1 indicates linear transposition patching as described in section 4.6.18 of the MPEG-4 AAC standard (such as that used in high-quality SBRs or low-power SBRs); sbrPatchingMode[ch]=0 indicates harmonic SBR patching as described in section 7.5.3 or 7.5.4 of the MPEG USAC standard.
[0064] The value "sbrOversamplingFlag[ch]" indicates the use of signal-adaptive frequency-domain oversampling in eSBR combined with DFT-based harmonic SBR patching, as described in Section 7.5.3 of the MPEG USAC standard. This flag controls the size of the DFT used in the transposer: 1 indicates that signal-adaptive frequency-domain oversampling is enabled, as described in Section 7.5.3.1 of the MPEG USAC standard; 0 indicates that signal-adaptive frequency-domain oversampling is disabled, as described in Section 7.5.3.1 of the MPEG USAC standard.
[0065] The value "sbrPitchInBinsFlag[ch]" controls the interpretation of the sbrPitchInBins[ch] parameter: 1 indicates that the value of sbrPitchInBins[ch] is valid and greater than zero; 0 indicates that the value of sbrPitchInBins[ch] is set to zero.
[0066] The value "sbrPitchInBins[ch]" controls the addition of cross-product terms in the SBR harmonic transposer. The value sbrPitchinBins[ch] is an integer within the range [0, 127] and represents the distance measured at the frequency bin of the 1536-line DFT that acts on the sampling frequency of the core coder.
[0067] If an MPEG-4 AAC bitstream represents an SBR channel pair and those channels are not coupled (i.e., not a single SBR channel), the bitstream will show two instances of the above syntax (relating to harmonic or non-harmonic transpositions), one for each channel of sbr_channel_pair_element().
[0068] Harmonic transposition in eSBR tools typically improves the quality of decoded music signals at relatively low crossover frequencies. Non-harmonic transposition (i.e., legacy spectral patching) typically improves speech signals. Therefore, the starting point in determining which type of transposition is preferable for encoding particular audio content is to choose the transposition method based on speech / music detection, with harmonic transposition used for music content and spectral patching used for speech content.
[0069] The performance of pre-flattening during eSBR processing is controlled by the value of a single bit eSBR metadata parameter known as "bs_sbr_preprocessing," which determines whether or not pre-flattening is performed. When an SBR QMF patching algorithm, such as that described in section 4.6.18.6.3 of the MPEG-4 AAC standard, is used, a pre-flattening step may be performed (as indicated by the "bs_sbr_preprocessing" parameter) to avoid discontinuities in the shape of the spectral envelope of the high-frequency signal input to the subsequent envelope tuner (the envelope tuner performs another stage of eSBR processing). Pre-flattening typically improves the operation of subsequent envelope tuner stages, resulting in a high-band signal that is perceived as more stable.
[0070] The overall bitrate requirement for including eSBR metadata (such as harmonic transposition and pre-flattening) in an MPEG-4 AAC bitstream is expected to be on the order of several hundred bits per second, because only the differential control data required to perform the eSBR processing is transmitted according to some embodiments of the present invention. Legacy decoders can ignore this information because it is included in a backward-compatible manner (as described later). Therefore, the negative bitrate impact associated with including eSBR metadata can be ignored for many reasons, including: • The bitrate penalty (due to including eSBR metadata) is only a small fraction of the total bitrate because only the differential control data necessary to perform the eSBR processing is transmitted (not a simulcast of SBR control data); and • Adjustments to SBR-related control information are typically independent of transposition details. Examples of cases where control data depends on the operation of the transposer are described later in this application.
[0071] Accordingly, embodiments of the present invention provide means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner. This efficient transmission of eSBR control data does not have a substantial adverse effect on the bitrate, while reducing memory requirements in decoders, encoders, and transcoders using embodiments of the present invention. Furthermore, the complexity and processing conditions associated with performing eSBR according to embodiments of the present invention are also reduced because the SBR data only needs to be processed once and not simulcast, and simulcasting would occur if the eSBR were treated as a completely separate object type in MPEG-4AAC rather than being integrated in a backward-compatible manner into the MPEG-4AAC codec.
[0072] Next, with reference to Figure 7, elements of a block of an MPEG-4 AAC bitstream containing eSBR metadata ("raw_data_block") will be described according to several embodiments of the present invention. Figure 7 is a diagram of a block of an MPEG-4 AAC bitstream ("raw_data_block"), showing some of its segments.
[0073] A block of MPEG-4 AAC bitstream may contain at least one "single_channel_element()" (e.g., a single-channel element as shown in Figure 7) and / or at least one "channel_pair_element()" (not specifically shown in Figure 7, but may exist), which contains the audio data of an audio program. The block may also contain a number of "fill elements" (e.g., fill element 1 and / or fill element 2 in Figure 7) that contain data related to the program (e.g., metadata). Each "single_channel_element()" contains an identifier indicating the start of a single-channel element (e.g., "ID1" in Figure 7) and may contain audio data indicating different channels in a multi-channel audio program. Each "channel_pair_element()" contains an identifier indicating the start of a channel-pair element (not shown in Figure 7) and may contain audio data indicating two channels in the program.
[0074] The "fill_element" (referred to hereafter as the fill element) of an MPEG-4 AAC bitstream includes an identifier ("ID2" in Figure 7) indicating the start of the fill element, and the fill data following the identifier. The identifier ID2 can consist of three unsigned integers ("uimsbf") transmitted bit-first with a value of 0x6. The fill data can include an "extension_payload()" element (often referred to hereafter as the extension payload), the syntax of which is shown in Table 4.57 of the MPEG-4 AAC standard. Several types of extension payloads exist, identified by the "extension_type" parameter, which is a four-bit unsigned integer ("uimsbf") transmitted bit-first with a value of 0x6.
[0075] Fill data (e.g., its extension payload) can include a header or identifier (e.g., "Header 1" in Figure 7) that indicates a segment of the fill data that represents an SBR object (i.e., the header initializes an "SBR object" type called sbr_extension_data() in the MPEG-4 AAC standard). For example, a spectral band replication (SBR) extension payload is identified by a value of '1101' or '1110' for the extension_type field in the header, where the identifier '1101' identifies an extension payload that has SBR data, and '1110' identifies an extension payload that has SBR data along with cyclic redundancy check (CRC) to verify the validity of the SBR data.
[0076] When a header (e.g., the extension_type field) initializes an SBR object type, the header is followed by SBR metadata (referred to here as "spectral band replica data," and to the MPEG-4 AAC standard as "sbr_data()"), and at least one spectral band replica extension element (e.g., the "SBR extension element" in fill element 1 of Figure 7) may follow the SBR metadata. Such a spectral band replica extension element (a segment of the bitstream) is referred to as an "sbr_extension()" container in the MPEG-4 AAC standard. A spectral band replica extension element optionally includes a header (e.g., the "SBR extension header" in fill element 1 of Figure 7).
[0077] The MPEG-4AAC standard assumes that spectral band replication extensions can contain parametric stereo (PS) data for the program's audio data. The MPEG-4AAC standard assumes that the header of a fill element (e.g., the header of the extension payload) initializes an SBR object type (as shown in "Header 1" in Figure 7), and that if the spectral band replication element of the fill element contains PS data, the fill element (e.g., its extension payload) contains the spectral band replication data and a "bs_extension_id" parameter, the value of which (i.e., bs_extension_id=2) indicates that the PS data is contained within the spectral band replication extension of the fill element.
[0078] According to some embodiments of the present invention, eSBR metadata (e.g., a flag indicating whether or not enhanced spectral band replication (eSBR) processing should be performed on the audio content of a block) is included in the spectral band replication extension element of a fill element. For example, such a flag is shown in fill element 1 in Figure 7, in which case the flag appears after the header of the "SBR extension element" of fill element 1 (the "SBR extension header" of fill element 1). Optionally, such a flag and additional eSBR metadata are included in the spectral band replication extension element after the header of the spectral band replication extension element (the SBR extension element of fill element 1 in Figure 7, after the SBR extension header). According to some embodiments of the present invention, the fill element containing the SBR metadata also includes a "bs_extension_id" parameter, the value of which (e.g., bs_extension_id = 3) indicates that eSBR metadata is included in the fill element and that eSBR processing should be performed on the audio content of the relevant block.
[0079] According to some embodiments of the present invention, the eSBR metadata is contained within a fill element of the MPEG-4 AAC bitstream (e.g., fill element 2 in Figure 7), rather than within a spectral band replication extension element (SBR extension element) of the fill element. This is because a fill element containing extension_payload() having SBR data or SBR data with CRC does not contain any other extension payload of any other extension type. Thus, in embodiments in which the eSBR metadata is stored in its own extension payload, a separate fill element is used to store the eSBR metadata. Such a fill element includes an identifier indicating the start of the fill element (e.g., "ID2" in Figure 7) and fill data following that identifier. The fill data includes an extension_payload() element (often referred to as the extension payload in this application), the syntax of which is shown in Table 4.57 of the MPEG-4 AAC standard. The fill data (e.g., its extension payload) includes a header (e.g., Header 2 in Fill Element 2 in Figure 7) that indicates an eSBR object (i.e., the header initializes the Enhanced Spectral Band Reproduction (eSBR) object type), and the fill data (e.g., its extension payload) includes eSBR metadata after the header. For example, Fill Element 2 in Figure 7 includes such a header ("Header 2"), and after the header, includes eSBR metadata (i.e., the "Flags" of Fill Element 2, which indicate whether Enhanced Spectral Band Reproduction (eSBR) processing should be performed on the audio content of the block). Optionally, additional eSBR metadata is also included in the fill data of Fill Element 2 in Figure 7, after Header 2. In the embodiments described in this paragraph, the header (e.g., Header 2 in Figure 7) has an identifier value that indicates an eSBR extension payload rather than one of the conventional values specified in Table 4.57 of the MPEG-4 AAC standard (consequently, the extension_type field of the header indicates that the fill data includes eSBR metadata).
[0080] In a first-class embodiment, the present invention is an audio processing unit (e.g., a decoder): A memory configured to store at least one block of the encoded audio bitstream (for example, at least one block of the MPEG-4 AAC bitstream) (e.g., buffer 201 in Figure 3 or 4); A bitstream payload deformator (e.g., element 205 in Figure 3 or element 215 in Figure 4) coupled to memory and configured to demultiplex at least a portion of a block of bitstreams; and A decryption subsystem (e.g., elements 202 and 203 in Figure 3, or elements 202 and 213 in Figure 4) configured and coupled to decode at least a portion of the audio content of a block of bitstream; The block includes: A fill element comprising an identifier indicating the start of a fill element and fill data following the identifier (for example, the identifier "id_syn_ele" has the value 0x6 in Table 4.85 of the MPEG-4 AAC standard), wherein the fill data is: It includes at least one flag that identifies whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of a block (e.g., using eSBR metadata and spectral band replication data contained in the block).
[0081] The flag is eSBR metadata, and an example of a flag is the sbrPatchingModeflag. Another example of a flag is the harmonicSBR flag. Both of these flags indicate whether the basic form of spectral band replication, or an enhanced form of spectral replication, is performed on the audio data of a block. The basic form of spectral replication is spectral patching, and the enhanced form of spectral band replication is harmonic transposition.
[0082] In some embodiments, the fill data also includes additional eSBR metadata (i.e., eSBR metadata other than flags).
[0083] The memory may be a buffer memory (e.g., an implementation of buffer 201 in Figure 4) that stores at least one block of the encoded audio bitstream (e.g., in a non-temporary manner).
[0084] The performance details of eSBR processing by an eSBR decoder (using eSBR harmonic transposition and pre-flattening) during decoding an MPEG-4 AAC bitstream containing eSBR metadata (shown in the eSBR tools below) would be as follows (for typical decoding with specified parameters): ●Harmonic Transposition (16kbps, 14400 / 28800Hz) ○DFT-based: 3.68 WMOPS (weighted million operations per second) ○QMF base: 0.98 WMOPS ● QMF patch processing / pre-processing (pre-flattening): 0.1 WMOPS DFT-based transpositions are typically known to perform better than QMF-based transpositions for transients.
[0085] According to some embodiments of the present invention, a fill element (of an encoded audio bitstream) containing eSBR metadata is also, The value (for example, bs_extension_id=3) indicates that the eSBR metadata is included in the fill element and that eSBR processing should be performed on the audio content of the relevant block (for example, the bs_extension_id parameter), and / or The value (for example, bs_extension_id=2) includes a parameter (for example, the same "bs_extension_id" parameter) that indicates the fill element's sbr_extension() container contains PS data. For example, as shown in Table 1 below, a parameter with the value bs_extension_id=2 may indicate that the fill element's sbr_extension() container contains PS data, and a parameter with the value bs_extension_id=3 may indicate that the fill element's sbr_extension() container contains eSBR metadata. Table 1 [Table 1]
[0086] According to some embodiments of the present invention, the syntax for each spectral band replication extension element, including eSBR metadata and / or PS data, is as shown in Table 2 below (where "sbr_extension()" indicates the container which is the spectral band replication extension element, "bs_extension_id" is as shown in Table 1 above, "ps_data" indicates the PS data, and "esbr_data" indicates the eSBR metadata). Table 2 [Table 2]
[0087] In an exemplary embodiment, esbr_data() referenced in Table 2 above represents the values of the following metadata parameters: 1.1-bit metadata parameter "bs_sbr_preprocessing"; and 2. The above parameters for each channel ("ch") of the audio content of the encoded bitstream to be decoded: "sbrPatchingMode[ch]"; "sbrOversamplingFlag[ch]"; "sbrPitchInBinsFlag[ch]"; and "sbrPitchInBins[ch]". For example, in some embodiments, esbr_data() may have the syntax shown in Table 3 to represent these metadata parameters: Table 3 [Table 3-1] [Table 3-2]
[0088] The above syntax, as an extension to legacy decoders, enables the efficient implementation of enhanced forms of spectral band replication, such as harmonic transposition. Specifically, the eSBR data in Table 3 includes only the parameters required to perform the enhanced forms of spectral band replication that are not already supported in the bitstream and cannot be directly derived from parameters already supported in the bitstream. All other parameters and processing data required to perform the enhanced forms of spectral band replication are extracted from existing parameters at already defined locations in the bitstream.
[0089] For example, decoders compliant with MPEG-4HE-AAC or HE-AACv2 may be extended to include enhanced forms of spectral band duplication, such as harmonic transposition. This enhanced form of spectral band duplication is added to the basic form of spectral band duplication already supported by the decoder. In the case of decoders compliant with MPEG-4HE-AAC or HE-AACv2, the basic form of spectral band duplication is the QMF spectral patching SBR tool defined in section 4.6.18 of the MPEG-4AAC standard.
[0090] When performing enhanced spectral band replication, the enhanced HE-AAC decoder can reuse many bitstream parameters already included in the bitstream's SBR extension payload. Specific parameters that may be reused include, for example, various parameters that determine the master frequency band table. These parameters include bs_start_freq (a parameter that determines the start of the master frequency table parameters), bs_stop_freq (a parameter that determines the end of the master frequency table), bs_freq_scale (a parameter that determines the number of frequency bands per octave), and bs_alter_scale (a parameter that changes the scale of the frequency bands). Parameters that may be reused also include those that determine the noise band table (bs_noid_bands) and the limiter band table parameters (bs_limiter_bands). Thus, in various embodiments, at least some equivalent parameters specified in the USAC standard are omitted from the bitstream, thereby reducing control overhead in the bitstream. Typically, if a parameter specified in the AAC standard has an equivalent parameter specified in the USAC standard, the equivalent parameter specified in the USAC standard will have the same name as the parameter specified in the AAC standard, for example, envelope scale factor E OrigMapped (the density scale factor EOrigMapped However, the equivalent parameters specified in the USAC standard typically have different values, which are "adjusted" for the enhanced SBR processing specified in the USAC standard, rather than for the SBR processing specified in the AAC standard.
[0091] To improve the subjective quality of audio content with harmonic frequency structures and strong tonal characteristics, especially at low bitrates, the activation of enhanced SBR is recommended. The values of the corresponding bitstream elements (i.e., esbr_data()) may be determined in the encoder by controlling these tools and applying a signal-dependent classification mechanism. In general, using the harmonic patching method (sbrPatchingMode==1) is preferred when encoding musical signals at very low bitrates, in which case the core codec may be quite limited in audio bandwidth. This is especially true when these signals contain a prominent harmonic structure. Conversely, using the normal SBR patching method is preferred for speech and mixed signals because it provides better preservation of the temporal structure in speech.
[0092] To improve the performance of the harmonic transposer, a preprocessing step can be activated (bs_sbr_preprocessing ==1) that attempts to avoid introducing spectral discontinuities into the signal heading toward the subsequent envelope tuner. The tool's operation is beneficial for signal types where the coarse spectral envelope of the low-band signal used for high-frequency reconstruction exhibits large fluctuations in level.
[0093] To improve the transient response of harmonic SBR patching, it is possible to apply signal-adaptive frequency-domain oversampling (sbrOversamplingFlag==1). Although signal-adaptive frequency-domain oversampling increases the computational complexity of the transponder, it only benefits frames containing transients, so the use of this tool is controlled by the bitstream element, which is sent once per frame and once per independent SBR channel.
[0094] Decoders operating in the proposed enhanced SBR mode typically require the ability to switch between legacy and enhanced SBR patching. Therefore, depending on the decoder configuration, a delay may be introduced, which can have a duration equal to the duration of a single core audio frame. Typically, the delays for both legacy and enhanced SBR patching will be similar.
[0095] In addition to numerous parameters, when performing enhanced spectral band replication according to embodiments of the present invention, other data elements may also be reused by the enhanced HE-AAC decoder. For example, envelope data and noise floor data may also be extracted from bs_data_env (envelope scalefactors) and bs_noid_env (noise floor scalefactors) data and used during enhanced spectral band replication.
[0096] Essentially, these embodiments enable enhanced spectral band replication that requires as little extra transmit data as possible, by leveraging configuration parameters and envelope data already supported by legacy HE-AAC or HE-AACv2 decoders in the SBR extended payload. While the metadata was originally tailored to the basic HFR format (e.g., spectral transformation processing of an SBR), according to these embodiments, it is used for enhanced HFR formats (e.g., harmonic transposition of an eSBR). As previously mentioned, the metadata generally represents operating parameters intended and tailored for use with the basic HFR format (e.g., linear spectral transformation) (operating parameters include, for example, envelope scale factor, noise floor scale factor, time / frequency grid parameter, sinusoidal summation information, variable crossover frequency / band, inverse filtering mode, envelope resolution, smoothing mode, and frequency interpolation mode). However, this metadata can be combined with additional metadata parameters specific to HFR-enhanced formats (e.g., harmonic transposition) to be used to efficiently and effectively process audio data using HFR-enhanced formats.
[0097] Therefore, an extended decoder supporting the enhanced form of spectral band replication can be created in a highly efficient manner by relying on already defined bitstream elements (e.g., elements in the SBR extension payload) and adding only the parameters necessary to support the enhanced form of spectral band replication (to the fill element extension payload). This data reduction property is combined with placing newly added parameters in reserved data fields such as extension containers, which significantly reduces the barriers to producing decoders that support the enhanced form of spectral band replication by ensuring that the bitstream is backward compatible with legacy decoders that do not support the enhanced form of spectral band replication. Reserved data fields are understood to be backward compatible data fields, i.e., data fields already supported by previous decoders, such as legacy HE-AAC or HE-AACv2 decoders. Similarly, extension containers are backward compatible, i.e., extension containers already supported by previous decoders, such as legacy HE-AAC or HE-AACv2 decoders. In Table 3, the numbers in the right column indicate the number of bits for the corresponding parameter in the left column.
[0098] In some embodiments, the SBR object type defined in MPEG-4AAC is updated to include the features of the extended SBR (eSBR) tool and the SBR-Tool, as indicated by the SBR extension element (bs_extension_id==EXTENSION_ID_ESBR). When the decoder detects this SBR extension element, the decoder uses the notified features of the extended SBR tool.
[0099] In some embodiments, the present invention is a method comprising the step of encoding audio data to generate an encoded bitstream (e.g., an MPEG-4 AAC bitstream), wherein eSBR metadata is included in at least one segment of at least one block of the encoded bitstream, and audio data is included in at least one other segment of the block. In a typical embodiment, the method comprises the step of multiplexing the audio data and eSBR metadata in each block of the encoded bitstream. In a typical decoding of the encoded bitstream in an eSBR decoder, the decoder extracts the eSBR metadata from the bitstream (including separating and analyzing the eSBR metadata and audio data), processes the audio data using the eSBR metadata, and generates a stream of decoded audio data.
[0100] Another aspect of the present invention is an eSBR decoder configured to perform eSBR processing (for example, using at least one of the eSBR tools known as harmonic transposition or pre-flattening) during the decoding of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that does not contain eSBR metadata. An example of such a decoder is illustrated with reference to Figure 5.
[0101] The eSBR decoder (400) in Figure 5 includes a buffer memory 201 (same as memory 201 in Figures 3 and 4), a bitstream payload deformatter 215 (same as deformatter 215 in Figure 4), an audio decoding subsystem 202 (sometimes called the "core" decoding stage or "core" decoding subsystem, and the same as the core decoding subsystem 202 in Figure 3), an eSBR control data generation subsystem 401, and an eSBR processing stage 203 (same as stage 203 in Figure 3), connected as shown. Typically, the decoder 400 also includes other processing elements (not shown).
[0102] In the operation of the decoder 400, a sequence of blocks of the encoded audio bitstream (MPEG-4 AAC bitstream) received by the decoder 400 is asserted from buffer 201 to deformatter 215.
[0103] The deformator 215 is configured and combined to separate each block of the bitstream and extract SBR metadata (including quantized envelope data) and typically other metadata from it. The deformator 215 is configured to assert at least the SBR metadata to the eSBR processing stage 203. The deformator 215 is also configured and combined to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.
[0104] The audio decoding subsystem 202 of decoder 400 is configured to decode the audio data extracted by the deformator 215 (such decoding may be referred to as the “core” decoding process) to produce decoded audio data and assert the decoded audio data to the eSBR processing stage 203. Decoding is performed in the frequency domain. Typically, the final stage of processing in subsystem 202 applies a frequency-to-time domain conversion to the decoded frequency-domain audio data, so that the output of subsystem is the decoded audio data in the time domain. Stage 203 is configured to apply the SBR tools (and eSBR tools) indicated by the SBR metadata (extracted by the deformator 215) and the eSBR metadata generated in subsystem 401 to the decoded audio data (i.e., perform SBR and eSBR processing on the output of decoding subsystem 202 using the SBR and eSBR metadata) to produce fully decoded audio data output from decoder 400. Typically, the decoder 400 includes memory (accessible by subsystem 202 and stage 203) for storing deformatted audio data and metadata outputs from the deformatter 215 (and optionally subsystem 401), and stage 203 is configured to access the audio data and metadata as needed during SBR and eSBR processing. The SBR processing in stage 203 may be considered post-processing on the output of the core decoding subsystem 202. Optionally, the decoder 400 also includes a final upmixing subsystem (which can apply the parametric stereo (PS) tools defined in the MPEG-4 AAC standard using the PS metadata extracted by the deformatter 215), which is configured to perform upmixing on the output of stage 203 and combine to produce fully decoded upmixed audio output from the APU 210.
[0105] Parametric stereo is a coding tool that represents a stereo signal using a linear downmix of the left and right channels of the stereo signal and a set of spatial parameters that describe the stereo image. Parametric stereo typically uses three types of spatial parameters: (1) the inter-channel intensity difference (IID) that describes the intensity difference between channels; (2) the inter-channel phase difference (IPD) that describes the phase difference between channels; and (3) the inter-channel coherence (ICC) that describes the coherence (or similarity) between channels. Coherence may be measured as the maximum value of the cross-correlation as a function of time or phase. These three parameters generally enable a high-quality reconstruction of the stereo image. However, the IPD parameter only specifies the relative phase difference between the channels of the stereo input signal and does not show the distribution of these phase differences with respect to the left and right channels. Therefore, a fourth type of parameter that describes the overall phase offset or overall phase difference (OPD) may additionally be used. In the stereo reconstruction process, consecutive window segments of both the received downmix signal s[n] and the uncorrelated version d[n] of the received downmix are processed together with the spatial parameters to generate the reconstructed signals for the left (l k (n)) and right (r k (n)) as follows: l k (n)=H 11 (k,n)s k (n)+H 21 (k,n)d k (n) r k (n)=H 12 (k,n)s k (n)+H 22 (k,n)d k (n) 0]Here, H 11 , H 12 , H 21 and H 22 are defined by the stereo parameters. The signal l k (n) and the signal r k(n) is ultimately transformed into the time domain by a frequency-time conversion.
[0106] The control data generation subsystem 401 in Figure 5 is configured and combined to detect at least one characteristic of the encoded audio bitstream to be decoded and to generate eSBR control data (which in other embodiments of the invention may be or include any type of eSBR metadata contained in the encoded audio bitstream) in accordance with at least one result of the detection step. When a particular characteristic (or combination of characteristics) of the bitstream is detected, the eSBR control data is asserted to stage 203 to trigger the application of individual eSBR tools or combinations of eSBR tools, and / or to control the application of such eSBR tools. For example, to control the performance of eSBR processing using harmonic transposition, several embodiments of the control data generation subsystem 401 would include: a music detector that sets (and asserts to stage 203) the sbrPatchingMode[ch] parameter in response to detecting whether the bitstream represents music; a transient detector that sets (and asserts to stage 203) the sbrOversamplingFlag[ch] parameter in response to detecting the presence or absence of transients in the audio content represented by the bitstream; and / or a pitch detector that sets (and asserts to stage 203) the sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch] parameters in response to detecting the pitch of the audio content represented by the bitstream. Another aspect of the present invention is an audio bitstream decoding method performed by any embodiment of the decoder of the present invention described in this paragraph and the preceding paragraphs.
[0107] Aspects of the present invention include a coding or decoding method of a type configured (e.g., programmed) to be performed by any embodiment of the APU, system, or device of the present invention. Other aspects of the present invention include a system or device configured (e.g., programmed) to perform any embodiment of the method of the present invention, and a computer-readable medium (e.g., disk) for storing (e.g., in a non-temporary manner) code for performing any embodiment of the method or any step thereof of the present invention. For example, the system of the present invention may be or include a programmable general-purpose processor, digital signal processor, or microprocessor separately configured to perform any various operations (including embodiments of the method or steps thereof) on data programmed and / or programmed in software or firmware. Such a general-purpose processor may be or include a computer system (programmed to perform embodiments of the method (or steps thereof) of the present invention in response to data being asserted) including input devices, memory, and processing circuits.
[0108] Embodiments of the present invention can be implemented in hardware, firmware, or software, or a combination thereof (e.g., a programmable logic array). Unless otherwise specified, the algorithms or processes included as part of the present invention are not essentially associated with any particular computer or other device. In particular, various general-purpose machines can be used with programs written in accordance with the teachings herein, or it may be more meaningful to construct more specialized devices (e.g., integrated circuits) to perform the required method steps. Accordingly, the present invention can be implemented in one or more computer programs to run in one or more programmable computer systems (e.g., implementations of any of the elements of Figure 1, or encoder 100 (or its elements) in Figure 2, or decoder 200 (or its elements) in Figure 3, or decoder 210 (or its elements) in Figure 4, or decoder 400 (or its elements) in Figure 5), each including at least one processor, at least one data storage system (including volatile and non-volatile memory and / or memory elements), at least one input device or port, and at least one output device or port. The program code is applied to input data to perform the functions described in this application and generate output information. The output information is applied to one or more output devices in a known manner.
[0109] Each such program can be implemented in any desired computer language (including machine, assembly, or high-level procedural, logical, or object-oriented programming languages) to communicate with a computer system. In any case, the language may be a compiled or interpreted language.
[0110] For example, when implemented by computer software instruction sequences, various functions and steps of the embodiments of the present invention can be realized by multithreaded software instruction sequences operated on appropriate digital signal processing hardware, in which case various devices, steps, and functions of the embodiments may correspond to some of the software instructions.
[0111] Each such computer program is preferably stored or downloaded to a storage medium or device (e.g., solid-state memory or medium, or magnetic or optical medium) that can be read by a general-purpose or dedicated programmable computer in order to configure and operate the computer when the storage medium or device is read by the computer system to perform the procedures described herein. The system of the present invention can be implemented as a computer-readable storage medium configured with (i.e., stored with) computer programs, the storage medium thus configured to operate the computer system in a specific predetermined manner in order to perform the functions described herein.
[0112] Many embodiments of the present invention have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the invention. Many modifications and variations of the invention are possible in light of the above teachings. For example, to facilitate efficient implementation, a phase shift may be used in combination with a complex QMF analysis and synthesis filter bank. The analysis filter bank is responsible for filtering the time-domain low-band signal generated by the core decoder into multiple subbands (e.g., QMF subbands). The synthesis filter bank is responsible for generating a broadband output audio signal by combining the regenerated high band generated by the selected HFR technique with the decoded low band (as indicated by the received sbrPatchingMode parameter). However, a given filter bank implementation operating in a particular sample rate mode, e.g., normal dual-rate operation or down-sampling SBR mode, should not have a bitstream-dependent phase shift. The QMF bank used in SBR is a complex-exponential extension of the theory of cosine modulation filter banks. It can be shown that extending the cosine modulation filter bank with complex-exponential modulation means that alias cancellation constraints are no longer used. Therefore, in the SBR QMF bank, the analysis filter h k (n) and composite filter f k Both (n) can be defined as follows:
number
[0113] The coefficients p0(n) of the prototype filter can be defined with a length L of 640, as shown in Table 4 below. Table 4 [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5] [Table 4-6] The prototype filter, p0(n), can also be derived from Table 4 by one or more mathematical operations such as rounding, subsampling, interpolation, and decimation.
[0114] The tuning of control information related to the SBR typically does not depend on the transposition details (as described above), but in some embodiments, certain elements of the control data may be simulcast in an eSBR extension container (bs_extension_id==EXTENSION_ID_ESBR) to improve the quality of the regenerated signal. Some of the elements that are simulcast may include noise floor data (e.g., noise floor scale factor and parameters indicating the direction in either the frequency or time direction for delta coding for each noise floor), inverse filtering data (e.g., parameters indicating an inverse filtering mode selected from no inverse filtering, low-level inverse filtering, intermediate-level inverse filtering, and strong-level inverse inverse filtering), and missing harmonics data (e.g., parameters indicating whether a sine wave should be added to a particular frequency band of the regenerated high band). All of these elements depend on the composite emulation of the decoder's transposer performed within the encoder, and therefore, if properly tuned for the selected transposer, it is possible to increase the quality of the regenerated signal.
[0115] Specifically, in some embodiments, missing harmonic and inverse filtering control data (along with other bitstream parameters in Table 3) are transmitted in an eSBR extension container and adjusted for the eSBR's harmonic transposer. For the eSBR's harmonic transposer, the additional bitrate required to transmit these two classes of metadata is relatively small. Therefore, transmitting the adjusted missing harmonic and / or inverse filtering control data in an eSBR extension container will increase the audio quality produced by the transposer with minimal impact on the bitrate. To ensure backward compatibility with legacy decoders, parameters adjusted for the SBR's spectral transformation processing may be transmitted in the bitstream as part of the SBR control data using either implicit or explicit signaling.
[0116] Within the scope of the attached claims, it should be understood that the present invention may be carried out in ways other than those specifically described herein. Any reference numbers that may be included in the following claims are for illustrative purposes only and should not be used in any way to interpret or limit the claims. Various aspects of this disclosure will be understood from the exemplary forms (EEEs) listed below:
[0117] EEE1. A method for performing high-frequency reconstruction of an audio signal: A step of receiving an encoded audio bitstream, wherein the encoded audio bitstream includes audio data representing the low-band portion of the audio signal and high-frequency reconstructed metadata; and A step of decoding the audio data in order to generate a decoded lowband audio signal; A step of extracting the high-frequency reconstruction metadata from the encoded audio bitstream, wherein the high-frequency reconstruction metadata includes operational parameters for the high-frequency reconstruction process, the operational parameters include patching mode parameters located within the extended container of the encoded audio bitstream, the first value of the patching mode parameter indicates spectral transformation, and the second value of the patching mode parameter indicates harmonic transposition by phase vocoder frequency spread; A step of filtering the decoded lowband audio signal in order to generate a filtered lowband audio signal; A step of regenerating the high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein the regeneration includes spectral transformation when the patch processing mode parameter is a first value, and includes harmonic transposition by phase vocoder frequency spread when the patch processing mode parameter is a second value; and A method comprising the step of combining the filtered low-band audio signal and the regenerated high-band portion to form a broadband audio signal. Eez2. The method according to EEE1, wherein the extended container includes inverse filtering control data to be used when the patch processing mode parameter is equal to the second value. EEE3. The method according to any one of EEE1 to EEE2, wherein the extended container further includes missing harmonic control data to be used when the patch processing mode parameter is equal to the second value. EEE4. The method of any of the preceding EEEs, wherein the encoded audio bitstream includes a fill element (having an identifier indicating the beginning of the fill element) and fill data following the identifier, the fill data including the extension container. EEE5. The identifier is a 3-bit unsigned integer transmitted in most significant bit first, having the value 0x6, as described in EEE4. EEE6. The aforementioned fill data includes an extended payload, the extended payload includes spectral band replication extended data, the extended payload is identified by a 4-bit unsigned integer transmitted most significant bit first, having the value '1101' or '1110', and optionally the spectral band replication extended data is: Optional spectral band replication header, Spectral band replication data following the header, and Spectrum-band replication extension element after the aforementioned spectral-band replication data The method according to EEE4 or EEE5, comprising a flag in the spectral band replication extension element. EEE7. The method according to any one of EEE1 to 6, wherein the high-frequency reconstruction metadata includes an envelope scale factor, a noise floor scale factor, time / frequency grid information, or a parameter indicating a crossover frequency. Eee8. The filtering described above is an analysis filter h, which is a modulated version of the prototype filter p0(n). k The analysis filter bank containing (n) is executed according to the following formula:
number
[0118] [Patent Document 1] U.S. Patent Application Publication No. 2018 / 0025737
Claims
1. A method for performing high-frequency reconstruction of an audio signal: A step of receiving an encoded audio bitstream, wherein the encoded audio bitstream includes audio data representing the low-band portion of the audio signal and high-frequency reconstruction metadata, the high-frequency reconstruction metadata includes an envelope scale factor; A step of decoding the audio data in order to generate a decoded lowband audio signal; A step of extracting the high-frequency reconstruction metadata from the encoded audio bitstream, wherein the high-frequency reconstruction metadata includes operational parameters for a high-frequency reconstruction process, the operational parameters include patching mode parameters located within a backward-compatible extension container of the encoded audio bitstream, the first value of the patching mode parameter indicating spectral transformation, and the second value of the patching mode parameter indicating harmonic transposition by phase vocoder frequency spread; The steps of filtering the decoded lowband audio signal in order to generate a filtered lowband audio signal; and A step of regenerating the high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein the regeneration includes spectral transformation when the patching mode parameter is a first value, and includes harmonic transposition by phase vocoder frequency spread when the patching mode parameter is a second value; A method comprising, wherein certain elements of the control data are simulcast in the backward-compatible extension container, and the certain elements of the control data include at least one of inverse filtering control data, noise floor control data, and missing harmonic control data for use in regenerating the signal.
2. The method according to claim 1, wherein the backward-compatible extension container includes inverse filtering control data to be used when the patch processing mode parameter is equal to the second value.
3. The method according to claim 1, wherein the backward-compatible extension container further includes missing harmonic control data to be used when the patching mode parameter is equal to the second value.
4. The method according to claim 1, wherein the phase shift is added to the filtered lowband audio signal after the filtering and compensated for before synthesis to reduce the complexity of the method.
5. A non-temporary computer-readable storage medium containing instructions that, when executed by a processor, cause the method according to claim 1 to be performed.
6. An audio processing unit that performs high-frequency reconstruction of an audio signal, comprising: An input interface for receiving an encoded audio bitstream, wherein the encoded audio bitstream includes audio data representing the low-band portion of the audio signal and high-frequency reconstruction metadata, the high-frequency reconstruction metadata including an envelope scale factor; A core audio decoder that decodes the audio data to generate a decoded lowband audio signal; A deformator for extracting the high-frequency reconstruction metadata from the encoded audio bitstream, wherein the high-frequency reconstruction metadata includes operational parameters for the high-frequency reconstruction process, and the operational parameters include patching mode parameters located within a backward-compatible extension container of the encoded audio bitstream; An analysis filter bank for filtering the decoded lowband audio signal in order to generate a filtered lowband audio signal; and A high-frequency regeneration unit that regenerates the high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein the regeneration includes spectral transformation when the patch processing mode parameter is a first value, and includes harmonic transposition by phase vocoder frequency spread when the patch processing mode parameter is a second value; An audio processing unit comprising a backward-compatible extension container in which certain elements of control data are simulcast, wherein the certain elements of the control data include at least one of inverse filtering control data, noise floor control data, and missing harmonic control data for use in regenerating the signal.
Citation Information
Patent Citations
Apparatus and method for improved amplitude response and temporal alignment in a bandwidth expansion method based on a phase vocoder for audio signals.
JP2013521536A
Apparatus and method for processing audio signals using patch boundary matching
JP2013521538A
Apparatus and method for decoding encoded audio signals using low computational resources
JP2016539377A
Apparatus and method for generating enhanced signals with independent noise filling
JP2017526004A
Backward-compatible integration of high-frequency reconstruction techniques for audio signals
JP2021507316A