Method for performing high-frequency reconstruction of an audio signal and audio processing unit
The method improves high-frequency reconstruction in audio encoding by regenerating high-band audio signals using metadata, addressing inefficiencies in existing techniques and enhancing audio quality for music content.
Patent Information
- Application Number
- JP2023192034
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-01-26
- Filing Date
- 2023-11-10
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2039-01-28
AI Technical Summary
Existing audio encoding techniques, such as MPEG-4 AAC, face challenges with spectral band replication (SBR) that are not ideal for certain audio types, particularly music content with low crossover frequencies, necessitating improved methods for high-frequency reconstruction.
The method involves decoding an encoded audio bitstream, extracting high-frequency reconstruction metadata, filtering the low-band audio signal, and regenerating the high-band portion using the metadata, with options for spectral conversion or harmonic transposition based on flags, and combining the signals to form a wideband audio signal.
This approach enhances audio quality by efficiently reconstructing high-frequency components, improving compression efficiency and maintaining spectral characteristics, especially for music content, while being backward compatible with legacy decoders.
Smart Images

Figure 0007711898000017 
Figure 0007711898000018 
Figure 0007711898000019
Abstract
Description
Background Art
[0001] Cross - reference to related applications This application claims priority based on the following applications, which are incorporated herein by reference: U.S. Provisional Application No. 62 / 622,205, filed on January 26, 2018.
[0002] Technical field Embodiments relate to audio signal processing, and more specifically, to the encoding, decoding, or transcoding of an audio bitstream by control data that specifies that either a basic form of high-frequency reconstruction (HFR) or an enhanced form of HFR is to be performed on audio data.
[0003] Background of the invention A typical audio bitstream includes both audio data (e.g., encoded audio data) representing one or more channels of audio content and metadata representing at least one characteristic of the audio data or audio content. One well-known format for generating an encoded audio bitstream is the MPEG-4 Advanced Audio Coding (AAC) format, which is described in the MPEG standard ISO / IEC 14496-3:2009. In the MPEG4 standard, AAC means "Advanced Audio Coding" and HE-AAC means "High Efficiency Advanced Audio Coding".
[0004] The MPEG-4 AAC standard defines several audio profiles that determine the objects and encoding tools present in an encoder or decoder that conforms to the standard. Three of these audio profiles are: (1) the AAC profile, (2) the HE-AAC profile, and (3) the HE-AAC v2 profile. The AAC profile includes the AAC Low Complexity (or "AAC-LC") object type. The AAC-LC object corresponds to the MPEG-2 AAC Low Complexity profile with some adjustments and does not include the Spectral Band Replication ("SBR") object type nor the Parametric Stereo ("PS") object type. The HE-AAC profile is a superset of the AAC profile and additionally includes the SBR object type. The HE-AAC v2 profile is a superset of the HE-AAC profile and additionally includes the PS object type.
[0005] The SBR object type includes the Spectral Band Replication tool, which is an important High Frequency Reconstruction ("HFR") encoding tool that significantly improves the compression efficiency of perceptual audio coders. SBR reconstructs the high frequency components of the audio signal on the receiver side (e.g., in the decoder). Thus, the encoder only needs to encode and transmit the low frequency components, enabling very high audio quality at a low data rate. SBR is based on replicating a sequence of harmonics that have been truncated in advance to reduce the data rate from the control data obtained from the encoder and the available bandwidth limitation signal. The ratio between tonal and noise components is maintained by adaptive inverse filtering in addition to the selective addition of noise and sine waves. In the MPEG-4 AAC standard, the SBR tool performs spectral patching (also called linear or spectral transformation), where a number of consecutive Quadrature Mirror Filter (QMF) subbands are copied (or "patched") from the transmitted low band portion of the audio signal to the high band portion of the audio signal generated at the decoder.
[0006] Spectral patch processing or linear transformation may not be ideal for certain audio types, such as music content with relatively low crossover frequencies. Therefore, techniques for improving spectral band replication are needed. SUMMARY OF THE INVENTION
[0007] Regarding a first class of embodiments, a method for decoding an encoded audio bitstream is disclosed. The method includes receiving the encoded audio bitstream and decoding the audio data to generate a decoded low-band audio signal. The method further includes extracting high-frequency reconstruction metadata, filtering the decoded low-band audio signal with an analysis filter bank to generate a filtered low-band audio signal, extracting a flag indicating whether spectral conversion or harmonic transposition should be performed on the audio data, and regenerating the high-band portion of the audio signal using the high-frequency reconstruction metadata and the filtered low-band audio signal according to the flag. Finally, the method includes combining the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal.
[0008] The second class of embodiments relates to an audio decoder for decoding an encoded audio bit stream. The decoder includes an input interface for receiving the encoded audio bit stream (the encoded audio bit stream includes audio data representing the low-band portion of the audio signal), and a core decoder for decoding the audio data to generate a decoded low-band audio signal. The decoder also includes a demultiplexer for extracting from the encoded audio bit stream high-frequency reconstruction metadata (the high-frequency reconstruction metadata includes operating parameters for a high-frequency reconstruction process that linearly transforms a consecutive number of subbands from the low-band portion of the audio signal to the high-band portion of the audio signal), and an analysis filter bank for filtering the decoded low-band audio signal to generate a filtered low-band audio signal. The decoder further includes a demultiplexer for extracting from the encoded audio bit stream a flag indicating whether linear transformation or harmonic transposition is to be performed on the audio data, and a high-frequency regenerator for regenerating the high-band portion of the audio signal using the high-frequency reconstruction metadata and the filtered low-band audio signal according to the flag. Finally, the decoder includes a synthesis filter bank for combining the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal.
[0009] Other classes of embodiments relate to encoding and transcoding an audio bit stream that includes metadata for identifying whether an enhanced spectral band replication (eSBR) process is to be performed. BRIEF DESCRIPTION OF THE DRAWINGS
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
[0011] Notation and terms Throughout this disclosure, including the claims, the expression "performing a process on" a signal or data (e.g., filtering, scaling, converting, or applying a gain to the signal or data) is used in a broad sense to indicate performing a direct process on the signal or data, or a processed version of the signal or data (e.g., a version of the signal that has been pre-filtered or pre-processed prior to performing that process).
[0012] Throughout this disclosure, including the claims, the terms "audio processing unit" or "audio processor" are used broadly to refer to a system, device, or apparatus configured to process audio data. Examples of audio processing units include, but are not limited to, encoders, transcoders, decoders, codecs, preprocessing systems, postprocessing systems, and bitstream processing systems (often referred to as bitstream processing tools). Virtually all consumer electronics products, such as mobile phones, televisions, laptops, and tablet computers, incorporate an audio processing unit or audio processor.
[0013] Throughout this disclosure, including the claims, the terms "coupled" or "couples" are used broadly to mean a direct or indirect connection. Thus, when a first device is coupled to a second device, that connection may be through a direct connection or through an indirect connection via other devices and connections. Additionally, components integrated within or together with other components are also coupled to each other.
[0014] Detailed description of the embodiments of the invention The MPEG-4 AAC standard is assumed to include metadata that indicates each type of high-frequency reconstruction (HFR) processing to be applied (if any) by a decoder to decode the audio content of a bitstream of encoded MPEG-4 AAC, and / or controls such HFR processing, and / or indicates at least one characteristic or parameter of at least one HFR tool to be used to decode the audio content of the bitstream. Here, we use the expression "SBR metadata" to denote this type of metadata described or referred to in the MPEG-4 AAC standard for use in spectral band replication ("SBR"). As will be understood by those skilled in the art, SBR is a form of HFR.
[0015] SBR is preferably used as a dual-rate system, where the underlying codec operates at half the original sampling rate while SBR operates at the original sampling rate. The SBR encoder operates in parallel with the underlying core codec, albeit at a higher sampling rate. SBR is mainly post-processing in the decoder, but important parameters are extracted in the encoder to compensate for the highest accuracy high-frequency reconstruction in the decoder. The encoder estimates the spectral envelope of the SBR range for a time and frequency range / resolution suitable for the current input signal segment characteristics. The spectral envelope is estimated by complex QMF analysis and subsequent energy calculation. The time and frequency resolution of the spectral envelope can be selected with a high degree of freedom to ensure the optimal time-frequency resolution for a given input segment. Envelope estimation needs to take into account that transients located in the original region, mainly the high-frequency region (e.g., a high-hat), are slightly present in the SBR high band generated before envelope adjustment, because the high band in the decoder is based on the low band and the transients are judged to be much smaller compared to the high band. This aspect imposes different conditions on the time-frequency resolution of the spectral envelope data compared to normal spectral envelope estimation as used in other audio coding algorithms.
[0016] Apart from the spectral envelope, several additional parameters representing the spectral characteristics of the input signal for different time and frequency domains are extracted. The encoder, of course, has access not only to the original signal but also to information on how the SBR unit in the decoder generates the high band, so that not only in situations where the low band constitutes a strong harmonic series and the regenerated high band mainly consists of random signal components, but also in situations where there are strong tone components in the original high band for which there is no corresponding low band counterpart in the high band region, the system is able to handle. Furthermore, the SBR encoder operates in close relation to the underlying core codec and evaluates which frequency ranges should be covered by SBR at a given time. The SBR data, in the case of a stereo signal, is efficiently encoded prior to transmission by utilizing not only the channel dependence of the control data but also entropy coding.
[0017] The control parameter extraction algorithm typically requires careful adjustment to the underlying codec at a given bitrate and a given sampling rate. This is due to the fact that lower bitrates usually exhibit a larger SBR range compared to higher bitrates, and different sampling rates correspond to different temporal resolutions of the SBR frame.
[0018] The SBR decoder typically includes several different parts. These include a bitstream decoding module, a high-frequency reconstruction module (HFR), an additional high-frequency component module, and an envelope adjustment module. The system is based on a complex-valued QMF filter bank (for high-quality SBR) or a real-valued QMF filter bank (for low-power SBR). Embodiments of the present invention are applicable to both high-quality SBR and low-power SBR. In the bitstream extraction module, control data is read from and decoded from the bitstream. The time-frequency grid is obtained for the current frame before reading the envelope data from the bitstream. The underlying core decoder decodes the audio signal of the current frame (at a lower sampling rate) and generates audio samples in the time domain. The resulting frame of audio data is used for high-frequency reconstruction by the HFR module. The decoded low-band signal is then analyzed using a QMF filter bank. Thereafter, high-frequency reconstruction and envelope adjustment are performed on the subband samples of the QMF filter bank. The high frequencies are reconstructed from the low band in a flexible manner based on the given control parameters. Further, the reconstructed high band is adaptively filtered on a subband-channel basis according to the control data to ensure appropriate spectral characteristics in the given time / frequency region.
[0019] The top level of the MPEG-4 AAC bitstream is a sequence of data blocks ("raw_data_block" elements), each of which is a segment of data (hereinafter referred to as a "block") containing audio data (typically over a period of 1024 or 960 samples) and related information and / or other data. Here, we use the term "block" to denote a segment of the MPEG-4 AAC bitstream that contains one (and no more than one) "raw_data_block" element of audio data (and corresponding metadata and optionally other related data).
[0020] Each block of the MPEG-4 AAC bitstream can contain a number of syntax elements (each of which also appears in the bitstream as a segment of data). Seven types of such syntax elements are defined in the MPEG-4 AAC standard. Each syntax element is identified by a different value of the data element "id_syn_ele". Specific examples of syntax elements include "single_channel_element()", "channel_pair_element()", and "fill_element()". A single channel element is a container (mono audio signal) that contains audio data for a single audio channel. A channel pair element contains audio data for two audio channels (stereo audio signal).
[0021] The fill element is a container of information that contains an identifier (e.g., the value of the id_syn_ele element above) followed by data, which is referred to as "fill data". Historically, the fill element has been used to adjust the instantaneous bitrate of a bitstream transmitted over a channel at a constant rate. By adding an appropriate amount of fill data to each block, it is possible to achieve a constant data rate.
[0022] According to an embodiment of the present invention, the fill data may include one or more extension payloads that extend the types of data that can be transmitted in the bitstream (e.g., metadata). A decoder that receives the bitstream with fill data containing a new type of data may optionally be used by a device (e.g., a decoder) that receives the bitstream to extend the functionality of the device. Thus, as will be understood by those skilled in the art, the fill element is a special type of data structure and is different from the data structures typically used to transmit audio data (e.g., an audio payload containing channel data).
[0023] In some embodiments of the present invention, the identifier used to identify a fill element may consist of a three bit unsigned integer transmitted most significant bit first (uimsbf) having a value of 0x6. In one block, several instances of the same type of syntax element (e.g., several fill elements) may occur.
[0024] Another standard for encoding an audio bit stream is the MPEG (Unified Speech and Audio Coding: USAC) standard (ISO / IEC 23003-3:2012). The MPEG USAC standard describes the encoding and decoding of audio content using spectral band replication processing (including SBR processing as described in the MPEG-4 AAC standard, and also including other enhanced forms of spectral band replication processing). This processing applies an extended and enhanced version of the spectral band replication tool (often referred to herein as the "extended SBR tool" or "eSBR tool") of a group of SBR tools described in the MPEG-4 AAC standard. Thus, eSBR (as defined in the USAC standard) is an improvement over SBR (as defined in the MPEG-4 AAC standard).
[0025] Herein, the expression "enhanced SBR processing" (or "eSBR processing") is used to represent spectral band replication processing that uses at least one eSBR tool not described or referred to in the MPEG-4 AAC standard (e.g., at least one eSBR tool described or referred to in the MPEG USAC standard). Examples of such eSBR tools are harmonic transposition and QMF patch processing preprocessing or "pre-flattening".
[0026] A harmonic transposer of integer order T maps a sine wave of frequency ω to a sine wave of frequency Tω while maintaining the signal duration. Typically, three orders, T = 2, 3, and 4, are used in sequence to generate each part of the desired output frequency range using the smallest possible transposition order. If an output above the fourth-order transposition range is required, it may be generated by a frequency shift. If possible, a baseband time domain that is almost critically sampled is created for processing to minimize the computational complexity.
[0027] The harmonic transposer may be either QMF or DFT based. When using a QMF-based harmonic transposer, the bandwidth expansion of the core coder time domain signal is performed entirely within the QMF domain using a modified phase vocoder structure, and time stretching is performed after decimation for all QMF subbands. Transpositions using several transposition factors (e.g., T = 2, 3, 4) are performed in a common QMF analysis / synthesis transform stage. Since the QMF-based harmonic transposer does not feature signal-adaptive frequency domain oversampling, the corresponding flag (sbrOversamplingFlag[ch]) in the bitstream may be ignored.
[0028] When using a DFT-based harmonic transposer, transposers of factors 3 and 4 (third- and fourth-order transposers) are preferably incorporated into the factor-2 transposer (second-order transposer) by interpolation to reduce complexity. For each frame (corresponding to the core coder samples of coreCoderFrameLength), the nominal "full-size" transform size of the transposer is first determined by the signal-adaptive frequency domain oversampling flag (sbrOverSamplingFlag[ch]) in the bitstream.
[0029] When sbrPatchingMode == 1, it indicates that linear transposition should be used to generate the high band, and additional steps may be introduced to avoid discontinuities in the shape of the spectral envelope of the high-frequency signal input to the subsequent envelope adjuster. This improves the operation of the next envelope adjustment stage and, as a result, produces a high-band signal that is perceived as more stable. The additional preprocessing operation is beneficial for signal types where the rough spectral envelope of the low-band signal used for high-frequency reconstruction exhibits large level variations. However, the values of the bitstream elements may be determined within the encoder by applying any kind of signal-dependent classification. The additional preprocessing is preferably activated by a 1-bit per bitstream element bs_sbr_preprocessing. When bs_sbr_preprocessing is set to 1, the additional processing is enabled. When bs_sbr_preprocessing is set to zero, the additional processing is disabled. The additional processing is the low band X for each patch Low It is preferable to utilize the pre-gain curve used by the high-frequency generator to scale
Equation
Equation
Equation
[0030] A bitstream generated according to the MPEG USAC standard (often referred to as the "USAC bitstream" in this application) contains the encoded audio content and typically metadata indicating each type of spectral band replication process applied by a decoder to decode the audio content of the USAC bitstream, and / or metadata controlling such spectral band replication process and / or indicating at least one characteristic or parameter of at least one SBR tool and / or eSBR tool used to decode the audio content of the USAC bitstream.
[0031] In the present application, the expression "enhanced SBR metadata" (or eSBR metadata) is used to indicate metadata that is not described or mentioned in the MPEG4AAC standard, but which indicates each type of spectral band replication process applied by a decoder to decode the audio content of an encoded audio bitstream (e.g., a USAC bitstream), and / or controls such spectral band replication process, and / or indicates at least one characteristic or parameter of at least one SBR tool and / or eSBR tool used to decode such audio content. Specific examples of eSBR metadata include metadata (for indicating or controlling spectral band replication processes) that is described or mentioned in the MPEG USAC standard but not in the MPEG-4AAC standard. Accordingly, eSBR metadata in the present application refers to metadata that is not SBR metadata, and SBR metadata in the present application refers to metadata that is not eSBR metadata.
[0032] A USAC bitstream may contain both SBR metadata and eSBR metadata. More specifically, a USAC bitstream may contain eSBR metadata for controlling the performance of eSBR processing by a decoder and SBR metadata for controlling the performance of SBR processing by a decoder. According to an exemplary embodiment of the present invention, eSBR metadata (e.g., configuration data specific to eSBR) is included in an MPEG-4AAC bitstream (e.g., the sbr_extension() container at the end of the SBR payload) (according to the present invention).
[0033] During decoding of an encoded bitstream using an eSBR tool set (including at least one eSBR tool), execution of eSBR processing by a decoder regenerates a high-frequency band of an audio signal based on replication of a sequence of harmonics discarded during encoding. Such eSBR processing typically adjusts the spectral envelope of the generated high-frequency band, applies inverse filtering, and adds noise and sine wave components to reproduce the spectral characteristics of the original audio signal.
[0034] According to an exemplary embodiment of the present invention, eSBR metadata is included in one or more metadata segments of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that includes audio data encoded in other segments (audio data segments). Typically, at least one such metadata segment of each block of the bitstream is a fill element (including an identifier indicating the start of the fill element) (or includes it), and the eSBR metadata is included in the fill element after the identifier.
[0035] FIG. 1 is a block diagram of an exemplary audio processing chain (audio data processing system), and one or more elements of the system can be configured according to an embodiment of the present invention. The system includes elements that are coupled together as shown as encoder 1, distribution subsystem 2, decoder 3, and post-processing unit 4. In variations of the illustrated system, one or more elements are omitted or additional audio data processing units are included.
[0036] In some implementations, encoder 1 (optionally including a preprocessing unit) is configured to receive PCM (time domain) samples including audio content as input and output an encoded audio bitstream (having a format compliant with the MPEG-4 AAC standard) representing the audio content. The data of the bitstream representing the audio content is often referred to herein as "audio data" or "encoded audio data". When the encoder is configured according to a typical embodiment of the present invention, the audio bitstream output from the encoder includes (and typically also other metadata) eSBR metadata as well as audio data.
[0037] One or more encoded audio bitstreams output from encoder 1 may be asserted to an encoded audio distribution subsystem 2. Subsystem 2 is configured to store and / or distribute each encoded bitstream output from encoder 1. The encoded audio bitstream output from encoder 1 may be stored by subsystem 2 (e.g., in the form of a DVD or Blu-ray disc), or transmitted by subsystem 2 (capable of realizing a transmission link or network), or both stored and transmitted by subsystem 2.
[0038] Decoder 3 is configured to decode an encoded MPEG-4 AAC audio bitstream (generated by encoder 1) received via subsystem 2. In some embodiments, decoder 3 extracts eSBR metadata from each block of the bitstream, decodes the bitstream (including performing eSBR processing using the extracted eSBR metadata), and is configured to generate decoded audio data (e.g., a stream of decoded PCM audio samples). In some embodiments, decoder 3 extracts SBR metadata from the bitstream (ignoring eSBR metadata included in the bitstream), decodes the bitstream (including performing SBR processing using the extracted SBR metadata), and is configured to generate decoded audio data (e.g., a stream of decoded PCM audio samples). Typically, decoder 3 includes a buffer that stores (e.g., in a non-transitory manner) segments of the encoded audio bitstream received from subsystem 2.
[0039] The post-processing unit 4 of FIG. 1 is configured to receive a stream of decoded audio data (e.g., decoded PCM audio samples) from decoder 3 and perform post-processing thereon. The post-processing unit may also be configured to render the post-processed audio content (or the decoded audio received from decoder 3) for playback by one or more speakers.
[0040] Figure 2 is a block diagram of an encoder (100) which is an embodiment of the audio processing unit of the present invention. Any of the components or elements of encoder 100 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware and software. Encoder 100 includes an encoder 105, a stuffer / formatter stage 107, a metadata generation stage 106, and a buffer memory 109 connected as shown. Typically, encoder 100 also includes other processing elements (not shown). Encoder 100 is configured to convert an input audio bitstream into an encoded output MPEG-4 AAC bitstream.
[0041] The metadata generation unit 106 is configured and coupled to generate (and / or pass through stage 107) the metadata (including eSBR metadata and SBR metadata) to be included by stage 107 in the encoded bitstream to be output from encoder 100.
[0042] Encoder 105 is configured and coupled to assert to stage 107 to encode the input audio data (e.g., by performing compression thereon) and include the resulting encoded audio in the encoded bitstream to be output from stage 107.
[0043] Stage 107 is configured to multiplex the encoded audio from encoder 105 and the metadata (including eSBR metadata and SBR metadata) from generation unit 106 to generate an encoded bitstream output from stage 107, and preferably, the encoded bitstream is configured to have a format as specified by any of the embodiments of the present invention.
[0044] Buffer memory 109 is configured to store (e.g., in a non-transitory manner) at least one block of the encoded audio bitstream output from stage 107, and then a sequence of blocks of the encoded audio bitstream is asserted from buffer memory 109 as an output from encoder 100 to the delivery system.
[0045] FIG. 3 is a block diagram of a system including a decoder (200) that is an embodiment of the audio processing unit of the present invention and, optionally, a post-processor (300) coupled thereto. Any of the components or elements of decoder 200 and post-processor 300 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware and software. Decoder 200 includes buffer memory 201, bitstream payload deframer 205, audio decoding subsystem 202 (sometimes referred to as the “core” decoding stage or “core” decoding subsystem), eSBR processing stage 203, and control bit generation stage 204 connected in the form shown. Typically, decoder 200 also includes other processing elements (not shown).
[0046] Buffer memory (buffer) 201 stores at least one block of the encoded MPEG-4 AAC audio bitstream received by decoder 200 (e.g., in a non-transitory form). In the operation of decoder 200, a sequence of blocks of the bitstream is asserted from buffer 201 to deframer 205.
[0047] In a modification of the embodiment of FIG. 3 (or the embodiment of FIG. 4 described later), an APU that is not a decoder (for example, the APU 500 in FIG. 6) includes a buffer memory (for example, the same buffer memory as the buffer 201), and the buffer memory stores at least one block of the same type of encoded audio bitstream (for example, an MPEG-4 AAC audio bitstream) (i.e., an encoded audio bitstream including eSBR metadata) received by the buffer 201 in FIG. 3 or FIG. 4 (for example, in a non-transitory manner).
[0048] Referring to FIG. 3 again, the de-formatter 205 is configured and coupled to demultiplex (or separate) each block of the bitstream, extract SBR metadata (including quantized envelope data) and eSBR metadata (and typically other metadata) therefrom, assert at least the eSBR metadata and the SBR metadata to the eSBR processing stage 203, and typically also assert the other extracted metadata to the decoding subsystem 202 (and, optionally, also control the bit generation unit 204). The de-formatter 205 is also configured and coupled to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.
[0049] The system of FIG. 3 also includes, as an option, a post-processor 300. The post-processor 300 includes a buffer memory (buffer) 301 and other processing elements (not shown) including at least one processing element coupled to the buffer 301. The buffer 301 stores (in a non-transitory manner) at least one block (or frame) of the decoded audio data received by the post-processor 300 from the decoder 200. The processing elements of the post-processor 300 are configured and coupled to receive and adaptively process a sequence of blocks of decoded audio output from the buffer 301 using metadata output from the decoding subsystem 202 (and / or the de-formatter 205) and / or control bits output from stage 204 of the decoder 200.
[0050] The audio decoding subsystem 202 of the decoder 200 decodes the audio data extracted by the parser 205 (such decoding may be referred to as "core" decoding processing), generates the decoded audio data, and is configured to assert the decoded audio data to the eSBR processing stage 203. The decoding is performed in the frequency domain and typically includes inverse quantization followed by spectral processing. Typically, the final stage of processing in the subsystem 202 applies the conversion from the frequency domain to the time domain to the decoded frequency domain audio data, and as a result, the output of the subsystem is the decoded audio data in the time domain. The stage 203 applies the SBR tool and eSBR tool indicated by the eSBR metadata and eSBR (extracted by the parser 205) to the decoded audio data (i.e., performs SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata), and is configured to generate the fully decoded audio data output from the decoder 200 (e.g., to the post-processor 300). Typically, the decoder 200 includes a memory (accessible by the subsystem 202 and the stage 203) that stores the metadata and the de-formatted audio data from the de-formatter 205, and the stage 203 is configured to access the audio data and metadata (including the SBR metadata and eSBR metadata) as needed during the SBR and eSBR processing. The SBR processing and eSBR processing in the stage 203 may be considered as post-processing on the output of the core decoding subsystem 202.As an option, the decoder 200 is also configured to perform an upmix on the output of stage 203 and is coupled to generate fully decoded and upmixed audio output from the decoder 200, including a final upmixing subsystem (which can apply parametric stereo (PS) tools defined in the MPEG-4 AAC standard using the PS metadata extracted by the de-formatter 205 and / or the control bits generated by subsystem 204). Alternatively, the post-processor 300 is configured to perform an upmix on the output of the decoder 200 (e.g., using the PS metadata extracted by the de-formatter 205 and / or the control bits generated in subsystem 204).
[0051] In response to the metadata extracted by the de-formatter 205, the control bit generator 204 is capable of generating control data that is used within the decoder 200 (e.g., in the final upmixing subsystem) and / or can be asserted as an output of the decoder 200 (e.g., to the post-processor 300 for use in post-processing). In response to the metadata extracted from the input bitstream (and optionally in response to control data), stage 204 is capable of generating (and asserting to the post-processor 300) control bits indicating that the decoded audio data output from the eSBR processing stage 203 should undergo a particular type of post-processing. In some implementations, the decoder 200 is configured to assert the metadata extracted by the de-formatter 205 from the input bitstream to the post-processor 300, and the post-processor 300 is configured to perform post-processing on the decoded audio data output from the decoder 200 using the metadata.
[0052] FIG. 4 is a block diagram of an audio processing unit (``APU'') (210) which is another embodiment of the audio processing unit of the present invention. APU 210 is a legacy decoder not configured to perform eSBR processing. Any of the components or elements of APU 210 may be implemented as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit) in hardware, software, or a combination of hardware. APU 210 includes a buffer memory 201, a bitstream payload deformer (parser) 215, an audio decoding subsystem 202 (often referred to as the ``core'' decoding stage or the ``core'' decoding subsystem), and an SBR processing stage 213 connected in the form shown. Typically, APU 210 includes other processing elements (not shown). APU 210 may represent, for example, an audio encoder, decoder, or transcoder.
[0053] Elements 201 and 202 of APU 210 are identical to the similarly numbered elements of decoder 200 (of FIG. 3), and their above description will not be repeated. In the operation of APU 210, a sequence of blocks of an encoded audio bitstream (MPEG-4 AAC bitstream) received by APU 210 is asserted from buffer 201 to deformer 215.
[0054] The formatter 215 demultiplexes each block of the bitstream and extracts therefrom SBR metadata (including quantized envelope data) and typically other metadata, but is configured and coupled to ignore eSBR metadata that may be included in the bitstream according to some embodiments of the present invention. The formatter 215 is configured to assert at least the SBR metadata to the SBR processing stage 213. The formatter 215 is also configured and coupled to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.
[0055] The audio decoding subsystem 202 of decoder 200 decodes the audio data extracted by the de-formatter 215 to generate decoded audio data (such decoding may be referred to as "core" decoding processing), and is configured to assert the decoded audio data to the SBR processing stage 213. Decoding is performed in the frequency domain. Typically, the final stage of processing in subsystem 202 applies the conversion from the frequency domain to the time domain to the decoded frequency domain audio data, so that the output of the subsystem is the decoded audio data in the time domain. Stage 213 applies the SBR tools (but not eSBR tools) indicated by the SBR metadata (extracted by the de-formatter 215) to the decoded audio data (i.e., performs SBR processing on the output of the decoding subsystem 202 using the SBR metadata), and is configured to generate the fully decoded audio data output from the APU 210 (e.g., to the post-processor 300). Typically, the APU 210 includes a memory (accessible by the subsystem 202 and stage 213) that stores the de-formatted audio data and the metadata output from the de-formatter 215, and stage 213 is configured to access the audio data and metadata (including SBR metadata) as needed during SBR processing. The SBR processing in stage 213 may be considered as post-processing on the output of the core decoding subsystem 202. Optionally, the APU 210 is also configured to perform upmixing on the output of stage 213 to generate the fully decoded upmixed audio output from the APU 210, and includes a combined final upmixing subsystem (which can apply the parametric stereo (PS) tools defined in the MPEG-4 AAC standard using the PS metadata extracted by the de-formatter 215).Alternatively, the post-processor is configured to perform upmixing on the output of the APU 210 (e.g., using the PS metadata extracted by the de-formatter 215 and / or the control bits generated at the APU 210).
[0056] The various implementations of the encoder 100, the decoder 200, and the APU 210 are configured to perform the various embodiments of the method of the present invention.
[0057] According to some embodiments, the eSBR metadata is included in an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) (e.g., including a few control bits that are the eSBR metadata), such that a legacy decoder (not configured to analyze the eSBR metadata or use any eSBR tool related to the eSBR metadata) can ignore the eSBR metadata, yet still decode the bitstream to the extent possible without any significant penalty in the typically decoded audio quality, without using any eSBR tool or the eSBR metadata related thereto. However, an eSBR decoder configured to analyze the bitstream to identify the eSBR metadata and use at least one eSBR tool according to the eSBR metadata will enjoy the audio type that utilizes at least one such eSBR tool. Accordingly, embodiments of the present invention provide means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner.
[0058] Typically, the eSBR metadata in the bitstream indicates one or more of the following eSBR tools (which are those described in the MPEG USAC standard and may or may not be applied by the encoder during generation of the bitstream): (e.g., indicating at least one characteristic or parameter thereof). ·Harmonic transposition; and ·QMF patching additional pre-processing (pre-flattening).
[0059] For example, the eSBR metadata included in the bitstream may indicate parameter values such as sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch], and bs_sbr_preprocessing (these are described in the MPEG USAC standard and this disclosure).
[0060] Here, when X is some parameter, the notation X[ch] indicates that the parameter is related to the channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, we often omit the expression [ch] and assume that the relevant parameter is related to the channel of the audio content.
[0061] Here, when X is some parameter, the notation X[ch][env] indicates that the parameter is related to the SBR envelope ("env") of the channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, we often omit the expressions [env] and [ch] and assume that the relevant parameter is related to the SBR envelope of the channel of the audio content.
[0062] During decoding of the symbolized bitstream, the performance of harmonic transposition during the eSBR processing stage of decoding (for each channel "ch" of the audio content represented by the bitstream) is controlled by the following eSBR metadata parameters: sbrPatchingMode[ch]; sbrOversamplingFlag[ch]; sbrPitchInBinsFlag[ch]; and sbrPitchInBins[ch].
[0063] The value "sbrPatchingMode[ch]" indicates the type of transposer used in eSBR: sbrPatchingMode[ch]=1 indicates linear transposition patching as described in section 4.6.18 of the MPEG-4 AAC standard (such as used in high-quality SBR or low-power SBR); sbrPatchingMode[ch]=0 indicates harmonic SBR patching as described in section 7.5.3 or 7.5.4 of the MPEG USAC standard.
[0064] The value "sbrOversamplingFlag[ch]" indicates the use of signal-adaptive frequency-domain oversampling in eSBR in combination with DFT-based harmonic SBR patching as described in section 7.5.3 of the MPEG USAC standard. This flag controls the size of the DFT used by the transposer: 1 indicates that signal-adaptive frequency-domain oversampling is enabled as described in section 7.5.3.1 of the MPEG USAC standard; 0 indicates that signal-adaptive frequency-domain oversampling is disabled as described in section 7.5.3.1 of the MPEG USAC standard.
[0065] The value "sbrPitchInBinsFlag[ch]" controls the interpretation of the sbrPitchInBins[ch] parameter: 1 indicates that the value of sbrPitchInBins[ch] is valid and greater than zero; 0 indicates that the value of sbrPitchInBins[ch] is set to zero.
[0066] The value "sbrPitchInBins[ch]" controls the addition of cross product terms in the SBR harmonic transposer. The value sbrPitchinBins[ch] is an integer value within the range [0,127] and represents the distance measured in frequency bins of a 1536-line DFT that acts on the core coder's sampling frequency.
[0067] If the MPEG-4 AAC bitstream indicates an SBR channel pair and those channels are not combined (i.e., not a single SBR channel), the bitstream indicates two instances of the above syntax (for harmonic or non-harmonic transposition), one for each channel of sbr_channel_pair_element().
[0068] The harmonic transposition of the eSBR tool typically improves the quality of music signals decoded at relatively low crossover frequencies. Non-harmonic transposition (i.e., legacy spectral patch processing) typically improves speech signals. Thus, the starting point in determining which type of transposition is preferable for encoding a particular audio content is to select the transposition method depending on speech / music detection, with harmonic transposition being used for music content and spectral patch processing being used for speech content.
[0069] The performance of the pre-flattening during eSBR processing is controlled by the value of a 1-bit eSBR metadata parameter known as "bs_sbr_preprocessing" in the sense that pre-flattening is either performed or not depending on this single-bit value. When using the SBR QMF patch processing algorithm as described in section 4.6.18.6.3 of the MPEG-4 AAC standard, a pre-flattening step may be executed (when indicated by the "bs_sbr_preprocessing" parameter) to avoid discontinuities in the spectral envelope shape of the high-frequency signal input to the subsequent envelope shaper (the envelope shaper performs another stage of the eSBR processing). Pre-flattening typically improves the operation of the subsequent envelope shaping stage, resulting in a high-band signal that is perceived as more stable.
[0070] The overall bitrate conditions for including eSBR metadata indicating the above-described eSBR tools (such as harmonic transposition and pre-flattening) in the MPEG-4 AAC bitstream are expected to be on the order of several hundred bits per second because only the differential control data required to perform the eSBR processing is transmitted according to some embodiments of the present invention. Legacy decoders can ignore this information because it is included in a backward-compatible manner (as described later). Therefore, the negative impact on the bitrate associated with including eSBR metadata can be ignored for many reasons including: · The bitrate penalty (due to including eSBR metadata) is only a very small part of the total bitrate because only the differential control data required to perform the eSBR processing is transmitted (not simulcasting of SBR control data); and · The adjustment of SBR-related control information typically does not depend on the details of the transposition. Examples of cases where the control data depends on the operation of the transposer are described later in this application.
[0071] Accordingly, embodiments of the present invention provide means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner. This efficient transmission of eSBR control data does not have a substantial adverse impact on the bit rate, while reducing the memory requirements in decoders, encoders, and transcoders that use aspects of the present invention. Further, the complexity and processing conditions associated with performing eSBR according to embodiments of the present invention are also reduced, since SBR data only needs to be processed once instead of being simulcast, and in the case of simulcast, eSBR is not integrated in a backward-compatible manner with the MPEG-4 AAC codec, but rather is treated as a completely separate object type in MPEG-4 AAC.
[0072] Next, with reference to FIG. 7, elements of a block ("raw_data_block") of an MPEG-4 AAC bitstream that includes eSBR metadata according to some embodiments of the present invention are described. FIG. 7 is a diagram of a block ("raw_data_block") of an MPEG-4 AAC bitstream, showing some of its segments.
[0073] Blocks of an MPEG-4 AAC bitstream may contain at least one "single_channel_element()" (e.g., the single-channel element shown in FIG. 7) and / or at least one "channel_pair_element()" (not specifically shown in FIG. 7 but which may exist), which contain the audio data of an audio program. The block may also contain a number of "fill elements" (e.g., fill element 1 and / or fill element 2 in FIG. 7) containing program-related data (e.g., metadata). Each "single_channel_element()" contains an identifier indicating the start of a single-channel element (e.g., "ID1" in FIG. 7) and can contain audio data representing different channels of a multichannel audio program. Each "channe_pair_element()" contains an identifier (not shown in FIG. 7) indicating the start of a channel pair element and can contain audio data representing two channels of a program.
[0074] An "fill_element" of an MPEG-4 AAC bitstream (referred to as a fill element in this application) contains an identifier indicating the start of the fill element ("ID2" in FIG. 7) and the fill data following the identifier. The identifier ID2 can be composed of three unsigned integers ("uimsbf") transmitted most significant bit first with a value of 0x6. The fill data can contain an "extension_payload()" element (often referred to as an extended payload in this application) whose syntax is shown in Table 4.57 of the MPEG-4 AAC standard. There are several types of extended payloads, which are identified by an "extension_type" parameter that is a 4-bit unsigned integer ("uimsbf") transmitted most significant bit first.
[0075] The fill data (e.g., its extended payload) can include a header or identifier (e.g., "header 1" in FIG. 7) indicating a segment of the fill data that represents an SBR object (i.e., the header initializes the "SBR object" type called sbr_extension_data() in the MPEG-4 AAC standard). For example, a spectral band replication (SBR) extended payload is identified by a value of '1101' or '1110' for the extension_type field in the header, where the identifier '1101' identifies an extended payload with SBR data, and '1110' identifies an extended payload with SBR data along with a cyclic redundancy check (CRC) to verify the integrity of the SBR data.
[0076] When the header (e.g., the extension_type field) initializes the SBR object type, SBR metadata (referred to herein as "spectral band replication data" and called "sbr_data()" in the MPEG-4 AAC standard) follows the header, and at least one spectral band replication extension element (e.g., the "SBR extension element" of fill element 1 in FIG. 7) can follow the SBR metadata. Such a spectral band replication extension element (a segment of the bitstream) is called an "sbr_extension()" container in the MPEG-4 AAC standard. The spectral band replication extension element optionally includes a header (e.g., the "SBR extension header" of fill element 1 in FIG. 7).
[0077] The MPEG-4 AAC specification assumes that the spectral band replication extension element can contain PS (parametric stereo) data of the program's audio data. The MPEG-4 AAC specification assumes that when the header of the fill element (e.g., the header of the extended payload) initializes the SBR object type (such as "header 1" in FIG. 7) and the spectral band replication element of the fill element contains PS data, the fill element (e.g., its extended payload) contains spectral band replication data and the "bs_extension_id" parameter, and its value (i.e., bs_extension_id = 2) indicates that the PS data is included in the spectral band replication extension element of the fill element.
[0078] According to some embodiments of the present invention, eSBR metadata (e.g., a flag indicating whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of a block) is included in the spectral band replication extension element of the fill element. For example, such a flag is shown in fill element 1 of FIG. 7, in which case the flag occurs after the header of the "SBR extension element" of fill element 1 (the "SBR extension header" of fill element 1). Optionally, such a flag and additional eSBR metadata are included in the spectral band replication extension element after the header of the spectral band replication extension element (the SBR extension element of fill element 1 in FIG. 7 after the SBR extension header). According to some embodiments of the present invention, the fill element containing the SBR metadata also includes the "bs_extension_id" parameter, and its value (e.g., bs_extension_id = 3) indicates that the eSBR metadata is included in the fill element and that the eSBR processing should be performed on the audio content of the associated block.
[0079] According to some embodiments of the present invention, eSBR metadata is included not within the spectral band replication extension element (SBR extension element) of the fill element, but within the fill element (e.g., fill element 2 of FIG. 7) of the MPEG-4 AAC bitstream. This is because a fill element containing an extension_payload() with SBR data or SBR data with CRC does not contain any other extension payloads of any other extension type. Thus, in embodiments where eSBR metadata is stored in its own extension payload, a separate fill element is used to store the eSBR metadata. Such a fill element includes an identifier indicating the start of the fill element (e.g., "ID2" of FIG. 7) and fill data following that identifier. The fill data includes an extension_payload() element (sometimes referred to as an extension payload in this application), the syntax of which is shown in Table 4.57 of the MPEG-4 AAC standard. The fill data (e.g., its extension payload) includes a header (e.g., header 2 of fill element 2 of FIG. 7) indicating an eSBR object (i.e., the header initializes an enhanced spectral band replication (eSBR) object type), and the fill data (e.g., its extension payload) includes eSBR metadata following the header. For example, fill element 2 of FIG. 7 includes such a header ("header 2"), and following the header, includes eSBR metadata (i.e., the "flag" of fill element 2, which indicates whether an enhanced spectral band replication (eSBR) process should be performed on the audio content of the block). Optionally, additional eSBR metadata is also included in the fill data of fill element 2 of FIG. 7, following header 2. In the embodiments described in this paragraph, the header (e.g., header 2 of FIG. 7) has an identification value indicating an eSBR extension payload, rather than one of the conventional values defined in Table 4.57 of the MPEG-4 AAC standard (as a result, the extension_type field of the header indicates that the fill data contains eSBR metadata).
[0080] In a first class of embodiments, the present invention is an audio processing unit (e.g., a decoder): a memory (e.g., buffer 201 of FIG. 3 or 4) configured to store at least one block of an encoded audio bit stream (e.g., at least one block of an MPEG-4 AAC bit stream); a bit stream payload de-formatter (e.g., element 205 of FIG. 3 or element 215 of FIG. 4) coupled to the memory and configured to demultiplex at least a portion of the block of the bit stream; and a decoding subsystem (e.g., elements 202 and 203 of FIG. 3, or elements 202 and 213 of FIG. 4) coupled and configured to decode at least a portion of the audio content of the block of the bit stream; including, the block is: a fill element including an identifier indicating the start of the fill element and fill data after the identifier (e.g., the "id_syn_ele" identifier has the value 0x6 in Table 4.85 of the MPEG-4 AAC standard), the fill data is: including at least one flag (e.g., using eSBR metadata and spectral band replication data included in the block) for identifying whether enhanced spectral band replication (eSBR) processing is to be performed on the audio content of the block.
[0081] The flag is eSBR metadata, an example of the flag is the sbrPatchingModeflag. Another example of the flag is the harmonicSBR flag. Both of these flags indicate whether the basic form of spectral band replication, or the enhanced form of spectral replication, is to be performed on the audio data of the block. The basic form of spectral replication is spectral patching, and the enhanced form of spectral band replication is harmonic transposition.
[0082] In some embodiments, the fill data also includes additional eSBR metadata (i.e., eSBR metadata other than flags).
[0083] The memory may be a buffer memory (e.g., the implementation of buffer 201 in FIG. 4) that stores (e.g., in a non-transitory manner) at least one block of the encoded audio bitstream.
[0084] Details of the performance of eSBR processing (using eSBR harmonic transposition and pre-flattening) by an eSBR decoder during decoding of an MPEG-4 AAC bitstream that includes eSBR metadata (indicating the following eSBR tools) will be as follows (for typical decoding with the specified parameters): ● Harmonic transposition (16 kbps, 14400 / 28800 Hz) ○ DFT-based: 3.68 WMOPS (weighted million operations per second) ○ QMF-based: 0.98 WMOPS ● QMF patch processing - preprocessing (pre-flattening): 0.1 WMOPS DFT-based transposition is typically known to function better than QMF-based transposition for transients.
[0085] According to some embodiments of the present invention, the fill element (of the encoded audio bitstream) that includes eSBR metadata also parameters (e.g., the bs_extension_id parameter) whose value (e.g., bs_extension_id = 3) indicates that eSBR metadata is included in the fill element and that eSBR processing should be performed on the audio content of the relevant block, and / or The value (e.g., bs_extension_id = 2) includes a parameter (e.g., the same "bs_extension_id" parameter) indicating that the sbr_extension() container of the fill element contains PS data. For example, as shown in Table 1 below, a parameter having a value of bs_extension_id = 2 may indicate that the sbr_extension() container of the fill element contains PS data, and a parameter having a value of bs_extension_id = 3 may indicate that the sbr_extension() container of the fill element contains eSBR metadata. Table 1 [Table 1]
[0086] According to some embodiments of the present invention, the syntax of each spectral band replication extension element including eSBR metadata and / or PS data is as shown in Table 2 below ("sbr_extension()" indicates a container that is a spectral band replication extension element, "bs_extension_id" is as shown in Table 1 above, "ps_data" indicates PS data, and "esbr_data" indicates eSBR metadata). Table 2 [Table 2]
[0087] In an exemplary embodiment, esbr_data() referred to in Table 2 above indicates the values of the following metadata parameters: 1.1 bit metadata parameter "bs_sbr_preprocessing"; and 2. For each channel ("ch") of the audio content of the encoded bitstream to be decoded, the above-mentioned parameters: "sbrPatchingMode[ch]"; "sbrOversamplingFlag[ch]"; "sbrPitchInBinsFlag[ch]"; and "sbrPitchInBins[ch]". For example, in some embodiments, to indicate these metadata parameters, esbr_data() may have the syntax shown in Table 3: Table 3
Table 3-1
Table 3-2
[0088] The above syntax enables an efficient implementation of an enhanced form of spectral band replication, such as harmonic transposition, as an extension to legacy decoders. Specifically, the eSBR data in Table 3 includes only the parameters required to perform an enhanced form of spectral band replication that is not already supported in the bitstream and cannot be directly derived from parameters already supported in the bitstream. All other parameters and processing data required to perform the enhanced form of spectral band replication are extracted from existing parameters in predefined locations within the bitstream.
[0089] For example, a decoder compliant with MPEG-4 HE-AAC or HE-AACv2 may be extended to include an enhanced form of spectral band replication such as harmonic transposition. This enhanced form of spectral band replication is added to the basic form of spectral band replication already supported in the decoder. In the case of a decoder compliant with MPEG-4 HE-AAC or HE-AACv2, the basic form of spectral band replication is the QMF spectral patch processing SBR tool defined in section 4.6.18 of the MPEG-4 AAC standard.
[0090] When performing the enhanced form of spectral band replication, the extended HE-AAC decoder can reuse many of the bitstream parameters already included in the SBR extension payload of the bitstream. Specific parameters that may be reused include, for example, various parameters that determine the master frequency band table. These parameters include bs_start_freq (a parameter that determines the start of the master frequency table parameters), bs_stop_freq (a parameter that determines the end of the master frequency table), bs_freq_scale (a parameter that determines the number of frequency bands per octave), and bs_alter_scale (a parameter that changes the scale of the frequency bands). Parameters that may be reused include the parameters that determine the noise band table (bs_noid_bands) and the limiter band table parameters (bs_limiter_bands). Thus, in various embodiments, at least some of the equivalent parameters specified in the USAC standard are omitted from the bitstream, thereby reducing the control overhead in the bitstream. Typically, when the parameters specified in the AAC standard have equivalent parameters specified in the USAC standard, the equivalent parameters specified in the USAC standard have the same name as the parameters specified in the AAC standard, for example, the envelope scale factor E OrigMapped (the envelope scalefactor EOrigMapped ) However, the equivalent parameters specified in the USAC standard usually have different values, which are "adjusted" for the enhanced SBR process defined in the USAC standard rather than for the SBR process defined in the AAC standard.
[0091] To improve the subjective quality for audio content with harmonic frequency structure and strong tone characteristics, especially at low bitrates, the activation of enhanced SBR is recommended. The value of the corresponding bitstream element (i.e., esbr_data()) may be determined in the encoder by controlling these tools and applying a signal-dependent classification mechanism. Generally, using the harmonic patching method (sbrPatchingMode == 1) is preferred when encoding music signals at very low bitrates, where the core codec may be quite limited in the audio bandwidth. This is especially true when these signals contain a prominent harmonic structure. Conversely, using the normal SBR patching method is preferred for speech and mixed signals because it provides better preservation of the temporal structure in speech.
[0092] To improve the performance of the harmonic transposer, it is possible to activate a preprocessing step that attempts to avoid introducing spectral discontinuities in the signal going to the subsequent envelope adjuster (bs_sbr_preprocessing == 1). The operation of the tool is beneficial for signal types where the rough spectral envelope of the low-band signal used for high-frequency reconstruction exhibits large variations in level.
[0093] To improve the transient response of the harmonic SBR patch processing, it is possible to apply signal adaptive frequency domain oversampling (sbrOversamplingFlag==1). Signal adaptive frequency domain oversampling increases the computational complexity of the transponder, but since it only benefits frames containing transients, the use of this tool is controlled by bitstream elements, which are transmitted once per frame and once per independent SBR channel.
[0094] A decoder operating in the proposed enhanced SBR mode typically needs to be switchable between legacy and enhanced SBR patch processing. Therefore, depending on the decoder settings, a delay that can be as long as the duration of one core audio frame may be introduced. Typically, the delays for both legacy and enhanced SBR patch processing will be similar.
[0095] In addition to a number of parameters, when performing an enhanced form of spectral band replication according to an embodiment of the present invention, other data elements may also be reused by the extended HE-AAC decoder. For example, envelope data and noise floor data may also be extracted from the bs_data_env (envelope scalefactors) and bs_noid_env (noise floor scalefactors) data and used during the enhanced form of spectral band replication.
[0096] Essentially, these embodiments enable an enhanced form of spectral band replication that requires as little extra transmitted data as possible by utilizing configuration parameters and envelope data already supported by a legacy HE-AAC or HE-AACv2 decoder in an SBR extended payload. Metadata was originally tailored to a basic form of HFR (such as the spectral conversion process of SBR), but according to the embodiments, it is used for an enhanced form of HFR (such as the harmonic transposition of eSBR). As described above, metadata generally represents operating parameters (such as envelope scale factors, noise floor scale factors, time / frequency grid parameters, sine wave addition information, variable crossover frequencies / bands, inverse filtering modes, envelope resolution, smoothing modes, frequency interpolation modes) that are intended to be used with a basic form of HFR (such as linear spectral conversion). However, this metadata can be combined with additional metadata parameters specific to an enhanced form of HFR (such as harmonic transposition) to be used to efficiently and effectively process audio data using the enhanced form of HFR.
[0097] Accordingly, an extended decoder that supports an enhanced form of spectral band replication can be created in a very efficient way by relying on already defined bitstream elements (e.g., elements within the SBR extension payload) and adding only the parameters necessary to support the enhanced form of spectral band replication (to the fill element extension payload). This data reduction property is combined by placing the newly added parameters in reserved data fields such as the extension container, ensuring that the bitstream is backward compatible with legacy decoders that do not support the enhanced form of spectral band replication, thus considerably reducing the barrier in generating a decoder that supports the enhanced form of spectral band replication. The reserved data fields are backward compatible data fields, i.e., data fields already supported by previous decoders such as legacy HE-AAC or HE-AACv2 decoders. Similarly, the extension container is backward compatible, i.e., an extension container already supported by previous decoders such as legacy HE-AAC or HE-AACv2 decoders. In Table 3, the numbers in the right column indicate the number of bits of the corresponding parameters in the left column.
[0098] In some embodiments, the SBR object type defined in MPEG-4 AAC is updated to include the features of the extended SBR (eSBR) tool and the SBR-Tool, as indicated by the SBR extension element (bs_extension_id == EXTENSION_ID_ESBR). When the decoder detects this SBR extension element, the decoder uses the notified features of the extended SBR tool.
[0099] In some embodiments, the present invention is a method that includes encoding audio data to generate an encoded bitstream (e.g., an MPEG-4 AAC bitstream), including eSBR metadata in at least one segment of at least one block of the encoded bitstream, and including audio data in at least one other segment of the block. In a typical embodiment, the method includes multiplexing audio data and eSBR metadata in each block of the encoded bitstream. In a typical decoding of the encoded bitstream in an eSBR decoder, the decoder extracts the eSBR metadata from the bitstream (including separating and analyzing the eSBR metadata and the audio data), uses the eSBR metadata to process the audio data, and generates a stream of decoded audio data.
[0100] Another aspect of the present invention is an eSBR decoder configured to perform eSBR processing (e.g., at least one of the eSBR tools known as harmonic transposition or pre-flagging is used) during the decoding of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that does not include eSBR metadata. An example of such a decoder is described with reference to FIG. 5.
[0101] The eSBR decoder (400) of FIG. 5 includes a buffer memory 201 (the same as the memory 201 of FIGS. 3 and 4), a bitstream payload deframer 215 (the same as the deframer 215 of FIG. 4), an audio decoding subsystem 202 (sometimes referred to as a "core" decoding stage or "core" decoding subsystem, the same as the core decoding subsystem 202 of FIG. 3), an eSBR control data generation subsystem 401, and an eSBR processing stage 203 (the same as the stage 203 of FIG. 3) connected in the manner shown. Typically, the decoder 400 also includes other processing elements (not shown).
[0102] In the operation of the decoder 400, a sequence of blocks of an encoded audio bitstream (MPEG-4 AAC bitstream) received by the decoder 400 is asserted from the buffer 201 to the de-formatter 215.
[0103] The de-formatter 215 is configured and coupled to separate each block of the bitstream and extract therefrom SBR metadata (including quantized envelope data) and typically other metadata. The de-formatter 215 is configured to assert at least the SBR metadata to the eSBR processing stage 203. The de-formatter 215 is also configured and coupled to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.
[0104] The audio decoding subsystem 202 of decoder 400 decodes the audio data extracted by the de-formatter 215 (such decoding may be referred to as "core" decoding processing), generates decoded audio data, and is configured to assert the decoded audio data to the eSBR processing stage 203. The decoding is performed in the frequency domain. Typically, the final stage of processing in subsystem 202 applies a conversion from the frequency domain to the time domain to the decoded frequency domain audio data, so the output of the subsystem is the decoded audio data in the time domain. Stage 203 is configured to apply the SBR tools (and eSBR tools) indicated by the SBR metadata (extracted by the de-formatter 215) and the eSBR metadata generated by subsystem 401 to the decoded audio data (i.e., perform SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata) to generate the fully decoded audio data output from decoder 400. Typically, decoder 400 includes a memory (accessible by subsystem 202 and stage 203) that stores the de-formatted audio data and metadata output from the de-formatter 215 (and optionally subsystem 401), and stage 203 is configured to access the audio data and metadata as needed during SBR and eSBR processing. The SBR processing in stage 203 may be considered a post-processing of the output of the core decoding subsystem 202. Optionally, decoder 400 also includes a final upmixing subsystem (which can apply parametric stereo (PS) tools defined in the MPEG-4 AAC standard using the PS metadata extracted by the de-formatter 215), which is configured to perform upmixing on the output of stage 203 and is configured and combined to generate the fully decoded upmixed audio output from the APU 210.
[0105] Parametric stereo is a coding tool that represents a stereo signal using a linear downmix of the left and right channels of the stereo signal and a set of spatial parameters that describe the stereo image. Parametric stereo typically uses three types of spatial parameters: (1) the inter-channel intensity difference (IID) that describes the intensity difference between channels; (2) the inter-channel phase difference (IPD) that describes the phase difference between channels; and (3) the inter-channel coherence (ICC) that describes the coherence (or similarity) between channels. Coherence may be measured as the maximum value of the cross-correlation as a function of time or phase. These three parameters generally enable a high-quality reconstruction of the stereo image. However, the IPD parameter only specifies the relative phase difference between the channels of the stereo input signal and does not show the distribution of these phase differences with respect to the left and right channels. Therefore, a fourth type of parameter that describes the overall phase offset or overall phase difference (OPD) may additionally be used. In the stereo reconstruction process, consecutive window segments of both the received downmix signal s[n] and the uncorrelated version d[n] of the received downmix are processed together with the spatial parameters to generate the reconstructed signals for the left (l k (n)) and right (r k (n)) as follows: l k (n)=H 11 (k,n)s k (n)+H 21 (k,n)d k (n) r k (n)=H 12 (k,n)s k (n)+H 22 (k,n)d k (n) where H 11 , H 12 , H 21 and H 22 are defined by the stereo parameters. The signal l k (n) and the signal r k(n) is finally converted to the time domain by frequency-time conversion.
[0106] The control data generation subsystem 401 of FIG. 5 detects at least one characteristic of the encoded audio bitstream to be decoded and is configured and coupled to generate eSBR control data (in other embodiments of the present invention, it may be any type of eSBR metadata included in the encoded audio bitstream or may include it) in response to at least one result of the detection step. When a specific characteristic (or combination of characteristics) of the bitstream is detected, the eSBR control data is asserted at stage 203 to trigger the application of individual eSBR tools or a combination of eSBR tools and / or to control the application of such eSBR tools. For example, to control the performance of eSBR processing using harmonic transposition, some embodiments of the control data generation subsystem 401 include: a music detector that sets the sbrPatchingMode[ch] parameter (and asserts the set parameter at stage 203) in response to detecting whether the bitstream represents music; a transience detector that sets the sbrOversamplingFlag[ch] parameter (and asserts the set parameter at stage 203) in response to detecting the presence or absence of transients in the audio content represented by the bitstream; and / or a pitch detector that sets the sbrPitchInBinsFlag[ch] and sbrPitchInBins[ch] parameters (and asserts the set parameters at stage 203) in response to detecting the pitch of the audio content represented by the bitstream. Another aspect of the present invention is an audio bitstream decoding method performed by any embodiment of the decoder of the present invention described in this paragraph and the preceding paragraph.
[0107] Aspects of the present invention include a type of encoding or decoding method configured (e.g., programmed) to be executed by any embodiment of the APU, system, or device of the present invention. Other aspects of the present invention include a system or device configured (e.g., programmed) to execute any embodiment of the method of the present invention, and a computer-readable medium (e.g., a disk) storing code (e.g., in a non-transitory manner) for executing any embodiment of the method of the present invention or its steps. For example, the system of the present invention can be or include a programmable general-purpose processor, a digital signal processor, or a microprocessor programmed with software or firmware and / or separately configured to perform any of a variety of operations (including embodiments of the method of the present invention or its steps) on data. Such a general-purpose processor can be or include a computer system (programmed to execute an embodiment of the method of the present invention (or its steps) in response to data being asserted) that includes an input device, memory, and a processing circuit.
[0108] Embodiments of the present invention can be implemented in hardware, firmware, software, or a combination of both (e.g., a programmable logic array). Unless otherwise specified, algorithms or processes included as part of the present invention are not inherently associated with a particular computer or other device. In particular, various general-purpose machines can be used with programs written in accordance with the teachings herein, or it may be more meaningful to construct a more specialized apparatus (e.g., an integrated circuit) to perform the required method steps. Accordingly, the present invention may be implemented in one or more programmable computer systems (e.g., an implementation of any of the elements of FIG. 1, or the encoder 100 (or its elements) of FIG. 2, or the decoder 200 (or its elements) of FIG. 3, or the decoder 210 (or its elements) of FIG. 4, or the decoder 400 (or its elements) of FIG. 5), each of which includes at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port. Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices in a known manner.
[0109] Each such program can be implemented in any desired computer language (including machine, assembly, or high-level procedural, logical, or object-oriented programming languages) for communicating with the computer system. In any case, the language can be a compiled or interpreted language.
[0110] For example, when implemented by a computer software instruction sequence, the various functions and steps of the embodiments of the present invention can be realized by a multi-threaded software instruction sequence operated by appropriate digital signal processing hardware. In that case, the various devices, steps, and functions of the embodiments may correspond to a part of the software instructions.
[0111] Each such computer program is preferably stored or downloaded to a storage medium or device (e.g., a solid-state memory or medium, or a magnetic or optical medium) readable by a general-purpose or special-purpose programmable computer for configuring and operating a computer when the storage medium or device is read by the computer system to execute the procedures described herein. The system of the present invention can be implemented as a computer-readable storage medium configured (i.e., storing) with a computer program, and the storage medium thus configured causes the computer system to operate in a specific predetermined manner to execute the functions described herein.
[0112] Numerous embodiments of the present invention have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the invention. Many modifications and variations of the present invention are possible in light of the above teachings. For example, a phase shift may be used in combination with a complex QMF analysis and synthesis filter bank to facilitate an efficient implementation. The analysis filter bank serves to filter the time-domain low-band signal generated by the core decoder into a plurality of subbands (e.g., QMF subbands). The synthesis filter bank serves to generate a broadband output audio signal by combining the reproduced high-band generated by the selected HFR technique with the decoded low-band (as indicated by the received sbrPatchingMode parameter). However, a given filter bank implementation operating in a particular sample rate mode, e.g., normal dual-rate operation or downsampling SBR mode, should not have a phase shift that is dependent on the bitstream. The QMF bank used in SBR is a complex-exponential extension of the theory of cosine-modulated filter banks. It can be shown that when a cosine-modulated filter bank is extended with complex-exponential modulation, the alias cancellation constraint is no longer used. Thus, in the SBR QMF bank, both the analysis filter h k (n) and the synthesis filter f k (n) can be defined as follows:
Equation
[0113] The coefficients p0(n) of the prototype filter can be defined with a length L of 640, as shown in Table 4 below. Table 4 [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5] [Table 4-6] The prototype filter, p0(n), can also be derived from Table 4 by one or more mathematical operations such as rounding, subsampling, interpolation, and decimation.
[0114] The adjustment of control information related to SBR typically (as described above) does not depend on the details of transposition, but in some embodiments, certain elements of the control data may be multicast in an eSBR extension container (bs_extension_id == EXTENSION_ID_ESBR) to improve the quality of the regenerated signal. Some of the elements to be multicast may include noise floor data (e.g., a parameter indicating a noise floor scale factor and a direction in either the frequency or time domain for delta coding for each noise floor), inverse filtering data (e.g., a parameter indicating an inverse filtering mode selected from no inverse filtering, low-level inverse filtering, intermediate-level inverse filtering, and strong-level inverse inverse filtering), and missing harmonics data (e.g., a parameter indicating whether a sine wave should be added to a specific frequency band of the regenerated high band). All of these elements depend on the synthesis emulation of the decoder's transposer executed within the encoder, and thus, if appropriately adjusted for the selected transposer, it is possible to increase the quality of the regenerated signal.
[0115] Specifically, in some embodiments, the missing harmonic and inverse filtering control data is transmitted in an eSBR extension container (along with the other bitstream parameters of Table 3) and adjusted for the eSBR harmonic transposer. For the eSBR harmonic transposer, the additional bitrate required to transmit these two classes of metadata is relatively small. Thus, transmitting the adjusted missing harmonic and / or inverse filtering control data in the eSBR extension container will increase the audio quality generated by the transposer with minimal impact on the bitrate. To ensure backward compatibility with legacy decoders, the parameters adjusted for the spectral conversion process of SBR may be transmitted in the bitstream as part of the SBR control data using either implicit or explicit signaling.
[0116] It should be understood that within the scope of the appended claims, the invention may be practiced in ways other than as specifically described herein. Any reference numbers that may be included in the following claims are for illustrative purposes only and should not be used to construe or limit the claims in any way. Various aspects of the present disclosure will be understood from the exemplary forms (EEEs) listed below:
[0117] EEE1. A method for performing high-frequency reconstruction of an audio signal, comprising: Receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata; and Decoding the audio data to generate a decoded low-band audio signal; Extracting the high-frequency reconstruction metadata from the encoded audio bitstream, wherein the high-frequency reconstruction metadata includes operation parameters of a high-frequency reconstruction process, the operation parameters include patch processing mode parameters located within an extension container of the encoded audio bitstream, the patch processing mode parameters of a first value indicate spectral conversion, and the patch processing mode parameters of a second value indicate harmonic transposition by phase vocoder frequency spreading; Filtering the decoded low-band audio signal to generate a filtered low-band audio signal; Regenerating a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, wherein the regeneration includes spectral conversion when the patch processing mode parameter is the first value, and the regeneration includes harmonic transposition by phase vocoder frequency spreading when the patch processing mode parameter is the second value; and A method comprising synthesizing the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal. EEE2. The method according to EEE1, wherein the extension container includes inverse filtering control data to be used when the patch processing mode parameter is equal to the second value. EEE3. The method according to any one of EEE1 to 2, wherein the extension container further includes missing harmonic control data to be used when the patch processing mode parameter is equal to the second value. EEE4. The encoded audio bitstream includes a fill element (having an identifier indicating the start of the fill element) and fill data following the identifier, and the fill data includes the extension container, according to any of the methods described in the preceding EEEs. EEE5. The identifier is a 3-bit unsigned integer transmitted most significant bit first and has a value of 0x6, according to the method described in EEE4. EEE6. The fill data includes an extended payload, the extended payload includes spectral band replication extension data, the extended payload is identified by a 4-bit unsigned integer transmitted most significant bit first and has a value of '1101' or '1110', and optionally, the spectral band replication extension data includes: An optional spectral band replication header, The spectral band replication data after the header, and The spectral band replication extension element after the spectral band replication data including, and the spectral band replication extension element includes a flag, according to the method described in EEE4 or EEE5. EEE7. The high-frequency reconstruction metadata includes an envelope scale factor, a noise floor scale factor, time / frequency grid information, or a parameter indicating a crossover frequency, according to any one of the items in EEE1 to EEE6. EEE8. The filtering is performed according to the following equation by an analysis filter bank including an analysis filter h k (n) which is a modulated version of the prototype filter p0(n):
Equation
Prior Art Documents
Patent Documents
[0118]
Patent Document 1
Claims
【Claim 1】 A method for performing high-frequency reconstruction of an audio signal, comprising: receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata; decoding the audio data to generate a decoded low-band audio signal; extracting the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata including a first set of operation parameters located within a backward-compatible extension container of the encoded audio bitstream, the first set of operation parameters including a patch processing mode parameter, a first value of the patch processing mode parameter indicating spectral conversion, and a second value of the patch processing mode parameter indicating harmonic transposition by phase vocoder frequency spreading; filtering the decoded low-band audio signal to generate a filtered low-band audio signal; regenerating a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, the regeneration including spectral conversion when the patch processing mode parameter is the first value, and the regeneration including harmonic transposition by phase vocoder frequency spreading when the patch processing mode parameter is the second value; combining the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal; wherein when the regeneration includes harmonic transposition by phase vocoder frequency spreading, the operation parameters include a second set of operation parameters not located within a backward-compatible extension container of the encoded audio bitstream. A method in which specific elements of control data are multicast in the backward-compatible extended container, and the specific elements of the control data include at least one of inverse filtering control data, noise floor control data, and missing harmonic control data for use in regenerating a signal. **Claim 2** A computer program for causing a processor to execute the method according to claim 1. **Claim 3** An audio processing unit for performing high-frequency reconstruction of an audio signal, comprising: An input interface for receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata; A core audio decoder for decoding the audio data to generate a decoded low-band audio signal; A de-formatter for extracting the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata including a first set of operation parameters located within a backward-compatible extended container of the encoded audio bitstream, the first set of operation parameters including patch processing mode parameters, a first value of the patch processing mode parameter indicating spectral conversion, and a second value of the patch processing mode parameter indicating harmonic transposition by phase vocoder frequency spreading; An analysis filter bank for filtering the decoded low-band audio signal to generate a filtered low-band audio signal; A high-frequency regeneration unit for regenerating a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, the regeneration including spectral conversion when the patch processing mode parameter is the first value, and the regeneration including harmonic transposition by phase vocoder frequency spreading when the patch processing mode parameter is the second value; A synthesis filter bank that synthesizes the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal; including, when the regeneration includes harmonic transposition by phase vocoder frequency spreading, the operating parameters including a second set of operating parameters that are not located within a backward-compatible extension container of the encoded audio bitstream, an audio processing unit in which specific elements of control data are simulcast in the backward-compatible extension container, the specific elements of the control data including at least one of inverse filtering control data, noise floor control data, and missing harmonic control data for use when regenerating a signal.
Citation Information
Patent Citations
Aliasing Reduction Using a Complex Exponential Modulation Filterbank
JP2004533155A
Multi-channel audio decoding device
JP2007178684A
How to update an encoder using filter interpolation
JP2011529578A
Apparatus and method for improved amplitude response and temporal alignment in a bandwidth expansion method based on a phase vocoder for audio signals.
JP2013521536A
Apparatus and method for processing audio signals using patch boundary matching
JP2013521538A