Integration of high frequency reconstruction techniques with reduced post-processing delay

By decoding and regenerating high-frequency components based on metadata, the method addresses the limitations of existing audio encoding techniques, enhancing audio quality and supporting backward compatibility.

JP2025111688AActive Publication Date: 2025-07-30DOLBY INTERNATIONAL AB
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025073977
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-04-25
Filing Date
2025-04-28
Publication Date
2025-07-30
Estimated Expiration
2039-04-25

AI Technical Summary

Technical Problem

Existing audio encoding techniques, such as MPEG-4 AAC, face challenges with spectral band replication (SBR) that are not ideal for certain audio types, particularly music content with low crossover frequencies, necessitating improved methods for high-frequency reconstruction.

Method used

The method involves decoding an encoded audio bitstream, extracting high-frequency reconstruction metadata, filtering the low-band audio signal, and regenerating the high-band portion based on the metadata, with options for spectral conversion or harmonic transposition, and combining the signals to form a wide-band audio signal.

Benefits of technology

This approach enhances the quality of audio reproduction by accurately reconstructing high-frequency components, improving the audio quality at low data rates, and supports backward compatibility with legacy decoders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111688000001_ABST
    Figure 2025111688000001_ABST
Patent Text Reader

Abstract

To disclose a method for decoding an encoded audio bitstream.SOLUTION: A method includes receiving an encoded audio bitstream and decoding audio data to generate a decoded lowband audio signal. The method further includes extracting high frequency reconstruction metadata and filtering the decoded lowband audio signal with an analytical filterbank to generate a filtered lowband audio signal. The method also includes extracting a flag indicating whether either spectral translation or harmonic transposition is to be performed on the audio data, and regenerating a highband portion of the audio signal using the filtered lowband audio signal and the high frequency reconstruction metadata in accordance with the flag. The high frequency regeneration is performed as a post-processing operation with a delay of 3010 samples per audio channel.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 662,296, filed Apr. 25, 2018, the entire disclosure of which is hereby incorporated by reference.

[0002] Embodiments relate to audio signal processing, and more specifically, to the encoding, decoding, or transcoding of an audio bitstream having control data that specifies whether basic-form high frequency reconstruction (“HFR”) or enhanced-form HFR should be performed on audio data.

Background Art

[0003] A typical audio bitstream includes both audio data (e.g., encoded audio data) representing one or more channels of audio content and metadata representing at least one characteristic of the audio data or audio content. One well-known format for generating an encoded audio bitstream is the MPEG-4 Advanced Audio Coding (AAC) format described in the MPEG standard ISO / IEC 14496-3:2009. In the MPEG-4 standard, AAC represents “Advanced Audio Coding” and HE-AAC represents “High-Efficiency Advanced Audio Coding”.

[0004] The MPEG-4 AAC standard defines several audio profiles that determine which objects and encoding tools are present in a compliant encoder or decoder. Three of those audio profiles are (1) the AAC profile, (2) the HE-AAC profile, and (3) the HE-AAC v2 profile. The AAC profile includes the AAC low complexity (i.e., “AAC-LC”) object type. The AAC-LC object corresponds to the MPEG-2 AAC low complexity profile with some modifications and does not include either the spectral band replication (“SBR”) object type or the parametric stereo (“PS”) object type. The HE-AAC profile is a superset of the AAC profile and further includes the SBR object type. The HE-AAC v2 profile is a superset of the HE-AAC profile and further includes the PS object type.

[0005] The SBR object type includes a spectral band replication tool, which is an important high-frequency reconstruction ("HFR") coding tool that significantly improves the compression efficiency of perceptual audio coders. SBR reconstructs the high-frequency components of an audio signal on the receiver side (e.g., within a decoder). Thus, the encoder only needs to encode and transmit the low-frequency components, enabling much higher audio quality at a low data rate. SBR is based on the replication of a sequence of harmonics that were previously discarded to reduce the data rate from the control data obtained from the encoder and the signal of the limited available bandwidth. The ratio between tonal components and noise-like components is maintained by adaptive inverse filtering and optionally the addition of noise and sine waves. In the MPEG-4 AAC standard, the SBR tool performs spectral patching (also called linear transformation or spectral transformation), in which a number of consecutive quadrature mirror filter (QMF) subbands are replicated (or "patched") from the transmitted low-band portion of the audio signal to the high-band portion of the audio signal, which is generated within the decoder.

[0006] Spectral patching or linear transformation may not be ideal for certain audio types, such as music content with a relatively low crossover frequency. Thus, techniques for improving spectral band replication are desired. SUMMARY OF THE INVENTION

[0007] The first class of embodiments relates to a method for decoding an encoded audio bitstream. The method includes receiving the encoded audio bitstream, decoding the audio data to generate a decoded low-band audio signal. The method further includes extracting high-frequency reconstruction metadata, filtering the decoded low-band audio signal with an analysis filter bank to generate a filtered low-band audio signal. The method further includes extracting a flag indicating whether spectral conversion or harmonic transposition should be performed on the audio data, and regenerating a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata according to the flag. Finally, the method includes combining the filtered low-band audio signal and the regenerated high-band portion to form a wide-band audio signal.

[0008] The second class of embodiments relates to an audio decoder that decodes an encoded audio bitstream. The decoder includes an input interface that receives the encoded audio bitstream, where the encoded audio bitstream includes audio data representing a low-band portion of an audio signal, and a core decoder that decodes the audio data to generate a decoded low-band audio signal. The decoder also includes a demultiplexer that extracts high-frequency reconstruction metadata from the encoded audio bitstream, where the high-frequency reconstruction metadata includes operation parameters for a high-frequency reconstruction process that linearly transforms a number of consecutive subbands from a low-band portion of the audio signal to a high-band portion of the audio signal, and an analysis filter bank that filters the decoded low-band audio signal to generate a filtered low-band audio signal. The decoder further includes a demultiplexer that extracts a flag from the encoded audio bitstream indicating whether linear transformation or harmonic transposition is to be performed on the audio data, and a high-frequency regenerator that regenerates a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata according to the flag. Finally, the decoder includes a synthesis filter bank that combines the filtered low-band audio signal and the regenerated high-band portion to form a wideband audio signal.

[0009] Other classes of embodiments relate to encoding and transcoding an audio bitstream that includes metadata specifying whether an enhanced spectral band replication (eSBR) process is to be performed. BRIEF DESCRIPTION OF THE DRAWINGS

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Embodiments for Carrying Out the Invention

[0011] Notation and Terminology System Throughout this disclosure, including in the claims, the expression "performing a process on" a signal or data (e.g., filtering, scaling, converting, or applying a gain to the signal or data) is used in a broad sense to mean performing the process directly on the signal or data or on a processed version of the signal or data (e.g., on a version of the signal that has received preliminary filtering or preprocessing before the process is performed).

[0012] Throughout this disclosure, including in the claims, the expressions "audio processing unit" or "audio processor" are used in a broad sense to denote a system, device, or apparatus configured to process audio data. Examples of audio processing units include, but are not limited to, encoders, transcoders, decoders, codecs, preprocessing systems, postprocessing systems, and bitstream processing systems (which may also be referred to as bitstream processing tools). Nearly all household appliances, such as mobile phones, televisions, laptops, and tablet computers, include an audio processing unit or an audio processor.

[0013] Throughout this disclosure, including in the claims, the terms "coupled" or "coupled to" are used in a broad sense to mean any direct or indirect connection. Thus, when a first device is coupled to a second device, that connection may be through a direct connection or through an indirect connection via other devices and connections. Also, components integrated into or integrated with other components are coupled to each other.

[0014] Detailed Description of Embodiments of the Invention The MPEG-4 AAC standard contemplates that an encoded MPEG-4 AAC bitstream includes the following metadata, namely, metadata indicating each type of high-frequency reconstruction (“HFR”) processing to be applied by a decoder (if it should be applied) to decode the audio content of the bitstream, and / or controlling such HFR processing, and / or indicating at least one characteristic or parameter of at least one HFR tool to be used to decode the audio content of the bitstream. Herein, the expression “SBR metadata” is used to represent this type of metadata described or referred to in the MPEG-4 AAC standard with respect to use in spectral band replication (“SBR”). As will be understood by those skilled in the art, SBR is a form of HFR.

[0015] SBR is preferably used as a dual-rate system, where the underlying codec operates at half the original sampling rate while SBR operates at the original sampling rate. The SBR encoder operates in parallel with the underlying core codec, although at a higher sampling rate. SBR is mainly a post-processing in the decoder, but important parameters are extracted in the encoder to ensure the most accurate high-frequency reconstruction in the decoder. The encoder estimates the spectral envelope of the SBR range with respect to a time and frequency range / resolution suitable for the current input signal segment characteristics. The spectral envelope is estimated by complex QMF analysis and subsequent energy calculation. The time and frequency resolution of the spectral envelope can be selected with a high degree of freedom to ensure the time-frequency resolution most suitable for a given input segment. The envelope estimation needs to take into account that the original transient components (e.g., hi-hat), mainly located in the high-frequency region, are slightly present in the SBR-generated high band before envelope adjustment. This is because the high band in the decoder is based on the low band where the transient components are much less prominent compared to the high band. This aspect poses different requirements regarding the time-frequency resolution of the spectral envelope data compared to the normal spectral envelope estimation used in other audio coding algorithms.

[0016] Leaving aside the spectral envelope, several additional parameters are extracted that represent the spectral characteristics of the input signal in different time and frequency domains. The encoder naturally has access to information on how the SBR unit in the decoder creates the high band given a particular set of control parameters as well as the original signal, so that the system can handle the following situations, namely, situations where the low band consists of a strong harmonic series and the recreated high band consists mainly of random signal components, and situations where there are strong tonal components in the original high band that have no counterparts in the underlying low band on which the high band region is based. Further, the SBR encoder operates in close conjunction with the underlying core codec to determine which frequency range should be covered by the SBR at a given point in time. The SBR data is efficiently coded before transmission by utilizing entropy coding and, in the case of stereo signals, the channel dependence of the control data.

[0017] The control parameter extraction algorithm typically needs to be carefully adjusted to match the underlying codec at a given bitrate and a given sampling rate. This is due to the fact that a lower bitrate usually means a larger SBR range compared to a higher bitrate and different sampling rates correspond to different time resolutions of the SBR frame.

[0018] The SBR decoder typically includes several different parts. It includes a bitstream decoding module, a high-frequency reconstruction (HFR) module, an additional high-frequency component module, and an envelope adjustment module. The system is based on a complex-valued QMF filter bank (for high-quality SBR) or a real-valued QMF filter bank (for low-power SBR). Embodiments of the invention are applicable to both high-quality SBR and low-power SBR. In the bitstream extraction module, control data is read from and decoded from the bitstream. Before reading the envelope data from the bitstream, a time-frequency grid is obtained for the current frame. The underlying core decoder decodes the audio signal of the current frame (at a lower sampling rate) to generate time-domain audio samples. The resulting frame of audio data is used for high-frequency reconstruction by the HFR module. Then, the decoded low-band signal is analyzed using a QMF filter bank. Subsequently, high-frequency reconstruction and envelope adjustment are performed on the subband samples of the QMF filter bank. The high frequencies are reconstructed from the low band in a flexible manner based on given control parameters. Furthermore, the reconstructed high band is adaptively filtered on a subband channel basis according to the control data to ensure appropriate spectral characteristics in the given time / frequency domain.

[0019] The top level of the MPEG-4 AAC bitstream is a sequence of data blocks ("raw_data_block" elements), each of which is a segment of data (hereafter referred to as a "block") containing audio data (typically over a period of 1024 or 960 samples) as well as related information and / or other data. Here, the term "block" is used to represent a segment of the MPEG-4 AAC bitstream having one (and no more than one) "raw_data_block" element's worth of audio data (along with the corresponding metadata, and optionally other related data).

[0020] Each block of an MPEG-4 AAC bitstream can contain a number of syntax elements (each of which is also embodied within the bitstream as a segment of data). Seven types of such syntax elements are defined in the MPEG-4 AAC standard. Each syntax element is identified by a different value of the data element "id_syn_ele". Examples of syntax elements include "single_channel_element()", "channel_pair_element()", "fill_element()". A single channel element is a container that contains audio data for a single audio channel (mono audio signal). A channel pair element contains audio data for two audio channels (i.e., a stereo audio signal).

[0021] A fill element is a container of information that contains an identifier (e.g., the value of the above element "id_syn_ele") followed by data (referred to as "fill data"). Fill elements have historically been used to adjust the instantaneous bitrate of a bitstream that is to be transmitted over a constant rate channel. By adding an appropriate amount of fill data to each block, a constant data rate can be achieved.

[0022] According to an embodiment of the invention, the fill data can include one or more extension payloads that extend the types of data (e.g., metadata) that can be transmitted in the bitstream. A decoder that receives a bitstream having fill data that includes a new type of data can optionally be used by the device (e.g., the decoder) that receives this bitstream to extend the functionality of the device. Thus, as can be understood by those skilled in the art, a fill element is a special type of data structure that is different from the data structures typically used to transmit audio data (e.g., an audio payload including channel data).

[0023] In some embodiments of the invention, the identifier used to identify the padding element may be composed of a 3-bit unsigned integer with the value of 0x6 and the most significant bit (“uimsbf”) transmitted first. Within one block, several instances of the same type of syntax element (e.g., several padding elements) may occur.

[0024] Another standard for encoding an audio bit stream is the MPEG USAC (Unified Speech and Audio Coding) standard (ISO / IEC 23003-3:2012). The MPEG USAC standard describes the encoding and decoding of audio content using spectral band replication processing (including the SBR processing described in the MPEG-4 AAC standard and other enhanced forms of spectral band replication processing). This processing applies an extended and enhanced version of the spectral band replication tool of the SBR toolset described in the MPEG-4 AAC standard (hereafter sometimes referred to as the “enhanced SBR tool” or “eSBR tool”). Thus, eSBR (defined in the USAC standard) is an improvement over SBR (defined in the MPEG-4 AAC standard).

[0025] Here, the expression “enhanced SBR processing” (or “eSBR processing”) is used to represent spectral band replication processing using at least one eSBR tool not described or mentioned in the MPEG-4 AAC standard (e.g., at least one eSBR tool described or mentioned in the MPEG USAC standard). Examples of such eSBR tools are harmonic transposition and additional preprocessing or “pre-flattening” by QMF patching.

[0026] A harmonic transposer of integer order T maps a sine wave of frequency ω to a sine wave of frequency Tω while maintaining the signal duration. Typically, three orders, T = 2, 3, 4, are used in sequence to generate each part of the desired output frequency range using the smallest possible transposition order. If an output of a transposition range above order 4 is required, it can be generated by a frequency shift. When possible, a substantially critically sampled baseband time domain is created for processing to minimize computational complexity.

[0027] The harmonic transposer may be either QMF-based or DFT-based. When using a QMF-based harmonic transposer, the bandwidth expansion of the core coder time domain signal is fully performed within the QMF domain using an improved phase vocoder structure, performing decimation and subsequent time stretching for all QMF subbands. Transpositions using several transposition factors (e.g., T = 2, 3, 4) are performed in a common QMF analysis / synthesis transform stage. Since the QMF-based harmonic transposer does not feature signal-adaptive frequency domain oversampling, the corresponding flag (sbrOversamplingFlag[ch]) in the bitstream can be ignored.

[0028] When using a DFT-based harmonic transposer, preferably, transposers of factors 3 and 4 (third and fourth order transposers) are integrated into a factor 2 transposer (second order transposer) by interpolation to reduce complexity. For each frame (corresponding to coreCoderFrameLength core coder samples), first, a transposer of a nominal "full size" transform size is determined by the signal-adaptive frequency domain oversampling flag (sbrOverSamplingFlag[ch]) in the bitstream.

[0029] When sbrPatchingMode == 1, it indicates that linear transposition should be used to generate the high band, and additional steps can be introduced to avoid discontinuities in the shape of the spectral envelope of the high-frequency signal input to the subsequent envelope adjuster. This improves the processing of the subsequent envelope adjustment stage and results in a high-band signal that feels more stable. The operation of this additional preprocessing is beneficial for signal types where the rough spectral envelope of the low-band signal used for high-frequency reconstruction exhibits large level fluctuations. However, the value of the bitstream element can be determined by the encoder by applying some kind of signal-dependent classification. This additional preprocessing is preferably activated by bs_sbr_preprocessing, which is a 1-bit bitstream element. When bs_sbr_preprocessing is set to 1, this additional processing is enabled. When bs_sbr_preprocessing is set to zero, this additional preprocessing is disabled. This additional processing preferably utilizes the preGain curve used by the high-frequency generator to scale the low-band X Low For example, the preGain curve can be calculated according to

Number

Number

Number

[0030] A bitstream generated according to the MPEG USAC standard (sometimes referred to herein as the "USAC" bitstream) includes the encoded audio content and typically metadata indicating each type of spectral band replication processing applied by a decoder to decode the audio content of the USAC bitstream, and / or metadata indicating at least one characteristic or parameter of at least one SBR tool and / or at least one eSBR tool that controls such spectral band replication processing and / or is used to decode the audio content of the USAC bitstream.

[0031] Here, the expression "enhanced SBR metadata" (or "eSBR metadata") is used to represent metadata that describes each type of spectral band replication process applied by a decoder to decode the audio content of an encoded audio bitstream (e.g., a USAC bitstream), and / or controls such spectral band replication process, and / or indicates at least one characteristic or parameter of at least one SBR tool and / or at least one eSBR tool used to decode such audio content, but is not described or referred to in the MPEG-4 AAC standard. An example of eSBR metadata is metadata (indicating or controlling spectral band replication process) that is not described or referred to in the MPEG-4 AAC standard but is described or referred to in the MPEG USAC standard. Thus, eSBR metadata represents, here, metadata that is not SBR metadata, and SBR metadata represents, here, metadata that is not eSBR metadata.

[0032] A USAC bitstream may include both SBR metadata and eSBR metadata. More specifically, a USAC bitstream may include eSBR metadata that controls the execution of eSBR processing by a decoder and SBR metadata that controls the execution of SBR processing by a decoder. According to an exemplary embodiment of the present invention, eSBR metadata (e.g., configuration data specific to eSBR) is included in an MPEG-4 AAC bitstream (e.g., within the sbr_extension() container at the end of the SBR payload) (in accordance with the present invention).

[0033] The execution of eSBR processing during the decoding of an encoded bitstream using an eSBR toolset (having at least one eSBR tool) by a decoder reproduces the high-frequency band of the audio signal based on the replication of a sequence of harmonics discarded during encoding. Such eSBR processing typically adjusts the spectral envelope of the generated high-frequency band, applies inverse filtering, and adds noise components and sine wave components in order to reproduce the spectral characteristics of the original audio signal.

[0034] According to an exemplary embodiment of the invention, the eSBR metadata is included in one or more of a plurality of metadata segments of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that also includes audio data encoded within other segments (audio data segments). Typically, at least one such metadata segment of each block of the bitstream is a padding element (including an identifier indicating the start of the padding element) (or includes), and the eSBR metadata is included in the padding element following the identifier.

[0035] FIG. 1 is a block diagram of an exemplary audio processing chain (audio data processing system) in which one or more of the elements of the system may be configured in accordance with an embodiment of the present invention. This system includes the following elements coupled together as shown, namely, an encoder 1, a delivery subsystem 2, a decoder 3, and a post-processing unit 4. In variations of the illustrated system, one or more of these elements may be omitted, or additional audio data processing units may be included.

[0036] In some implementations, encoder 1 (which optionally includes a preprocessing unit) is configured to receive as input PCM (time domain) samples having audio content and output an encoded audio bitstream (having a format compliant with the MPEG-4 AAC standard) representing the audio content. Of the bitstream, the data representing the audio content may herein be referred to as "audio data" or "encoded audio data". When the encoder is configured according to a typical embodiment of the present invention, the audio bitstream output from the encoder includes the audio data as well as eSBR metadata (and typically other metadata too).

[0037] One or more encoded audio bitstreams output from encoder 1 may be asserted to an encoded audio delivery subsystem 2. Subsystem 2 is configured to store and / or deliver each of the encoded bitstreams output from encoder 1. The encoded audio bitstreams output from encoder 1 may be stored by subsystem 2 (e.g., in the form of a DVD or Bluray (registered trademark) disk), or may be transmitted by subsystem 2, or may be stored and transmitted by subsystem 2.

[0038] Decoder 3 is configured to decode an encoded MPEG-4 AAC audio bitstream received via subsystem 2 (generated by encoder 1). In some embodiments, decoder 3 extracts eSBR metadata from each block of the bitstream and decodes the bitstream (including by performing eSBR processing using the extracted eSBR metadata) to generate decoded audio data (e.g., a stream of decoded PCM audio samples). In some embodiments, decoder 3 extracts SBR metadata from the bitstream (but ignores eSBR metadata included in the bitstream) and decodes the bitstream (including by performing SBR processing using the extracted SBR metadata) to generate decoded audio data (e.g., a stream of decoded PCM audio samples). Typically, decoder 3 includes a buffer for storing (e.g., non-transitorily) segments of the encoded audio bitstream received from subsystem 2.

[0039] The post-processing unit 4 of FIG. 1 is configured to receive a stream of decoded audio data (e.g., decoded PCM audio samples) from decoder 3 and perform post-processing thereon. The post-processing unit may also be configured to render the post-processed audio content (or the decoded audio received from decoder 3) for playback by one or more speakers.

[0040] Figure 2 is a block diagram of an encoder (100) which is an embodiment of the inventive audio processing unit. Any of the components or elements of encoder 100 can be implemented in hardware, in software, or in a combination of hardware and software, as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit). Encoder 100 includes an encoder 105, a buffer / formatter stage 107, a metadata generation stage 106, and a buffer memory 109, connected as shown. Typically, encoder 100 also includes other processing elements (not shown). Encoder 100 is configured to convert an input audio bitstream into an encoded output MPEG-4 AAC bitstream.

[0041] Metadata generator 106 is coupled and configured to generate (and / or pass to stage 107) metadata (including eSBR metadata and SBR metadata) to be included by stage 107 in the encoded bitstream output from encoder 100.

[0042] Encoder 105 is coupled and configured to encode the input audio data (e.g., by performing compression thereon) and assert to stage 107 the resulting encoded audio for inclusion in the encoded bitstream output from stage 107.

[0043] Stage 107 is configured to multiplex the encoded audio from encoder 105 and the metadata (including eSBR metadata and SBR metadata) from generator 106 to produce an encoded bitstream output from stage 107, preferably having a format specified by one of the embodiments of the present invention.

[0044] Buffer memory 109 is configured to store (e.g., non - transiently) at least one block of the encoded audio bitstream output from stage 107, and a series of blocks of the encoded audio bitstream are asserted from buffer memory 109 as an output from encoder 100 to the delivery system.

[0045] FIG. 3 is a block diagram of a system including a decoder (200) which is one embodiment of the inventive audio processing unit and optionally including a post - processor (300) coupled thereto. Any of the components or elements of decoder 200 and post - processor 300 may be implemented in hardware, in software, or in a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit). Decoder 200 has a buffer memory 201, a bitstream payload de - formatter (parser) 205, an audio decoding subsystem 202 (also referred to as the “core” decoding stage or “core” decoding subsystem), an eSBR processing stage 203, and a control bit generation stage 204, connected as shown. Typically, decoder 200 also includes other processing elements (not shown).

[0046] Buffer memory (buffer) 201 stores (e.g., non - transiently) at least one block of the encoded MPEG - 4 AAC audio bitstream received by decoder 200. In the operation of decoder 200, a series of blocks of the bitstream are asserted from buffer 201 to de - formatter 205.

[0047] In a variation of the embodiment of FIG. 3 (or the embodiment of FIG. 4 described later), an APU that is not a decoder (e.g., the APU 500 of FIG. 6) includes a buffer memory (e.g., the same buffer memory as buffer 201) that (e.g., non-temporarily) stores at least one block of the same type of encoded audio bitstream received by buffer 201 of FIG. 3 or FIG. 4 (i.e., an encoded audio bitstream including eSBR metadata) (e.g., an MPEG-4 AAC audio bitstream).

[0048] Referring again to FIG. 3, the de-formatter 205 is coupled and configured to demultiplex each block of the bitstream, extract therefrom SBR metadata (including quantized envelope data) and eSBR metadata (and typically also other metadata), assert at least the eSBR metadata and the SBR metadata to the eSBR processing stage 203, and also typically assert the other extracted metadata to the decoding subsystem 202 (and optionally also to the control bit generator 204). The de-formatter 205 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.

[0049] The system of FIG. 3 also optionally includes a post-processor 300. The post-processor 300 includes a buffer memory (buffer) 301 and other processing elements (not shown) including at least one processing element coupled to the buffer 301. The buffer 301 stores at least one block (or frame) of the decoded audio data received by the post-processor 300 from the decoder 200. The processing elements of the post-processor 300 are coupled and configured to receive a series of blocks (or frames) of the decoded audio output from the buffer 301 and adaptively process it using metadata output from the decoding subsystem 202 (and / or the de-formatter 205) and / or control bits output from stage 204 of the decoder 200.

[0050] The audio decoding subsystem 202 of decoder 200 decodes the audio data extracted by parser 205 (such decoding may be referred to as "core" decoding process), generates the decoded audio data, and is configured to assert the decoded audio data to eSBR processing stage 203. This decoding is performed in the frequency domain and typically includes inverse quantization followed by spectral processing. Typically, the final stage of processing in subsystem 202 applies a frequency domain - time domain conversion to the decoded frequency domain audio data so that the output of subsystem 202 is the decoded audio data in the time domain. Stage 203 applies the SBR tools and eSBR tools indicated by the SBR metadata and eSBR metadata (extracted by parser 205) to the decoded audio data (that is, uses the SBR and eSBR metadata to perform SBR and eSBR processing on the output of decoding subsystem 202) to generate the fully decoded audio data output from decoder 200 (for example, to post - processor 300). Typically, decoder 200 includes a memory (accessible by subsystem 202 and stage 203) that stores the formatted audio data and metadata output from deframer 205, and stage 203 is configured to access the audio data and metadata (including SBR metadata and eSBR metadata) as needed during SBR and eSBR processing. The SBR processing and eSBR processing in stage 203 can be regarded as post - processing on the output of core decoding subsystem 202. Optionally, decoder 200 is also combined and configured to perform upmixing on the output of stage 203 to generate the fully decoded and upmixed audio output from decoder 200. This final upmixing subsystem (which may apply parametric stereo ("PS") tools defined in the MPEG - 4 AAC standard using the PS metadata extracted by deframer 205 and / or the control bits generated by subsystem 204) is included.Alternatively, the post-processor 300 is configured to perform upmixing on the output of the decoder 200 (e.g., using the PS metadata extracted by the de-formatter 205 and / or the control bits generated by the subsystem 204).

[0051] In response to the metadata extracted by the de-formatter 205, the control bit generator 204 can generate control data that can be used within the decoder 200 (e.g., in the final upmixing subsystem) and / or asserted as the output of the decoder 200 (e.g., to the post-processor 300 for post-processing use). In response to the metadata extracted from the input bitstream (and optionally in response to the control data as well), stage 204 can generate (and assert to the post-processor 300) control bits indicating that the decoded audio data output from the eSBR processing stage 203 should undergo a particular type of post-processing. In some implementations, the decoder 200 is configured to assert the metadata extracted by the de-formatter 205 from the input bitstream to the post-processor 300, and the post-processor 300 is configured to use the metadata to perform post-processing on the decoded audio data output from the decoder 200.

[0052] Figure 4 is a block diagram of an audio processing unit (“APU”) (210), which is another embodiment of the inventive audio processing unit. The APU 210 is a legacy decoder that is not configured to perform eSBR processing. Any of the components or elements of the APU 210 may be implemented in hardware, in software, or in a combination of hardware and software as one or more processes and / or one or more circuits (e.g., ASIC, FPGA, or other integrated circuit). The APU 210 has a buffer memory 201, a bitstream payload de-formatter (parser) 215, an audio decoding subsystem 202 (also referred to as the “core” decoding stage or “core” decoding subsystem), and an SBR processing stage 213, connected as shown. Typically, the APU 210 also includes other processing elements (not shown). The APU 210 may represent, for example, an audio encoder, decoder, or transcoder.

[0053] Elements 201 and 202 of the APU 210 are the same as the similarly numbered elements of the decoder 200 (of FIG. 3) and will not be repeated in their description above. In operation of the APU 210, a series of blocks of the encoded audio bitstream (MPEG-4 AAC bitstream) received by the APU 210 are asserted from the buffer 201 to the de-formatter 205.

[0054] The de-formatter 215 demultiplexes each block of the bitstream and then extracts SBR metadata (including quantized envelope data) and typically also other metadata, but is coupled and configured to ignore eSBR metadata that may be included in the bitstream according to any embodiment of the present invention. The de-formatter 215 is configured to assert at least the SBR metadata to the SBR processing stage 213. The de-formatter 215 is also coupled and configured to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.

[0055] The audio decoding subsystem 202 of the APU 210 decodes the audio data extracted by the de-formatter 215 (such decoding may be referred to as "core" decoding process), generates the decoded audio data, and is configured to assert the decoded audio data to the SBR processing stage 213. This decoding is performed in the frequency domain. Typically, as the output of the subsystem 202 is the decoded audio data in the time domain, the final stage of the processing in the subsystem 202 applies a frequency-domain to time-domain conversion to the decoded audio data in the frequency domain. The stage 213 applies the SBR tools indicated by the SBR metadata (extracted by the de-formatter 215) to the decoded audio data (without applying the eSBR tool) (i.e., uses the SBR metadata to perform SBR processing on the output of the decoding subsystem 202) to generate the fully decoded audio data output from the APU 210 (e.g., to the post-processor 300). Typically, the APU 210 includes a memory (accessible by the subsystem 202 and the stage 213) that stores the formatted audio data and metadata output from the de-formatter 215, and the stage 213 is configured to access the audio data and metadata (including the SBR metadata) as needed during the SBR processing. The SBR processing in the stage 213 can be regarded as a post-processing on the output of the core decoding subsystem 202. Optionally, the APU 210 is also combined and configured with a final upmixing subsystem that performs upmixing on the output of the stage 213 to generate the fully decoded and upmixed audio output from the APU 210 (this may apply the parametric stereo ("PS") tools defined in the MPEG-4 AAC standard using the PS metadata extracted by the de-formatter 205). Alternatively, the post-processor is configured to perform upmixing on the output of the APU 210 (e.g., using the PS metadata extracted by the de-formatter 215 and / or the control bits generated by the APU 210).

[0056] Various implementations of the encoder 100, decoder 200, and APU 210 are configured to execute different embodiments of the inventive method.

[0057] According to some embodiments, eSBR metadata is included in an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) (e.g., a few control bits that are eSBR metadata), but a legacy decoder (which is not configured to parse the eSBR metadata or use eSBR tools related to the eSBR metadata) can ignore the eSBR metadata and, nevertheless, be able to decode the bitstream to the extent possible, typically without a significant penalty in decoded audio quality, without the use of the eSBR metadata or eSBR tools related to the eSBR metadata. On the other hand, an eSBR decoder configured to parse the bitstream to identify the eSBR metadata and use at least one eSBR tool in response to the eSBR metadata will enjoy the benefits of using at least one such eSBR tool. Accordingly, embodiments of the invention provide a means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner.

[0058] Typically, the eSBR metadata in the bitstream indicates one or more of the following eSBR tools (which are described in the MPEG USAC standard and may or may not be applied by the encoder during generation of the bitstream): · Harmonic transposition, and · Additional preprocessing (preflattening) by QMF patching.

[0059] For example, the eSBR metadata included in the bitstream may indicate the values of parameters such as sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBins[ch], sbrPitchInBins[ch], and bs_sbr_preprocessing (described in the MPEG USAC standard and this disclosure).

[0060] Here, assuming X is some parameter, the notation X[ch] indicates that the parameter relates to the channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, the expression [ch] may be omitted, and it may be assumed that the corresponding parameter relates to the channel of the audio content.

[0061] Here, assuming X is some parameter, the notation X[ch][env] indicates that the parameter relates to the SBR envelope ("env") of the channel ("ch") of the audio content of the encoded bitstream to be decoded. For simplicity, the expressions [env] and [ch] may be omitted, and it may be assumed that the corresponding parameter relates to the SBR envelope of the channel of the audio content.

[0062] In decoding the encoded bitstream, the execution of harmonic transposition during the decoding eSBR processing stage (for each channel "ch" of the audio content indicated by the bitstream) is controlled by the eSBR metadata parameters sbrPatchingMode[ch], sbrOversamplingFlag[ch], sbrPitchInBinsFlag[ch], and sbrPitchInBins[ch].

[0063] The value “sbrPatchingMode[ch]” indicates the type of transposer used in eSBR. sbrPatchingMode[ch]=1 indicates linear transposition patching (used for either high-quality SBR or low-power SBR) as described in section 4.6.18 of the MPEG-4 AAC standard, and sbrPatchingMode[ch]=0 indicates harmonic SBR patching as described in section 7.5.3 or 7.5.4 of the MPEG USAC standard.

[0064] The value “sbrOversamplingFlag[ch]” indicates the use of signal adaptive frequency domain oversampling in eSBR in combination with DFT-based harmonic SBR patching as described in section 7.5.3 of the MPEG USAC standard. This flag controls the size of the DFT used in the transposer. 1 indicates that signal adaptive frequency domain oversampling is enabled as described in section 7.5.3.1 of the MPEG USAC standard, and 0 indicates that signal adaptive frequency domain oversampling is disabled as described in section 7.5.3.1 of the MPEG USAC standard.

[0065] The value “sbrPitchInBinsFlag[ch]” controls the interpretation of the sbrPitchInBins[ch] parameter. 1 indicates that the value of sbrPitchInBins[ch] is valid and greater than zero, and 0 indicates that the value of sbrPitchInBins[ch] is set to zero.

[0066] The value “sbrPitchInBins[ch]” controls the addition of outer product terms in the SBR harmonic transposer. The value sbrPitchinBins[ch] is an integer value within the range [0,127] and represents the distance measured in frequency bins of a 1536-line DFT that acts on the sampling frequency of the core coder.

[0067] When the MPEG-4 AAC bitstream indicates SBR channel pairs whose channels are not combined (rather than a single SBR channel), the bitstream indicates two instances of the above syntax (relating to harmonic or non-harmonic transposition) for each channel of sbr_channel_pair_element().

[0068] Harmonic transposition of the eSBR tool typically improves the quality of the decoded music signal at relatively low crossover frequencies. Non-harmonic transposition (i.e., legacy spectral patching) typically improves speech signals. Thus, the starting point in the decision on which type of transposition is preferable for encoding a particular audio content is to select the transposition method according to audio / music detection, assuming that harmonic transposition is used for music content and spectral patching is used for speech content.

[0069] The execution of preflattening in eSBR processing is controlled by the value of a 1-bit eSBR metadata parameter known as “bs_sbr_preprocessing” (in the sense that preflattening is either executed or not executed according to this single-bit value). When the SBR QMF patching algorithm described in section 4.6.18.6.3 of the MPEG-4 AAC standard is used, the post-preflattening step can be executed (when indicated by the “bs_sbr_preprocessing” parameter) as part of an effort to avoid discontinuities in the shape of the spectral envelope of the high-frequency signal input to the subsequent envelope regulator (the envelope regulator performs another stage of eSBR processing). Preflattening typically improves the processing of the subsequent envelope adjustment stage and results in a high-band signal that feels more stable.

[0070] The overall bitrate requirement for including eSBR metadata (harmonic transposition and pre-flattening) indicating the above eSBR tool in an MPEG-4 AAC bitstream is expected to be on the order of several hundred bits per second according to some embodiments of the invention, since only the differential control data required to perform the eSBR processing is transmitted. Since this information is included in a backward-compatible manner (as will be described later), legacy decoders can ignore this information. Thus, the adverse impact on bitrate associated with including eSBR metadata can be ignored for several reasons including the following: · Since only the differential control data required to perform the eSBR processing is transmitted (not the simultaneous transmission of SBR control data), the bitrate penalty (due to including eSBR metadata) is a very small part of the overall bitrate, and · The adjustment (tuning) of SBR-related control information typically does not depend on the details of the transposition. An example of the case where the control data depends on the operation of the transposer will be described later in this application.

[0071] Accordingly, embodiments of the invention provide means for efficiently transmitting enhanced spectral band replication (eSBR) control data or metadata in a backward-compatible manner. This efficient transmission of eSBR control data reduces the memory requirements in decoders, encoders, and transcoders that adopt aspects of the invention without having a visible adverse impact on bitrate. Further, the complexity and processing requirements associated with performing eSBR according to embodiments of the invention are also reduced. This is because the SBR data only needs to be processed once (as opposed to being transmitted simultaneously as would be the case if eSBR were treated as a completely separate object type in MPEG-4 AAC instead of being integrated into the MPEG-4 AAC codec in a backward-compatible manner).

[0072] Next, referring to FIG. 7, elements of a block (“raw_data_block”) of an MPEG-4 AAC bitstream that includes eSBR metadata are described according to some embodiments of the present invention. FIG. 7 is a diagram of a block (“raw_data_block”) of an MPEG-4 AAC bitstream showing a portion of that segment.

[0073] A block of an MPEG-4 AAC bitstream can include at least one “single_channel_element()” (e.g., the single-channel element shown in FIG. 7) and / or at least one “channel_pair_element()” (not specifically shown in FIG. 7 but may be present) that includes audio data of an audio program. This block can also include a plurality of “padding elements” (e.g., padding element 1 and / or padding element 2 in FIG. 7) that include data related to that program (e.g., metadata). Each “single_channel_element()” includes an identifier (e.g., “ID1” in FIG. 7) indicating the start of a single-channel element and can include audio data indicating different channels of a multi-channel audio program. Each “channe_pair_element()” includes an identifier (not shown in FIG. 7) indicating the start of a channel pair element and can include audio data indicating two channels of a program.

[0074] The fill_element (referred to here as the fill element) of the MPEG-4 AAC bitstream includes an identifier ("ID2" in Figure 7) indicating the start of the fill element and the fill data following the identifier. The identifier ID2 can be composed of a 3-bit unsigned integer with the most significant bit ("uimsbf") transmitted first and having a value of 0x6. The fill data can include an extension_peyload() element (which may also be referred to here as the extended payload), and its syntax is shown in Table 4.57 of the MPEG-4 AAC standard. There are several types of extended payloads, which are identified via an "extension_type" parameter that is a 4-bit unsigned integer with the most significant bit ("uimsbf") transmitted first.

[0075] The fill data (e.g., its extended payload) can include a header or identifier (e.g., "header 1" in Figure 7) indicating a segment of the fill data representing an SBR object (i.e., the header starts an "SBR object" type that is referred to as sbr_extension_data() in the MPEG-4 AAC standard). For example, with values of '1101' or '1110' in the extension_type field within the header, a spectral band replication (SBR) extended payload is identified, with the identifier '1101' identifying an extended payload with SBR data and '1110' identifying an extended payload with SBR data with a cyclic redundancy check (CRC) to verify the accuracy of the SBR data.

[0076] When a header (e.g., the extension_type field) starts an SBR object type, SBR metadata (referred to as "sbr_data()" in the MPEG-4 AAC standard and may be referred to here as "spectral band replication data") follows the header, and at least one spectral band replication extension element (e.g., the "SBR extension element" of padding element 1 in FIG. 7) can follow the SBR metadata. Such a spectral band replication extension element (a segment of the bitstream) is called the "sbr_extension()" container in the MPEG-4 AAC standard. The spectral band replication extension element optionally includes a header (e.g., the "SBR extension header" of padding element 1 in FIG. 7).

[0077] The MPEG-4 AAC standard contemplates that a spectral band replication extension element can include PS (parametric stereo) data regarding the audio data of a program. The MPEG-4 AAC standard contemplates that when a header of a padding element (e.g., the header of its extended payload) starts an SBR object type and the spectral band replication extension element of the padding element includes PS data (as is the case for "header 1" in FIG. 7), the padding element (e.g., its extended payload) includes spectral band replication data and a "bs_extension_id" parameter having a value indicating that PS data is included in the spectral band replication extension element of the padding element (i.e., bs_extension_id = 2).

[0078] According to some embodiments of the present invention, eSBR metadata (e.g., a flag indicating whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of a block) is included in the spectral band replication extension element of a stuffing element. For example, such a flag is shown in stuffing element 1 of FIG. 7, where the flag occurs after the header of the "SBR extension element" of stuffing element 1 (the "SBR extension header" of stuffing element 1). Optionally, such a flag and additional eSBR metadata are included in the spectral band replication extension element after the header of the spectral band replication extension element (e.g., in FIG. 7, the SBR extension element of stuffing element 1 after the SBR extension header). According to some embodiments of the present invention, a stuffing element containing eSBR metadata also includes a "bs_extension_id" parameter having a value (e.g., bs_extension_id = 3) indicating that the eSBR metadata is included in the stuffing element and that eSBR processing should be performed on the audio content of the corresponding block.

[0079] According to some embodiments of the present invention, the eSBR metadata is included in padding elements other than the spectral band replication extension element (SBR extension element) of the padding elements in the MPEG-4 AAC bitstream (e.g., padding element 2 in FIG. 7). This is because padding elements containing extension_payload() with SBR data or SBR data with CRC do not contain any other extension payloads of other extension types. Therefore, in embodiments where the eSBR metadata stores its own extension payload, a separate padding element is used to store the eSBR metadata. Such a padding element includes an identifier indicating the start of the padding element (e.g., "ID2" in FIG. 7) and padding data following the identifier. The padding data can include an extension_payload() element (sometimes referred to as an extension payload herein), and its syntax is shown in Table 4.57 of the MPEG-4 AAC standard. The padding data (e.g., its extension payload) includes a header indicating the eSBR object (e.g., "header 2" of padding element 2 in FIG. 7) (i.e., this header starts the enhanced spectral band replication (eSBR) object type), and the padding data (e.g., its extension payload) includes the eSBR metadata after the header. For example, padding element 2 in FIG. 7 includes such a header ("header 2") and, after the header, includes eSBR metadata (i.e., a "flag" within padding element 2 indicating whether enhanced spectral band replication (eSBR) processing should be performed on the audio content of the block). Optionally, additional eSBR metadata is also included in the padding data of padding element 2 in FIG. 7 after header 2. In the embodiments described in this paragraph, the header (e.g., header 2 in FIG. 7) has an identification value that is not one of the conventional values defined in Table 4.57 of the MPEG-4 AAC standard, and instead indicates the eSBR extension payload (such that the extension_type field of the header indicates that the padding data includes eSBR metadata).

[0080] In a first class of embodiments, the invention is an audio processing unit (e.g., a decoder), the audio processing unit comprising: a memory (e.g., buffer 201 of FIG. 3 or FIG. 4) configured to store at least one block of an encoded audio bitstream (e.g., at least one block of an MPEG-4 AAC bitstream); a bitstream payload de-formatter (e.g., element 205 of FIG. 3, or element 215 of FIG. 4) coupled to the memory and configured to demultiplex at least one portion of the block of the bitstream; a decoding subsystem (e.g., elements 202 and 203 of FIG. 3, or elements 202 and 213 of FIG. 4) coupled and configured to decode at least one portion of the audio content of the block of the bitstream, the block comprising: a stuffing element, the stuffing element comprising an identifier (e.g., an “id_syn_ele” identifier having a value of 0x6 in Table 4.85 of the MPEG-4 AAC standard) indicating the start of the stuffing element and stuffing data following the identifier; at least one flag for specifying whether enhanced spectral band replication (eSBR) processing should be performed (e.g., using spectral band replication data and eSBR metadata included in the block) on the audio content of the block; and.

[0081] This flag is eSBR metadata, an example of the flag being the sbrPatchingMode flag. Another example of the flag is the harmonicSBR flag. Both of these flags indicate whether basic form spectral band replication should be performed on the audio data of the block or whether enhanced form spectral band replication should be performed. Basic form spectral band replication is spectral patching and enhanced form spectral band replication is harmonic transposition.

[0082] In some embodiments, the padding data also includes additional eSBR metadata (i.e., eSBR metadata other than the above flags).

[0083] The memory can be a buffer memory (e.g., an implementation of buffer 201 in FIG. 4) that stores (e.g., non-transitorily) at least one block of the encoded audio bitstream.

[0084] It is estimated that the complexity of the execution of eSBR processing (using eSBR harmonic transposition and pre-flattening) by an eSBR decoder during decoding of an MPEG-4 AAC bitstream that includes eSBR metadata (indicating these eSBR tools) is as follows (with respect to typical decoding using the indicated parameters): · Harmonic transposition (16 kbps, 14400 / 28800 Hz) 〇 DFT-based: 3.68 WMOPS (weighted million operations per second) 〇 QMF-based: 0.98 WMOPS · QMF patching preprocessing (pre-flattening): 0.1 WMOPS It is known that DFT-based transposition typically functions better than QMF-based transposition for transient signals.

[0085] According to some embodiments of the present invention, the stuffing element (of the encoded audio bitstream) containing eSBR metadata also has a parameter (e.g., the "bs_extension_id" parameter) with a value (e.g., bs_extension_id = 3) that signals that the value contains eSBR metadata in the stuffing element and that eSBR processing should be performed on the audio content of the corresponding block, and / or a parameter (e.g., the same "bs_extension_id" parameter) with a value (e.g., bs_extension_id = 2) that signals that the sbr_extension() container of the stuffing element contains PS data. For example, as shown in Table 1 below, the fact that such a parameter has a value of bs_extension_id = 2 can signal that the sbr_extension() container of the stuffing element contains PS data, and the fact that such a parameter has a value of bs_extension_id = 3 can signal that the sbr_extension() container of the stuffing element contains eSBR metadata.

Table 1

[0086] According to some embodiments of the invention, the syntax of each spectral band replication extension element containing eSBR metadata and / or PS data is as shown in Table 2 below ("sbr_extension()" represents a container that is a spectral band replication extension element, "bs_extension_id" is as described in Table 1 above, "ps_data" represents PS data, and "esbr_data" represents eSBR metadata).

Table 2

[0087] For example, in some embodiments, esbr_data() may have the syntax shown in Table 3 to indicate these metadata parameters.

Table 3-1

Table 3-2

[0088] The above syntax, as an extension to legacy decoders, enables an efficient implementation of enhanced forms of spectral band replication, such as harmonic transposition. Specifically, the eSBR data in Table 3 includes only the parameters necessary to perform an enhanced form of spectral band replication that is neither already supported in the bitstream nor directly derivable from parameters already supported in the bitstream. All other parameters and processing data necessary to perform the enhanced form of spectral band replication are extracted from parameters that are pre - existing at a predefined location within the bitstream.

[0089] For example, a decoder compliant with MPEG-4 HE-AAC or HE-AAC v2 can be extended to include enhanced forms of spectral band replication such as harmonic transposition. This enhanced form of spectral band replication is in addition to the basic form of spectral band replication already supported by the decoder. In the context of a decoder compliant with MPEG-4 HE-AAC or HE-AAC v2, this basic form of spectral band replication is the QMF spectral patching SBR tool defined in section 4.6.18 of the MPEG-4 AAC standard.

[0090] When performing enhanced spectral band replication, the extended HE-AAC decoder may reuse many of the bitstream parameters already included in the SBR extension payload of the bitstream. Specific parameters that may be reused include, for example, various parameters that determine the master frequency band table. Those parameters include bs_start_freq (a parameter that specifies the start of the master frequency table parameters), bs_stop_freq (a parameter that specifies the end of the master frequency table), bs_freq_scale (a parameter that specifies the number of frequency bands per octave), and bs_alter_scale (a parameter that changes the scale of the frequency bands). The parameters that may be reused also include the parameters that determine the noise band table (bs_noise_bands) and the limiter band table (bs_limiter_bands). Thus, in various embodiments, at least some of the parameters equivalent to those defined in the USAC standard are omitted from the bitstream, thereby reducing the control overhead in the bitstream. Typically, when the parameters defined in the AAC standard have equivalent parameters defined in the USAC standard, the equivalent parameters defined in the USAC standard have the same name as the parameters defined in the AAC standard, for example, envelope scalefactor EOrigMapped. However, the equivalent parameters defined in the USAC standard typically have different values that are "tuned" to the enhanced SBR processing defined in the USAC standard rather than to the SBR processing defined in the AAC standard.

[0091] To improve the subjective quality of audio content that has harmonic frequency structure and strong tonal characteristics, especially at low bitrates, the activation of enhanced SBR is recommended. The values of the corresponding bitstream elements that control those tools (i.e., esbr_data()) can be determined at the encoder by applying a signal-dependent classification mechanism. In general, for encoding music signals at very low bitrates, the use of the harmonic patching method (sbrPatchingMode == 1) is preferred, in which case the core codec can be quite limited in the audio bandwidth. This is especially true when these signals contain a prominent harmonic structure. In contrast, for speech signals and mixed signals, the use of the normal SBR patching method is preferred. This is because it provides better preservation of the temporal structure in speech.

[0092] To improve the performance of the harmonic transposer, a preprocessing step (bs_sbr_preprocessing == 1) can be activated that aims to avoid introducing spectral discontinuities in the signal entering the subsequent envelope adjuster. The operation of this tool is beneficial for signal types that exhibit large level variations by using the rough spectral envelope of the low-band signal for high-frequency reconstruction.

[0093] To improve the transient response of harmonic SBR patching, signal-adaptive frequency-domain oversampling (sbrOversamplingFlag == 1) can be applied. Signal-adaptive frequency-domain oversampling increases the complexity of the transposer's calculations, but since it only benefits frames that contain transient components, the use of this tool is controlled by a bitstream element that is transmitted once per independent SBR channel and per frame.

[0094] A decoder operating in the proposed enhanced SBR mode typically needs to be able to switch between legacy SBR patching and enhanced SBR patching. Thus, depending on the decoder settings, a delay may be introduced that can be as long as the duration of one core audio frame. Typically, this delay is the same for both legacy SBR patching and enhanced SBR patching.

[0095] In addition to these numerous parameters, other data elements can also be reused by the extended HE-AAC decoder when performing enhanced spectral band replication according to embodiments of the invention. For example, envelope data and noise floor data can also be extracted from bs_data_env (envelope scale factor) and bs_noise_env (noise floor scale factor) data and used during enhanced spectral band replication.

[0096] In essence, these embodiments utilize the configuration parameters and envelope data already supported by the legacy HE-AAC or HE-AAC v2 decoder within the SBR extension payload to enable enhanced spectral band replication that requires as little additional transmission data as possible. The metadata is originally tuned for the basic form of HFR (e.g., the spectral conversion operation of SBR), but according to the embodiments, it is used for enhanced form of HFR (e.g., the harmonic transposition of eSBR). As described above, the metadata generally represents the operating parameters (e.g., envelope scale factor, noise floor scale factor, time / frequency grid parameters, sine wave addition information, variable crossover frequency / band, inverse filtering mode, envelope resolution, smoothing mode, frequency interpolation mode) intended and tuned to be used in the basic form of HFR (e.g., linear spectral conversion). However, this metadata can be combined with additional metadata parameters specific to the enhanced form of HFR (e.g., harmonic transposition) and used to efficiently and effectively process audio data using the enhanced form of HFR.

[0097] Therefore, by relying on already defined bitstream elements (e.g., those within the SBR extension payload) and only adding the parameters necessary to support enhanced spectral band replication, an extended decoder that supports enhanced spectral band replication can be created very efficiently. This data reduction feature, in combination with placing the newly added parameters in a reserved data field such as an extended container, guarantees that the bitstream is backward compatible with legacy decoders that do not support enhanced spectral band replication, substantially reducing the barrier to creating a decoder that supports enhanced spectral band replication.

[0098] In Table 3, the numbers in the right column indicate the number of bits of the corresponding parameters in the left column.

[0099] In some embodiments, the SBR object type defined in MPEG-4 AAC is updated to include aspects of the SBR-Tool and the enhanced SBR (eSBR) tool signaled by an SBR extension element (bs_extension_id == EXTENSION_ID_ESBR). When a decoder supports and detects this SBR extension element, the decoder uses the signaled aspects of the enhanced SBR tool. The SBR object type updated in this way is referred to as an SBR enhancement.

[0100] In one embodiment, the invention is a method comprising encoding audio data to produce an encoded bitstream (e.g., an MPEG-4 AAC bitstream), including eSBR metadata in at least one segment of at least one block of the encoded bitstream, and including audio data in at least one other segment of the block. In a typical embodiment, the method includes multiplexing the audio data with the eSBR metadata in each block of the encoded bitstream. In a typical decoding of the encoded bitstream in an eSBR decoder, the decoder extracts the eSBR metadata from the bitstream (including by parsing and demultiplexing the eSBR metadata and the audio data), and processes the audio data using the eSBR metadata to produce a stream of decoded audio data.

[0101] Another aspect of the invention is an eSBR decoder configured to perform eSBR processing (e.g., using at least one of eSBR tools known as harmonic transposition or pre-flattening) during decoding of an encoded audio bitstream (e.g., an MPEG-4 AAC bitstream) that does not include eSBR metadata. An example of such a decoder will be described with reference to FIG. 5.

[0102] The eSBR decoder (400) of FIG. 5 includes a buffer memory 201 (the same as the memory 201 of FIGS. 3 and , a bitstream payload deformatter 215 (the same as the deformatter 215 of FIG. 4), an audio decoding subsystem 202 (also referred to as a "core" decoding stage or a "core" decoding subsystem, the same as the core decoding subsystem 202 of FIG. 3), an eSBR control data generation subsystem 401, and an eSBR processing stage 203 (the same as the stage 203 of FIG. 3), connected as shown. Typically, decoder 400 also includes other processing elements (not shown).

[0103] In the operation of decoder 400, a series of blocks of the encoded audio bitstream (MPEG-4 AAC bitstream) received by decoder 400 are asserted from buffer 201 to de-formatter 215.

[0104] The de-formatter 215 demultiplexes each block of the bitstream, extracts the SBR metadata (including the quantized envelope data) therefrom, and typically also extracts other metadata. The de-formatter 215 is configured to assert at least the SBR metadata to the eSBR processing stage 203. The de-formatter 215 is also combined and configured to extract audio data from each block of the bitstream and assert the extracted audio data to the decoding subsystem (decoding stage) 202.

[0105] The audio decoding subsystem 202 of the decoder 400 decodes the audio data extracted by the de-formatter 215 (such decoding may be referred to as "core" decoding process), generates the decoded audio data, and is configured to assert the decoded audio data to the eSBR processing stage 203. This decoding is performed in the frequency domain. Typically, the final stage of processing in the subsystem 202 applies a frequency-domain to time-domain conversion to the decoded frequency-domain audio data such that the output of the subsystem 202 is the decoded audio data in the time domain. The stage 203 applies the SBR tools (and eSBR tools) indicated by the SBR metadata (extracted by the de-formatter 215) and the eSBR metadata generated by the subsystem 401 to the decoded audio data (i.e., performs SBR and eSBR processing on the output of the decoding subsystem 202 using the SBR and eSBR metadata) to generate the fully decoded audio data output from the decoder 400. Typically, the decoder 400 includes a memory (accessible by the subsystem 202 and the stage 203) that stores the de-formatted audio data and metadata output from the de-formatter 215 (and optionally the subsystem 401 as well), and the stage 203 is configured to access the audio data and metadata as needed during the SBR and eSBR processing. The SBR processing in the stage 203 can be regarded as a post-processing of the output of the core decoding subsystem 202. Optionally, the decoder 400 is also combined and configured to perform upmixing on the output of the stage 203 to generate the fully decoded and upmixed audio output from the decoder 400 (this may apply the parametric stereo ("PS") tools defined in the MPEG-4 AAC standard using the PS metadata extracted by the de-formatter 215).

[0106] Parametric stereo is an encoding tool that represents a stereo signal using linear downmixing of the left and right channels of the stereo signal and a set of spatial parameters that describe the stereo image. Parametric stereo typically uses three types of spatial parameters: (1) inter-channel intensity differences (IID) that describe the intensity difference between channels, (2) inter-channel phase differences (IPD) that describe the phase difference between channels, and (3) inter-channel coherence (ICC) that describes the coherence (or similarity) between channels. Coherence can be measured as the maximum of the cross-correlation as a function of time or phase. These three parameters generally enable high-quality reconstruction of the stereo image. However, the IPD parameter only describes the relative phase difference between the channels of the stereo input signal and does not show the distribution of these phase differences across the left and right channels. Therefore, a fourth type of parameter that describes the overall phase offset or overall phase difference can be additionally used. In the stereo reconstruction process, consecutive window segments of both the received downmixed signal s[n] and the decorrelated version d[n] of the received downmixing are processed together with the spatial parameters, l k (n)=H 11 (k,n)s k (n)+H 21 (k,n)d k (n) r k (n)=H 12 (k,n)s k (n)+H 22 (k,n)d k (n) Accordingly, a left reconstruction signal (l k (n)) and a right reconstruction signal (r k (n)) are generated, where H 11 , H 12 , H 21 and H 22is defined by stereo parameters. Signal l k (n) and signal r k (n) are finally converted back to the time domain by frequency-time conversion.

[0107] The control data generation subsystem 401 of FIG. 5 detects at least one characteristic of the encoded audio bitstream to be decoded and is coupled and configured to generate eSBR control data (which is or may include any type of eSBR metadata included in the encoded audio bitstream according to other embodiments of the invention) in response to at least one result of the detection step. The eSBR control data is asserted at stage 203 to trigger and / or control the application of individual eSBR tools or combinations of eSBR tools in response to detecting a particular characteristic (or combination of characteristics) of the bitstream. For example, to control the execution of eSBR processing using harmonic transposition, some embodiments of the control data generation subsystem 401 set the sbrPatchingMode[ch] parameter (and assert the set parameter at stage 203) in response to detecting whether the bitstream represents music, a transient detector that sets the sbrOversamplingFlag[ch] parameter (and asserts the set parameter at stage 203) in response to detecting the presence or absence of transients in the audio content represented by the bitstream, and / or a pitch detector that sets the sbrPitchInsFlag[ch] and sbrPitchIns[ch] parameters (and asserts the set parameters at stage 203) in response to detecting the pitch of the audio content represented by the bitstream. Another aspect of the invention is an audio bitstream decoding method performed by any of the embodiments of the invention decoder described in this and the previous paragraph.

[0108] Aspects of the invention include a type of encoding or decoding method configured (e.g., programmed) to be executed by an embodiment of an inventive APU, system, or apparatus. Other aspects of the invention include a system or apparatus configured (e.g., programmed) to execute an embodiment of the inventive method, and a computer-readable medium (e.g., a disk) storing (e.g., non-transitorily) code for implementing an embodiment of the inventive method or its steps. For example, the inventive system can be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed (or otherwise configured) in software or firmware to perform any of a variety of processes on data, including an embodiment of the inventive method or its steps. Such a general-purpose processor can be or be included in a computer system including an input device, a memory, and a processing circuit programmed (and / or otherwise configured) to perform an embodiment of the inventive method (or its steps) in response to data asserted thereto.

[0109] Embodiments of the present invention can be implemented in hardware, firmware, software, or a combination of both (e.g., a programmable logic array). Unless otherwise stated, an algorithm or process included as part of the invention is not inherently related to a particular computer or other device. In particular, various general-purpose machines can be used with the programs described according to the teachings herein, or it may be more convenient to construct a more specialized device (e.g., an integrated circuit) to perform the required method steps. Accordingly, the invention can be implemented in one or more programmable computer systems each having at least one processor, at least one data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or port, and at least one output device or port (e.g., one implementing any of the elements of FIG. 1, or the encoder 100 (or its elements) of FIG. 2, or the decoder 200 (or its elements) of FIG. 3, or the decoder 210 (or its elements) of FIG. 4, or the decoder 400 (or its elements) of FIG. 5) by one or more computer programs executed thereon. The program code is applied to input data to perform the functions described herein and generate output information. The output information is provided to one or more output devices in a known manner.

[0110] Each such program can be implemented in any desired computer language (including machine language, assembly language, or high-level procedural, logical, or object-oriented programming languages) to communicate with the computer system. In any case, the language can be a compiled language or an interpreted language.

[0111] For example, when implemented by a computer software instruction sequence, the various functions and steps of the embodiments of the invention may be implemented by a multi-threaded software instruction sequence running on appropriate digital signal processing hardware. In that case, the various devices, steps, and functions of the embodiments may correspond to a part of the software instructions.

[0112] Each such computer program is preferably stored or downloaded to a storage medium or storage device (e.g., solid state memory or medium, or magnetic or optical medium) readable by a general-purpose or special-purpose programmable computer. When the storage medium or storage device is read by a computer system, the computer is configured and operated to execute the procedures described herein. The inventive system may also be implemented as a computer-readable storage medium comprising (i.e., storing) a computer program. A storage medium configured as such causes a computer system to operate in a specific predefined manner so as to execute the functions described herein.

[0113] Numerous embodiments of the invention have been described. However, it is to be understood that various changes can be made without departing from the spirit and scope of the invention. In light of the teachings above, numerous changes and modifications of the present invention are possible. For example, in order to assist with efficient implementation, a phase shift may be used in combination with a complex QMF analysis and synthesis filter bank. The analysis filter bank is responsible for filtering the time-domain low-band signal generated by the core decoder into a plurality of subbands (e.g., QMF subbands). The synthesis filter bank is responsible for combining the regenerated high-band generated by a selected HFR technique (indicated by the received sbrPatchingMode parameter) with the decoded low-band to generate a wideband output audio signal. A given filter bank implementation operating in a particular sample rate mode, such as normal dual-rate operation or downsampling SBR mode, however, should not have a phase shift that depends on the bitstream. The QMF bank used in SBR is a complex exponential extension of the theory of cosine-modulated filter banks. It can be shown that when extending a cosine-modulated filter bank using complex exponential modulation, the alias cancellation constraint is not used. Therefore, in the SBR QMF bank, both the analysis filter h k (n) and the synthesis filter f k (n) are [Number] It can be defined by, where p0(n) is a real-valued symmetric or asymmetric prototype filter (typically, a low-pass prototype filter), M represents the number of channels, and N is the prototype filter order. The number of channels used in the analysis filter bank can be different from the number of channels used in the synthesis filter bank. For example, the analysis filter bank can have 32 channels, and the synthesis filter bank can have 64 channels. When operating the synthesis filter bank in the downsampling mode, the synthesis filter bank can have only 32 channels. Since the subband samples from the filter bank are complex-valued, additional phase shift steps that may be channel-dependent can be added to the analysis filter bank. These additional phase shifts need to be compensated for before the synthesis filter bank. In principle, the phase shift terms can be set to any value without disrupting the operation of the QMF analysis / synthesis chain, but they may also be restricted to specific values for compliance verification. The SBR signal will be affected by the choice of the phase factor, but the low-pass signal coming from the core decoder will not be affected. The sound quality of the output signal is not affected.

[0114] The coefficients of the prototype filter p0(n) can be defined with a length L of 640, as shown in Table 4 below.

Table 4-1

Table 4-2

Table 4-3

Table 4-4

Table 4-5

[0115] The tuning of SBR-related control information typically (as described above) does not depend on the details of the transposition. However, in some embodiments, to improve the quality of the regenerated signal, certain elements of the control data may be transmitted simultaneously within the eSBR extension container (bs_extension_id == EXTENSION_ID_ESBR). Some of the elements transmitted simultaneously may include noise floor data (e.g., a noise floor scale factor and a parameter indicating the direction in either the frequency or time domain of the delta coding for each noise floor), inverse filtering data (e.g., a parameter indicating an inverse filtering mode selected from no inverse filtering, low-level inverse filtering, intermediate-level inverse filtering, and strong-level inverse filtering), and missing harmonic data (e.g., a parameter indicating whether a sine wave should be added to a specific frequency band of the regenerated high band). All of these elements rely on the synthesis emulation of the decoder's transposer executed by the encoder and can thus enhance the quality of the regenerated signal when appropriately adjusted for the selected transposer.

[0116] Specifically, in some embodiments, the missing harmonic and inverse filtering control data are transmitted within the eSBR extension container (along with other bitstream parameters in Table 3) and adjusted for the harmonic transposer of eSBR. The additional bitrate required to transmit these two classes of metadata for the harmonic transposer of eSBR is relatively low. Therefore, sending the adjusted missing harmonic and / or inverse filtering control data in the eSBR extension container will enhance the quality of the audio generated by the transposer with minimal impact on the bitrate. To ensure backward compatibility with legacy decoders, parameters adjusted for the spectral conversion process of SBR can also be sent in the bitstream as part of the SBR control data using either implicit or explicit signaling.

[0117] The complexity of the decoder with SBR enhancement described in this application must be limited so as not to significantly increase the overall computational complexity of the implemented one. Preferably, when using the eSBR tool, the PCU (MOP) of the SBR object type is 4.5 or less, and when using the eSBR tool, the RCU of the SBR object type is 3 or less. The processing power by approximation is given in Processor Complexity Units (PCU) defined by the number of integer MOPS. The RAM usage by approximation is given in RAM Complexity Units (RCU) defined by the number of integer kWords (1000 words). The number of RCUs does not include the working buffer that can be shared among different objects and / or channels. Also, the PCU is proportional to the sampling frequency. The PCU value is given in MOPS (Million Operations per Second) per channel, and the RCU value is given in kWords per channel.

[0118] Special care is required for compressed data, such as HE-AAC encoded audio, that can be decoded by different decoder configurations. In this case, decoding can be performed in a backward-compatible (AAC only) and enhanced (AAC+SBR) manner. When the compressed data allows both backward-compatible decoding and enhanced decoding, and the decoder is operating in an enhanced manner, such that it uses a post-processor (e.g., the SBR post-processor in HE-AAC) that inserts some additional delay, it must be ensured that this additional time delay that occurs for the backward-compatible mode, as described by the corresponding value of n, is taken into account when presenting the synthesis unit. To ensure that the synthesis timestamps are handled correctly (so that the audio remains synchronized with other media), the additional delay introduced by post-processing, given by the number of samples (per audio channel) at the output sample rate, is 3010 when the decoder operation mode includes the SBR enhancements (including eSBR) described in this application. Therefore, in the audio synthesis unit, when the decoder operation mode includes the SBR enhancements described in this application, the synthesis time is applied to the 3011th audio sample within the synthesis unit.

[0119] To improve the subjective quality of audio content, particularly at low bitrates and with harmonic frequency structure and strong tonal characteristics, enhanced SBR should be activated. The value of the corresponding bitstream element that controls those tools (i.e., esbr_data()) can be determined at the encoder by applying a signal-dependent classification mechanism.

[0120] In general, for encoding music signals at very low bitrates, it is preferable to use the harmonic patching method (sbrPatchingMode==0), in which case the core codec can be considerably restricted in the audio bandwidth. This applies especially when these signals contain a prominent harmonic structure. In contrast, for speech and mixed signals, it is preferable to use the normal SBR patching method, since it provides better preservation of the temporal structure in speech.

[0121] To improve the performance of the harmonic transposer, a preprocessing step (bs_sbr_preprocessing==1) can be activated to avoid introducing spectral discontinuities in the signal entering the subsequent envelope adjuster. The operation of this tool is beneficial for signal types showing large level variations, which make extensive use of the coarse spectral envelope of the low-band signal for high-frequency reconstruction.

[0122] To improve the transient response of harmonic SBR patching (sbrPatchingMode==0), signal-adaptive frequency-domain oversampling (sbrOversamplingFlag==1) can be applied. Signal-adaptive frequency-domain oversampling increases the complexity of the transposer calculations, but the use of this tool is controlled by the bitstream elements that are transmitted once per independent SBR channel and per frame, since it only benefits frames containing transient components.

[0123] Typical bitrate settings recommended for HE-AACv2 with SBR enhancement (i.e., enabling the harmonic transposer of the eSBR tool) correspond to 20 - 32 kbps for stereo audio content at either a 44.1 kHz or 48 kHz sampling rate. The relative subjective quality gain of SBR enhancement increases towards the lower bitrate boundary, and a properly configured encoder enables extending this range to even lower bitrates. The bitrates presented above are only recommendations and can be adapted to specific service requirements.

[0124] A decoder operating in the proposed enhanced SBR mode should typically be able to switch between legacy SBR patching and enhanced SBR patching. Thus, depending on the decoder settings, a delay that can be as long as the duration of one core audio frame may be introduced. Typically, this delay is equal for both legacy SBR patching and enhanced SBR patching.

[0125] It should be understood that within the scope of the appended claims, the invention may be practiced otherwise than as specifically described herein. Any reference signs in the following claims are included for illustrative purposes only and should not be used to construe or limit the claims in any way.

[0126] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEE).

[0127] EEE1. A method for performing high-frequency reconstruction of an audio signal, the method comprising: receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata; decoding the audio data to generate a decoded low-band audio signal; Extract the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata includes operation parameters for a high-frequency reconstruction process, the operation parameters include patching mode parameters placed within a backward compatibility extension container of the encoded audio bitstream, a first value of the patching mode parameters indicates spectral conversion, and a second value of the patching mode parameters indicates harmonic transposition by phase vocoder frequency spreading. Filter the decoded low-band audio signal to generate a filtered low-band audio signal. Use the filtered low-band audio signal and the high-frequency reconstruction metadata to regenerate a high-band portion of the audio signal, the regenerating including spectral conversion when the patching mode parameter is the first value, and the regenerating including harmonic transposition by phase vocoder frequency spreading when the patching mode parameter is the second value. Combine the filtered low-band audio signal with the regenerated high-band portion to form a wide-band audio signal. comprising: The filtering, the regenerating, and the combining are performed as post-processing operations with a delay of 3010 samples or less per audio channel, and the spectral conversion includes maintaining a ratio between a tonal component and a noise-like component by adaptive inverse filtering. method.

[0128] EEE2. The encoded audio bitstream further includes padding elements, the padding elements have an identifier indicating the start of the padding elements and padding data after the identifier, and the padding data includes the backward compatibility extension container, the method of EEE1.

[0129] EEE3. The identifier is a 3-bit unsigned integer with the most significant bit transmitted first and having a value of 0x6, the method of EEE2.

[0130] EEE4. The filling data includes an extended payload, the extended payload includes spectrum band replication extension data, and the extended payload is identified by a 4-bit unsigned integer with the most significant bit transmitted first and having a value of '1101' or '1110'. Optionally, the spectrum band replication extension data includes an optional spectrum band replication header, spectrum band replication data after the header, and a spectrum band replication extension element after the spectrum band replication data, the spectrum band replication extension element including a flag, and includes the method of EEE2 or 3.

[0131] EEE5. The high-frequency reconstruction metadata includes an envelope scale factor, a noise floor scale factor, time / frequency grid information, or a parameter indicating a crossover frequency, and is any one of the methods of EEE1 to EEE4.

[0132] EEE6. The backward compatibility extension container further includes a flag indicating whether additional preprocessing is used to avoid discontinuities in the shape of the spectrum envelope of the high-band portion when the patching mode parameter is equal to the first value, the first value of the flag enabling the additional preprocessing, and the second value of the flag disabling the additional preprocessing, and is any one of the methods of EEE1 to EEE5.

[0133] EEE7. The additional preprocessing includes calculating a pre-gain curve using linear prediction filter coefficients, and is the method of EEE6.

[0134] EEE8. The rear interchangeable expansion container further includes a flag indicating whether signal adaptive frequency domain oversampling should be applied when the patching mode parameter is equal to the second value, where a first value of the flag enables the signal adaptive frequency domain oversampling and a second value of the flag disables the signal adaptive frequency domain oversampling, according to any one of EEE1 to EEE5.

[0135] EEE9. The signal adaptive frequency domain oversampling according to EEE8 is applied only to frames including transient signals.

[0136] EEE10. The harmonic transposition by phase vocoder frequency spreading is performed with an estimated complexity of 4.5 million operations per second and 3k words of memory or lower, according to any one of EEE1 to EEE9.

[0137] EEE11. A non-transitory computer-readable medium including instructions that, when executed by a processor, perform any one of the methods of EEE1 to EEE10.

[0138] EEE12. A computer program product having instructions that, when executed by a computing device or system, cause the computing device or system to perform any one of the methods of EEE1 to EEE10.

[0139] EEE13. An audio processing unit that performs high-frequency reconstruction of an audio signal, the audio processing unit includes an input interface that receives an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata, and a core audio decoder that decodes the audio data to generate a decoded low-band audio signal. A de-formatter that extracts the high-frequency reconstruction metadata from the encoded audio bit stream, the high-frequency reconstruction metadata includes operation parameters for a high-frequency reconstruction process, the operation parameters include patching mode parameters placed in a backward compatibility extension container of the encoded audio bit stream, a first value of the patching mode parameters indicates spectral conversion, and a second value of the patching mode parameters indicates harmonic transposition by phase vocoder frequency spreading, and a de-formatter; An analysis filter bank that filters the decoded low-band audio signal to generate a filtered low-band audio signal; A high-frequency regenerator that reconstructs a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, the reconstructing includes spectral conversion when the patching mode parameter is the first value, and the reconstructing includes harmonic transposition by phase vocoder frequency spreading when the patching mode parameter is the second value, and a high-frequency regenerator; A synthesis filter bank that combines the filtered low-band audio signal with the regenerated high-band portion to form a wide-band audio signal; having The analysis filter bank, the high-frequency regenerator, and the synthesis filter bank are executed by a post-processor with a delay of 3010 samples or less per audio channel, and the spectral conversion has maintaining a ratio between a pitch component and a noise-like component by adaptive inverse filtering. An audio processing unit.

[0140] EEE14. The harmonic transposition by phase vocoder frequency spreading of the audio processing unit of EEE13 is executed with an estimated complexity of 4.5 million operations per second and 3k words of memory or lower.

Claims

1. A method for performing high-frequency reconstruction of an audio signal, the method comprising: Receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata, the high-frequency reconstruction metadata including a noise scale factor; Decoding the audio data to generate a decoded low-band audio signal; Extracting the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata including operation parameters for a high-frequency reconstruction process, the operation parameters including patching mode parameters placed within a backward compatibility extension container of the encoded audio bitstream, a first value of the patching mode parameters indicating spectral conversion, and a second value of the patching mode parameters indicating harmonic transposition by phase vocoder frequency spreading; Filtering the decoded low-band audio signal to generate a filtered low-band audio signal; Using the filtered low-band audio signal and the high-frequency reconstruction metadata to regenerate a high-band portion of the audio signal, the regenerating including spectral conversion when the patching mode parameters are the first value, and the regenerating including harmonic transposition by phase vocoder frequency spreading when the patching mode parameters are the second value; Having; The filtering and the regenerating are performed as post-processing operations with a delay of 3010 samples per audio channel, and the spectral conversion includes maintaining a ratio between a pitch component and a noise-like component by adaptive inverse filtering. Method.

2. The method according to claim 1, wherein the harmonic transposition by phase vocoder frequency spreading is performed with an estimated complexity of 4.5 million operations per second or less and 3k words or less of memory.

3. A non-transitory computer-readable medium including instructions that, when executed by a processor, perform the method according to claim 1.

4. A computer program stored on a non-transitory computer-readable medium having instructions that, when executed by a computing device or system, cause the computing device or system to perform the method according to claim 1.

5. An audio processing unit that performs high-frequency reconstruction of an audio signal, the audio processing unit comprising: An input interface that receives an encoded audio bitstream, the encoded audio bitstream including audio data representing a low-band portion of the audio signal and high-frequency reconstruction metadata, the high-frequency reconstruction metadata including a noise scale factor; an input interface; A core audio decoder that decodes the audio data to generate a decoded low-band audio signal; A de-formatter that extracts the high-frequency reconstruction metadata from the encoded audio bitstream, the high-frequency reconstruction metadata including operation parameters for a high-frequency reconstruction process, the operation parameters including patching mode parameters placed within a backward compatibility extension container of the encoded audio bitstream, a first value of the patching mode parameters indicating spectral conversion, and a second value of the patching mode parameters indicating harmonic transposition by phase vocoder frequency spreading; a de-formatter; An analysis filter bank that filters the decoded low-band audio signal to generate a filtered low-band audio signal; A high-frequency regenerator that reconstructs a high-band portion of the audio signal using the filtered low-band audio signal and the high-frequency reconstruction metadata, the reconstructing including spectral conversion when the patching mode parameter is the first value, and the reconstructing including harmonic transposition by phase vocoder frequency spreading when the patching mode parameter is the second value; a high-frequency regenerator; comprising The analysis filter bank and the high-frequency regenerator are executed by a post-processor with a delay of 3010 samples per audio channel, and the spectral conversion has maintaining the ratio between the tonal component and the noise-like component by adaptive inverse filtering. Audio processing unit.

Citation Information

Patent Citations

  • Speech coder and method, speech decoder and method, speech band spreading apparatus and method

    JP2010020251A

  • Audio signal synthesizer and audio signal encoder

    JP2011527447A

  • Bandwidth expansion coding device, bandwidth expansion decoding device, and phase vocoder

    JP2012531632A

  • Improvement of harmonic transposition based on subbandblocking

    JP2013516652A

  • Apparatus and method for improved amplitude response and temporal alignment in a bandwidth expansion method based on a phase vocoder for audio signals.

    JP2013521536A