Low latency low frequency effect codec
By filtering and frequency domain representation of the LFE channel signal, combined with entropy decoding and silent frame indicators, the low latency and high efficiency coding problem of the IVAS codec in low frequency effect channel processing is solved, realizing low latency and high efficiency audio signal processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-01
- Publication Date
- 2026-03-24
AI Technical Summary
Existing IVAS codecs struggle to achieve low latency and efficient encoding when processing low-frequency effect (LFE) channels, especially in low/medium bitrate scenarios. They cannot effectively handle deep bass sounds from 20Hz to 120Hz, and cannot handle audio up to 400Hz in high bitrate scenarios.
A low-pass filter is used to filter the time-domain LFE channel signal and convert it into a frequency-domain representation. An entropy decoder is used to quantize the coefficients in the frequency band group. Encoding is performed by adjusting the scaling shift factor and the entropy decoder. A silent frame indicator is combined to reduce the bit rate, ensuring low latency and efficient coding.
A low-latency LFE codec is implemented, capable of carrying audio up to 300Hz in low/medium bit rate scenarios and up to 400Hz in high bit rate scenarios. It features a low algorithm latency and a low bit rate silent mode, with the bit rate fluctuating between 2,000 bits/second and 4,000 bits/second during the active frame period and 50 bits/second during the silent frame period.
Smart Images

Figure CN114424282B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 895,049, filed September 3, 2019, and U.S. Provisional Patent Application No. 63 / 069,420, filed August 24, 2020, each of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0003] The disclosure relates generally to audio signal processing, and in particular to processing a low frequency effects (LFE) channel. BACKGROUND
[0004] For example, a standardization effort for immersive services includes developing an Immersive Voice and Audio Services (IVAS) codec for sound, multi-stream teleconferencing, virtual reality (VR), user-generated live and non-live content streaming. The goal of the IVAS standard is to develop a single codec that is excellent in audio quality, low in latency, supports spatial audio coding, has proper bit rate range, high quality error resilience, and practical implementation complexity. To achieve this goal, it is desirable to develop an IVAS codec that can handle low latency LFE operation based on a device capable of IVAS or any other device capable of processing LFE signals. The LFE channel is used for deep bass tonal sound ranging from 20 Hz to 120 Hz and is typically sent to a speaker designed to reproduce low frequency audio content. SUMMARY
[0005] Embodiments of a configurable low latency LFE codec are disclosed.
[0006] In some embodiments, a method of encoding a low frequency effects (LFE) channel includes receiving, using one or more processors, a time domain LFE channel signal; filtering, using a low pass filter, the time domain LFE channel signal; converting, using the one or more processors, the filtered time domain LFE channel signal to a frequency domain representation of the LFE channel signal including a number of coefficients representing a spectrum of the LFE channel signal; arranging, using the one or more processors, coefficients into a number of sub-band groups corresponding to different frequency bands of the LFE channel signal; quantizing, using the one or more processors, coefficients in each sub-band group according to a frequency response curve of the low pass filter; encoding, using the one or more processors, the quantized coefficients in the sub-band groups using an entropy coder tuned for each sub-band group; generating, using the one or more processors, a bitstream including the encoded quantized coefficients; and storing, using the one or more processors, the bitstream on a storage device or streaming the bitstream to a downstream device.
[0007] In some implementations, quantizing the coefficients in each band group further includes generating a scaling shift factor based on a maximum number of available quantization points and a sum of absolute values of the coefficients, and quantizing the coefficients using the scaling shift factor.
[0008] In some implementations, if the quantized coefficients exceed the maximum number of quantization points, the scaling shift factor is decreased and the coefficients are quantized again.
[0009] In some implementations, the quantization points are different for each band group.
[0010] In some implementations, the coefficients in each band group are quantized according to a fine quantization scheme or a coarse quantization scheme, wherein more quantization points are allocated to a respective sub-band group with the fine quantization scheme than are allocated to the sub-band group according to the coarse quantization scheme.
[0011] In some implementations, the sign bits of the coefficients are coded separately from the coefficients.
[0012] In some implementations, there are four sub-band groups, and a first sub-band group corresponds to a first frequency range of 0 Hz to 100 Hz, a second sub-band group corresponds to a second frequency range of 100 Hz to 200 Hz, a third sub-band group corresponds to a third frequency range of 200 Hz to 300 Hz, and a fourth sub-band group corresponds to a fourth frequency range of 300 Hz to 400 Hz.
[0013] In some implementations, the entropy coder is an arithmetic entropy coder.
[0014] In some implementations, converting the filtered time-domain LFE channel signal into a frequency-domain representation of the LFE channel signal that includes a number of coefficients representing a spectrum of the LFE channel signal further includes determining a first step size for the LFE channel signal, specifying a first window size of a window function based on the first step size, applying the first window size to one or more frames of the time-domain LFE channel signal, and applying a modified discrete cosine transform (MDCT) to the windowed frames to generate the coefficients.
[0015] In some implementations, the method further includes determining a second step size for the LFE channel signal, specifying a second window size of the window function based on the second step size, and applying the second window size to the one or more frames of the time-domain LFE channel signal.
[0016] In some embodiments, the first step size is N milliseconds (ms), N is greater than or equal to 5 ms and less than or equal to 60 ms, the first window size is higher than or equal to 10 ms, the second step size is 5 ms and the second window size is 10 ms.
[0017] In some embodiments, the first step size is 20 milliseconds (ms), the first window size is 10 ms or 20 ms or 40 ms, the second step size is 10 ms and the second window size is 10 ms or 20 ms.
[0018] In some embodiments, the first step size is 10 milliseconds (ms), the first window size is 10 ms or 20 ms, the second step size is 5 ms and the second window size is 10 ms.
[0019] In some embodiments, the first step size is 20 milliseconds (ms), the first window size is 10 ms, 20 ms or 40 ms, the second step size is 5 ms and the second window size is 10 ms.
[0020] In some embodiments, the window function is a Kaiser-Bessel Derived (KBD) window function with a configurable fade length.
[0021] In some embodiments, the low pass filter is a fourth order Butterworth filter low pass filter with a cutoff frequency of about 130 Hz or lower than 130 Hz.
[0022] In some embodiments, the method further comprises: determining, using the one or more processors, whether an energy level of a frame of the LFE channel signal is below a threshold; generating a silence frame indicator to indicate to the decoder in accordance with the energy level being below a threshold level; inserting the silence frame indicator into metadata of the LFE channel bitstream; and reducing LFE channel bit rate when a silence frame is detected.
[0023] In some implementations, a method of decoding a low frequency effect (LFE) channel bitstream includes receiving, using one or more processors, an LFE channel bitstream including entropy coded coefficients representing a spectrum of a time domain LFE channel signal; decoding, using the one or more processors, the quantized coefficients using an entropy decoder; inverse quantizing, using the one or more processors, the inverse quantized coefficients, wherein the coefficients have been quantized in sub-band groups corresponding to frequency bands according to a frequency response curve of a low pass filter used to filter the time domain LFE channel signal in an encoder; converting, using the one or more processors, the inverse quantized coefficients to a time domain LFE channel signal; adjusting, using the one or more processors, a delay of the time domain LFE channel signal; and filtering, using a low pass filter, the delay adjusted LFE channel signal.
[0024] In some implementations, an order of the low pass filter is configured to ensure that a first total algorithmic delay due to encoding and decoding the LFE channel in a multi-channel audio signal including the LFE channel signal is less than or equal to a second total algorithmic delay due to encoding and decoding other channels.
[0025] In some implementations, the method further includes determining whether the second total algorithmic delay exceeds a threshold; and in accordance with the second total algorithmic delay exceeding the threshold, configuring the low pass filter as an Nth order low pass filter, where N is an integer greater than or equal to 2; and in accordance with the second total algorithmic delay not exceeding the threshold, configuring the order of the low pass filter to be less than N.
[0026] Other implementations disclosed herein relate to systems, devices, and computer readable media. The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
[0027] The particular embodiments disclosed herein provide one or more of the following advantages. The disclosed low-latency LFE codec: 1) targets primarily the LFE channel; 2) targets primarily the frequency range of 20 Hz to 120 Hz, but carries audio up to 300 Hz in low / medium bitrate scenarios and up to 400 Hz in high bitrate scenarios; 3) enables low bitrates by applying a quantization scheme according to the frequency response curve of the input lowpass filter; 4) has low algorithmic latency and is designed to operate in 20 millisecond (ms) steps and has a total algorithmic latency (including framing) of 33 msec; 5) can be configured to smaller steps and lower algorithmic latency to support other scenarios, including being configured to steps as low as 5 msec and a total algorithmic latency (including framing) of 13 msec; 6) automatically selects a lowpass filter at the decoder output based on the latency available to the LFE codec; 7) has a silence mode with a low bitrate of 50 bits per second (bps) during silence; and 8) during active frames, the bitrate fluctuates between 2 kilobits per second (kbps) to 4 kbps based on the quantization level used, and during silence frames the bitrate is 50 bps. BRIEF DESCRIPTION OF DRAWINGS
[0028] In the drawings, specific arrangements or orders of illustrative elements (e.g., elements that represent devices, units, instruction blocks, and data elements) are shown for ease of explanation. However, one of skill in the art will understand that the specific order or arrangement of the illustrative elements in the drawings is not intended to imply a particular processing order or sequence or process separation. Furthermore, inclusion of illustrative elements in the drawings is not intended to imply that such elements are required in all embodiments, or that the features represented by such elements can not be included in some embodiments or combined with other elements in some embodiments.
[0029] Furthermore, in the drawings, connection elements (e.g., lines or arrows) illustrating connections, relationships, or associations between or among two or more other illustrative elements are not intended to limit the scope of the disclosure to only such
[0030] Figure 1 FIG. 1 illustrates an IVAS and LFE codec for encoding and decoding IVAS and LFE bitstreams, in accordance with one or more embodiments.
[0031] Figure 2Ais a block diagram illustrating LFE encoding according to one or more embodiments.
[0032] Figure 2B is a block diagram illustrating LFE decoding according to one or more embodiments.
[0033] Figure 3 is a plot illustrating the frequency response of a fourth order Butterworth low pass filter with a corner cutoff point of 130 Hz according to one or more embodiments.
[0034] Figure 4 is a plot illustrating a Fielder window according to one or more embodiments.
[0035] Figure 5 illustrates the variation of fine quantization points with frequency according to one or more embodiments.
[0036] Figure 6 illustrates the variation of coarse quantization points with frequency according to one or more embodiments.
[0037] Figure 7 illustrates the probability distribution of quantized MDCT coefficients when fine quantization according to one or more embodiments.
[0038] Figure 8 illustrates the probability distribution of quantized MDCT coefficients when coarse quantization according to one or more embodiments.
[0039] Figure 9 is a flowchart of a process for encoding modified discrete cosine transform (MDCT) coefficients according to one or more embodiments.
[0040] Figure 10 is a flowchart of a process for decoding modified discrete cosine transform (MDCT) coefficients according to one or more embodiments.
[0041] Figure 11 is a block diagram of a system for implementing the reference Figures 1 to 10 features and processes described.
[0042] Like reference symbols in the various drawings indicate like elements. DETAILED DESCRIPTION
[0043] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described embodiments. It will be apparent, however, to one skilled in the art that the various described embodiments can be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features described below can each be used independently of one another or in any combination with any other feature.
[0044] Nomenclature
[0045] As used herein, the term "comprises" and variations thereof are to be construed as open-ended terms that mean "including, but not limited to." The term "or" is to be construed as "and / or" unless the context clearly indicates otherwise. The term "based on" is to be construed as "based at least in part on." The term "one embodiment" and "an embodiment" are to be construed as "at least one embodiment." The term "another embodiment" is to be construed as "at least one other embodiment." The term "determined" or "determining" is to be construed as obtaining, receiving, calculating, computing, estimating, predicting or deriving. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention belongs.
[0046] System Overview
[0047] Figure 1 FIG. illustrates an IVAS codec 100 for encoding and decoding an IVAS bitstream including an LFE channel bitstream, in accordance with one or more embodiments. At encoding, the IVAS codec 100 receives N+1 channels of audio data 101, where N channels of audio data 101 are input into a spatial analysis and downmix unit 102 and one LFE channel is input into an LFE channel encoding unit 105. The audio data 101 includes, but is not limited to, mono signals, stereo signals, binaural signals, spatial audio signals (e.g., multi-channel spatial audio objects), higher order Ambisonics (HOA), and any other audio data.
[0048] In some implementations, the spatial analysis and downmix unit 102 is configured to implement complex advanced coupling (CACPL) for analysis / downmixing of stereo audio data, and / or implement spatial reconstruction (SPAR) for analysis / downmixing of FoA audio data. In other implementations, the spatial analysis and downmix unit 102 implements other formats. The output of the spatial analysis and downmix unit 102 includes spatial metadata and 1 to N channels of audio data. The spatial metadata is input into the spatial metadata encoding unit 104, which is configured to quantize and entropy code the spatial metadata. In some implementations, the quantization can include fine, moderate, coarse, and extra coarse quantization strategies, and the entropy coding can include Huffman or arithmetic coding.
[0049] The 1 to N channels of audio data are input into the main audio channel encoding unit 103, which is configured to encode the 1 to N channels of audio data into one or more enhanced voice services (EVS) bitstreams. In some implementations, the main audio channel encoding unit 103 complies with 3GPP TS 26.445 and provides a variety of functionalities, such as enhancing quality and coding efficiency for narrowband (EVS-NB) and wideband (EVS-WB) speech services, enhancing quality using super wideband (EVS-SWB) speech, enhancing quality of mixed content and music in conversational applications, robustness to packet loss and delay jitter, and backward compatibility with AMR-WB codecs.
[0050] In some implementations, the main audio channel encoding unit 103 includes a pre-processing and mode selection unit that selects between a speech coder for encoding speech signals and a perceptual coder for encoding audio signals based on mode / bit rate control at a prescribed bit rate. In some implementations, the speech encoder is an improved variation of algebraic code excited linear prediction (ACELP) that extends specialized LP-based modes for different speech classes.
[0051] In some implementations, the audio encoder is a modified discrete cosine transform (MDCT) encoder that is improved in efficiency at low delay / low bit rate and designed to perform seamless and reliable switching between the speech encoder and the audio encoder.
[0052] As previously described, the LFE channel signal is used for deep bass tones ranging from 20 Hz to 120 Hz, and is typically sent to a speaker designed to reproduce low frequency audio content (e.g., a subwoofer). The LFE channel signal is input into the LFE channel signal encoding unit 105, which is configured to encode the LFE channel signal as described with reference to Figure 2A to the LFE channel signal.
[0053] In some embodiments, the IVAS decoder includes a spatial metadata decoding unit 106 configured to recover spatial metadata, and a main audio channel decoding unit 107 configured to recover 1 to N channel audio signals. The recovered spatial metadata and the recovered 1 to N channel audio signals are input into a spatial synthesis / upmix / render unit 109 configured to synthesize and render the 1 to N channel audio signals into N or more than N channel output audio signals using the spatial metadata for playback on speakers of various audio systems, including but not limited to: home theater systems, video conference room systems, virtual reality (VR) equipment, and any other audio system capable of rendering audio. Figure 2B
[0054] Although the example embodiments of LFE encoding / decoding described above are performed by an IVAS codec, the low-latency LFE codec described below can be a standalone LFE codec, or it can be included in any specialized or standardized audio codec that encodes and decodes low frequency signals in audio applications that require or desire low latency and configurability.
[0055] Figure 2A is a block diagram illustrating functional components of the LFE channel encoding unit 105 shown in Figure 1 is a block diagram illustrating functional components of the LFE channel decoding unit 108 shown in Figure 2B is a block diagram illustrating functional components of the LFE channel encoding unit 105 shown in Figure 1 is a block diagram illustrating functional components of the LFE channel decoding unit 108 shown in. The LFE channel decoder 108 includes an entropy decoding and inverse quantization unit 204, an inverse MDCT and windowing unit 205, a delay adjustment unit 206, and an output LPF 207. The delay adjustment unit 206 can be located before or after the LPF 207, and performs delay adjustment (e.g., by buffering the decoded LFE channel signal) to match the decoded LFE channel signal with the main codec decoded output. Hereinafter, the LFE channel encoding unit 105 and the LFE channel decoding unit 108 described with reference to Figure 2B are collectively referred to as an LFE codec.
[0056] The LFE channel encoding unit 105 includes an input low pass filter (LPF) 201, a windowing and MDCT unit 202, and a quantization and entropy coding unit 203. In an embodiment, the input audio signal is a pulse code modulated (PCM) audio signal, and the LFE channel encoding unit 105 expects an input audio signal with a 5 millisecond, 10 millisecond, or 20 millisecond stride. Intrinsically, the LFE channel encoding unit 105 operates on 5 millisecond or 10 millisecond subframes, and performs windowing and MDCT on a combination of these subframes. In an embodiment, the LFE channel encoding unit 105 operates with a 20 millisecond input stride and intrinsically divides this input into two subframes of equal length. The last subframe of the previous input frame to the LFE is concatenated with the first subframe of the current input frame to the LFE and windowed. The first subframe of the current input frame to the LFE is concatenated with the second subframe of the current input frame to the LFE and windowed. The MDCT is performed twice, once on each windowed block.
[0057] In an embodiment, the algorithmic delay (not including framing delay) is equal to 8 milliseconds plus the delay caused by the input LPF 103 plus the delay caused by the output LPF 207. In the case of a fourth order input LPF 201 and a fourth order output LPF 207, the total system latency is approximately 15 milliseconds. In the case of a fourth order input LPF 201 and a second order output LPF 207, the total LFE codec latency is approximately 13 milliseconds.
[0058] Figure 3 FIG. 2B is a graph illustrating the frequency response of an example input LPF 201 according to one or more embodiments. In the example shown, the LPF 201 is a fourth order Butterworth filter with a cutoff frequency of 130 Hz. Other embodiments can use different types of LPFs (e.g., Chebyshev, Bessel) with the same or different order and the same or different cutoff frequency.
[0059] Figure 4 FIG. 2C is a graph illustrating a Hanning window according to one or more embodiments. In an embodiment, the window function applied by the windowing and MDCT unit 202 is a Hanning window function with a fade length of 8 milliseconds. The Hanning window is a Kaiser-Bessel Derived (KBD) window with alpha = 5, which is constructed to satisfy the Princen-Bradley condition for the MDCT, and thus is a window used in the Advanced Audio Coding (AAC) digital audio format. Other window functions can also be used.
[0060] Quantization and entropy coding
[0061] In embodiments, quantization and entropy coding unit 203 implements a quantization strategy that conforms to the input LPF 201 frequency response curve to more efficiently quantize the MDCT coefficients. In embodiments, the frequency range is divided into 4 sub-band groups representing 4 frequency bands: 0 Hz to 100 Hz, 100 Hz to 200 Hz, 200 Hz to 300 Hz, and 300 Hz to 400 Hz. These frequency bands are examples, and more or less frequency bands can be used with the same or different frequency ranges. More specifically, the MDCT coefficients are quantized using scaling shift factors that are dynamically computed based on the MDCT coefficient values in a particular frame, and the quantization points are selected according to the LPF frequency response curve as shown in Figures 5 to 8
[0062] In embodiments, the following describes the quantization strategy for the F len len Any value given by 5*f ms can be taken, where 1<=f<=12.
[0063] First, the input PCM step size is divided into N sub-frames of equal length, each sub-frame width (S w ) = F len / N ms. N should be chosen such that each S w is a multiple of 5 ms (for example, if F len = 20 ms, then N can be 1, 2, or 4; if F len = 10 ms, then N can be 1 or 2; and if F len = 5 ms, then N equals 1). Let S i be the i-th sub-frame in any given frame, where i is an integer ranging from 0<=i<=N, with S0corresponding to the last sub-frame in the previous input frame to LFE encoding unit 105, and S1to SNare the N sub-frames in the current frame. N i i+1 Next, each S w and S mdct sub-frame is concatenated and windowed with a Hanning window (see Figure 4 ) are windowed, and then MDCT is performed on these windowed samples. This results in a total of N MDCTs per frame. The number of MDCT coefficients (num_coeffs) from each MDCT = sampling frequency * S w / 1000. The frequency resolution of each MDCT (width of each MDCT coefficient) (W mdct ) is approximately 1000 / (2*S w ) Hz. Given that subwoofers typically have an LPF cutoff of approximately 100 Hz to 120 Hz, and the post-LPF energy is typically very low after 400 Hz, the MDCT coefficients up to 400 Hz are quantized and sent to the LFE decoding unit 108, while the rest of the MDCT coefficients are quantized to 0. Sending the MDCT coefficients up to 400 Hz ensures high quality reconstruction up to 120 Hz at the LFE decoding unit 108. Thus, the total number of MDCT coefficients used for quantization and coding (N quant ) is equal to N*400 / W mdct .
[0065] Next, the MDCT coefficients are arranged in M sub-band groups, where the width of each sub-band group is a multiple of W mdct and the sum of the widths of all sub-band groups is equal to 400 Hz. The width of each sub-band is made to be SBW m Hz, where m is an integer ranging from 1 <= m <= M. At this width, the number of coefficients in the mthsub-band group = SN quant = N*SBW m / W mdct (i.e., SBW m / W mdct coefficients from each MDCT). Then, the MDCT coefficients in each sub-band group are scaled according to a shift scaling factor (shift) described below, which is determined by the sum or maximum of the absolute values of all N quant MDCT coefficients. The scaled MDCT coefficients in each sub-band group are then individually quantized and coded using a quantization scheme that conforms to the LPF curve at the encoder input. The quantized MDCT coefficients are coded with an entropy coder (e.g., arithmetic or Huffman coder). Different entropy coders are used to code each sub-band group, and each entropy coder uses an appropriate probability distribution model to efficiently code the respective sub-band group.
[0066] An example quantization strategy with a frame size (F len ) of 20 milliseconds (ms), 2 sub-frames (N = 2), and a sampling frequency = 48000 will now be described. Under this example input configuration, the sub-frame width S w= 10 ms and the number of MDCTs = N = 2. The first MDCT is performed on a 20 ms block. This block is formed by concatenating the 10 ms to 20 ms sub-frame in the previous 20 ms input with the 0 ms to 10 ms sub-frame in the current 20 ms input, and then windowing the 20 ms long Field window (see Figure 4 ). In the case of N = 1 and N = 4, the Field window is scaled accordingly and the fade length changes to 16 / N ms. The second MDCT is performed on a 20 ms block formed by windowing the current 20 ms input frame with a 20 ms long Field window. The number of MDCT coefficients (num_coeffs) per MDCT = 480, the width W mdct = 50 Hz, the total number of quantized and coded coefficients N quant = 16, and the total number of quantized and coded coefficients / MDCT = 16 / N = 8.
[0067] Next, the MDCT coefficients are arranged in 4 sub-band groups (M = 4), where each sub-band group corresponds to a 100 Hz frequency band (0 to 100, 100 to 200, 200 to 300, 300 to 400, SBW m = 100 Hz, the number of coefficients in each sub-band group = SN quant = N * SBW m / W mdct = 4). Let a1, a2, a3, a4, a5, a6, a7, a8 be the first 8 MDCT coefficients to be quantized from the first MDCT, and b1, b2, b3, b4, b5, b6, b7, b8 be the first 8 MDCT coefficients to be quantized from the second MDCT. The 4 sub-band groups are arranged to have the following coefficients:
[0068] Sub-band group 1 = {a1, a2, b1, b2},
[0069] Sub-band group 2 = {a3, a4, b3, b4},
[0070] Sub-band group 3 = {a5, a6, b5, b6},
[0071] Sub-band group 4 = {a7, a8, b7, b8},
[0072] where each sub-band group corresponds to a 100 Hz frequency band.
[0073] A frame with a gain of approximately -30 dB (or less than -30 dB) can have a value of greater than 10 -2 or 10 -1or lower, while a frame with full-scale gain can have MDCT coefficients with values of 20 or higher. To satisfy this wide range of values, a scaling shift factor (shift) is computed based on the sum of the absolute values of the available maximum quantization points (max_value) and the MDCT coefficients (lfe_dct_new), as follows:
[0074] shift = floor(shifts_per_double * log2(max_value / sum(abs(lfe_dct_new))).
[0075] In an implementation, lfe_dct_new is an array of 16 MDCT coefficients, shifts_per_double is a constant (e.g., 4), max_value is an integer selected for fine quantization (e.g., 63 quantization values) and coarse quantization (e.g., 31 quantization values), and the shift is limited to 5-bit values from 4 to 35 for fine quantization and 5-bit values from 2 to 33 for coarse quantization.
[0076] The quantized MDCT coefficients are then computed as follows:
[0077] vals = round(lfe_dct_new * (2^(shift / shifts_per_double))), where the round() operation rounds the result to the nearest integer value.
[0078] If the quantized values (vals) exceed the maximum allowed number of available quantization points (max_val), the scaling shift factor (shift) is decreased and the quantized values (vals) are computed again. In other implementations, instead of the sum function sum(abs(lfe_dct_new)), the maximum function max(abs(lfe_dct_new)) can be used to compute the scaling shift factor (shift), but using the max() function disperses the quantization values more, making it more difficult to design an efficient entropy coder.
[0079] In the quantization steps described above, the quantized values for each band group are computed together in one loop, but the quantization points for each band group are different. If the first band group exceeds the allowed range, the scaling shift factor is decreased. If any of the other band groups exceed the allowed range, the band group is truncated to max_value. The sign bits and the absolute values of the quantized MDCT coefficients are coded separately for each band group.
[0080] Figure 5FIG. illustrates the variation of fine quantization points with frequency, according to one or more embodiments. At fine quantization, sub-band group 1 (0 Hz to 100 Hz) has 64 quantization points, sub-band group 2 (100 Hz to 200 Hz) has 32 quantization points, sub-band group 3 (200 Hz to 300 Hz) has 8 quantization points, and sub-band group 4 (300 Hz to 400 Hz) has 2 quantization points. In an embodiment, each sub-band group is entropy coded with an entropy coder (e.g., arithmetic or Huffman entropy coder), where each entropy coder uses a different probability distribution. Thus, the primary 0 Hz to 100 Hz range is allocated the most quantization points.
[0081] Note that the allocation of quantization points for sub-band group 1 to sub-band group 4 follows the shape of the LPF frequency response curve, which has more information in the lower frequencies than in the higher frequencies, and no information outside the cutoff frequency. To properly reconstruct frequencies up to 130 Hz, the MDCT coefficients corresponding to frequencies above 130 Hz are also encoded to avoid or minimize aliasing. In some embodiments, the MDCT coefficients up to 400 Hz are encoded so that frequencies up to 130 Hz can be properly reconstructed at the decoding unit.
[0082] Figure 6 FIG. illustrates the variation of coarse quantization points with frequency, according to one or more embodiments. At coarse quantization, sub-band group 1 (0 Hz to 100 Hz) has 32 quantization points, sub-band group 2 (100 Hz to 200 Hz) has 16 quantization points, sub-band group 3 (200 Hz to 300 Hz) has 4 quantization points, and sub-band group 4 (300 Hz to 400 Hz) is not quantized and entropy coded. In an embodiment, each sub-band group is entropy coded with a separate entropy coder that uses a different probability distribution.
[0083] Figure 7 FIG. illustrates the probability distribution of quantized MDCT coefficients at fine quantization, according to one or more embodiments. The y-axis is the frequency of occurrence and the x-axis is the number of quantization points. Sgl is sub-band group 1 corresponding to quantized MDCT coefficients in the 0 Hz to 100 Hz band, Sg2 is sub-band group 2 corresponding to quantized MDCT coefficients in the 100 Hz to 200 Hz band. Sg3 is sub-band group 3 corresponding to quantized MDCT coefficients in the 200 Hz to 300 Hz band. Sg4 is sub-band group 4 corresponding to quantized MDCT coefficients in the 300 Hz to 400 Hz band.
[0084] Figure 8FIG. illustrates probability distribution of quantized MDCT coefficients at coarse quantization, according to one or more embodiments. The y-axis is the frequency of occurrence and the x-axis is the number of quantization points. Sgl is sub-band group 1 corresponding to quantized MDCT coefficients in the 0 Hz to 100 Hz frequency band, Sg2 is sub-band group 2 corresponding to quantized MDCT coefficients in the 100 Hz to 200 Hz frequency band. Sg3 is sub-band group 3 corresponding to quantized MDCT coefficients in the 200 Hz to 300 Hz frequency band. Sg4 is sub-band group 4 corresponding to quantized MDCT coefficients in the 300 Hz to 400 Hz frequency band.
[0085] Note that the primary frequency band (0 Hz to 100 Hz) is the band where the LFE effect is found most and thus is allocated more quantization points for greater resolution. However, the bits allocated to the primary frequency band in coarse quantization is less than in fine quantization. In an embodiment, whether a frame of MDCT coefficients is quantized using fine quantization or coarse quantization depends on the desired target bit rate set by the primary audio channel encoder 103. The primary audio channel encoder 103 sets this value once during initialization or dynamically on a frame by frame basis based on the bits needed or used to encode the primary audio channels in each frame.
[0086] Silence frame
[0087] In some embodiments, a signal is added in the LFE channel bitstream to indicate a silence frame. A silence frame is a frame with energy below a specified threshold. In some embodiments, a 1 bit is included in the LFE channel bitstream transmitted to the decoder (e.g., inserted in the frame header) to indicate a silence frame, and all MDCT coefficients in the LFE channel bitstream are set to 0. This technique can reduce the bit rate to 50 bps during a silence frame.
[0088] Decoder LPF
[0089] Two options are provided at the output of the LFE channel decoding unit 108 to implement the LPF 207 (see FIG. 2). Figure 2B The LPF 207 is selected based on the available delay (total delay of other audio channels minus LFE fade delay minus input LPF delay). Note that the other channels are expected to be encoded / decoded by the primary audio channel encoding unit 103 / primary audio channel decoding unit 107, and the delay of the channels depends on the algorithmic delay of the primary audio channel encoding unit 103 / primary audio channel decoding unit 107.
[0090] In implementations, if the available delay is less than 3.5 ms, a second order Butterworth LPF with a cutoff of 130 Hz is used; otherwise a fourth order Butterworth LPF with a cutoff of 130 Hz is used. Thus, at the LFE channel decoding unit 108, a tradeoff must be made between removing aliasing energy above the cutoff frequency and algorithmic delay. In some implementations, the LPF 207 can be removed entirely, as subwoofers typically have an LPF. The LPF 207 helps reduce aliasing energy above the cutoff of the LFE decoder output itself, and can help efficient post-processing.
[0091] Example process
[0092] Figure 9 is a flowchart of a process 900 of encoding MDCT coefficients according to one or more implementations. The process 900 can be implemented using, for example, the system 1100 described with reference to Figure 11 The process 900 can be implemented using, for example, the system 1100 described with reference to
[0093] The process 900 includes the following steps: receiving a time domain LFE channel signal (901); filtering the time domain LFE channel signal using a low pass filter (902); converting the filtered time domain LFE channel signal to a frequency domain representation of the LFE channel signal including a number of coefficients representing a spectrum of the LFE channel signal (903); arranging the coefficients into a number of subband groups corresponding to different frequency bands of the LFE channel signal (904); quantizing the coefficients in each subband group using a scaling shift factor according to a frequency response curve of the low pass filter (905); encoding the quantized coefficients in each subband group using an entropy coder configured for the subband group (906); generating a bitstream including the encoded quantized coefficients (907); and storing the bitstream on a storage device or streaming the bitstream to a downstream device (908).
[0094] Figure 10 is a flowchart of a process 1000 of decoding MDCT coefficients according to one or more implementations. The process 1000 can be implemented using, for example, the system 1100 described with reference to Figure 11 The process 1000 can be implemented using, for example, the system 1100 described with reference to
[0095] The process 1000 includes the following steps: receiving an LFE channel bitstream (1001), wherein the LFE channel bitstream includes entropy coded coefficients representing a spectrum of a time domain LFE channel signal; decoding and inverse quantizing the coefficients (1002), wherein the coefficients are quantized in subband groups corresponding to different frequency bands according to a frequency response curve of a low pass filter using a scaling shift factor; converting the decoded and inverse quantized coefficients to a time domain LFE channel signal (1003); adjusting a delay of the time domain LFE channel signal (1004); and filtering the delay-adjusted LFE channel signal using a low pass filter (1005). In an embodiment, an order of the low pass filter can be configured based on a total algorithmic delay that can be derived from a primary codec that can be used to encode / decode full bandwidth channels of a multi-channel audio signal including the time domain LFE channel signal. In some implementations, the decoding unit only needs to know whether the encoding unit encoded the MDCT coefficients with fine quantization or coarse quantization. The quantization type can be indicated using bits in an LFE bitstream header or any other suitable signaling mechanism.
[0096] In some implementations, the decoding of the inverse quantized coefficients to time domain PCM samples is performed in the following way. The inverse quantized coefficients in each subband group are rearranged into N groups (N is the number of MDCTs operated at the encoding unit), where each group has coefficients corresponding to a respective MDCT. According to the example implementation described above, the encoding unit encodes the following 4 subband groups:
[0097] Subband group 1 = {a1, a2, b1, b2},
[0098] Subband group 2 = {a3, a4, b3, b4},
[0099] Subband group 3 = {a5, a6, b5, b6},
[0100] Subband group 4 = {a7, a8, b7, b8}.
[0101] The decoding unit decodes and rearranges the 4 subband groups back to {a1, a2, a3, a4, a5, a6, a7, a8} and {b1, b2, b3, b4, b5, b6, b7, b8}, and then pads zeros to the groups to get the desired inverse MDCT (iMDCT) input length. N iMDCTs are performed to inverse transform the MDCT coefficients in each group to time domain blocks. In this example, each block is 2*Sw ms wide, where Sw is the subframe width defined above. Next, the decoded LFE channel signal is obtained by Figure 4The same window of the Field window used by the LFE encoding unit shown in the middle is used to window this block. Each sub-frame S is reconstructed by properly superimposing the windowed data of the previous iMDCT output with the current iMDCT output i (i is an integer between 1 <= i <= N). Finally, the output of (1003) is reconstructed by concatenating all N sub-frames.
[0102] Example System Architecture
[0103] Figure 11 is used to implement a reference Figures 1 to 10 A block diagram of a system 1100 described features and processes in accordance with one or more implementations. The system 1100 includes one or more server computers or any client devices, including but not limited to: a calling server, a user equipment, a conference room system, a home theater system, virtual reality (VR) equipment, and immersive content ingestion devices. The system 1100 includes any consumer device, including but not limited to: a smartphone, a tablet computer, a wearable computer, a vehicle computer, a game console, a surround sound system, a kiosk, etc.
[0104] As shown, the system 1100 includes a central processing unit (CPU) 1101, which is capable of executing various processes according to a program stored in, for example, a read only memory (ROM) 1102 or a program loaded from, for example, a storage unit 1108 to a random access memory (RAM) 1103. Data required when the CPU 1101 executes various processes is also stored in the RAM 1103 as necessary. The CPU 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0105] The following components are connected to the I / O interface 1105: an input unit 1106, which can include a keyboard, a mouse, etc.; an output unit 1107, which can include a display such as a liquid crystal display (LCD) and one or more speakers; a storage unit 1108, which includes a hard disk or another suitable storage device; and a communication unit 1109, which includes a network interface card such as a network card (e.g., wired or wireless).
[0106] In some implementations, the input unit 1106 includes one or more microphones in different locations (depending on the host device) and capable of capturing audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0107] In some implementations, the output unit 1107 includes a system with various numbers of speakers. The output unit 1107 can reproduce audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats), depending on the capabilities of the host device.
[0108] The communication unit 1109 is configured to communicate with other devices (e.g., via a network). The driver 1110 is also connected to the I / O interface 1105 as necessary. A removable media 1111 (e.g., a disk, an optical disk, a magneto-optical disk, a flash drive, or another suitable removable media) is installed on the driver 1110 so that a computer program read from the removable media 1111 is installed in the storage unit 1108 as necessary. It will be understood by those skilled in the art that although the system 1100 is described as including the components described above, in actual applications, some of these components can be added, removed, and / or replaced, and all such modifications or changes are within the scope of the present application.
[0109] According to example embodiments of the present application, the processes described above can be implemented as computer software programs or embodied on computer-readable storage media. For example, embodiments of the present application include a computer program product having a computer program tangibly embodied on a machine-readable medium, the computer program including program code for executing methods. In such embodiments, the computer program can be downloaded from a network via the communication unit 1309 and installed, and / or installed from the removable media 1111.
[0110] In general, the various example embodiments of the present application can be implemented as hardware or special-purpose circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the units discussed above can be executed by a CPU in combination with other components of the Figure 11 The control circuitry can thus perform the actions described in the present application. Some aspects can be implemented as hardware, while other aspects can be implemented as firmware or software executable by a controller, microprocessor, or other computing device (e.g., control circuitry). Although various aspects of example embodiments of the present application can be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein can be implemented in, as non-limiting examples, hardware, software, firmware, special-purpose circuitry or logic, general purpose hardware or controller, or other computing devices or some combination thereof.
[0111] Also, various blocks in the flowchart diagrams can be viewed as methods, and / or as operations implemented by a computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s), and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present application include computer program products comprising computer program code embodied on a tangible medium, the computer program code comprising instructions configured to carry out the methods described above.
[0112] In the context of the present application, a machine / computer readable medium can be any tangible medium that can contain or store program(s) for use by or in connection with an instruction execution system, apparatus, or device. The machine / computer readable medium can be a machine / computer readable signal medium or a machine / computer readable storage medium. The machine / computer readable medium can be non-transitory and can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine / computer readable storage medium will include one or more of an electrical connection having one or more wires, a portable computer diskette, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0113] Computer program code to carry out operations of the present application can be written in any combination of one or more programming languages. These computer program codes might be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program code, when executed by the processor of the computer or other programmable data processing apparatus, enables the computer or other programmable data processing apparatus to implement the functions / acts specified in the flowchart diagrams and / or block diagrams. The computer program code can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operations to be performed on the computer or other programmable data processing apparatus to produce a computer implemented process such that the computer program code which implements the process can be executed on the computer or other programmable data processing apparatus to cause a series of operations to be performed on the computer or other programmable data processing apparatus to produce a computer implemented process.
[0114] While this document contains many specific implementation details, these should not be construed as limiting the scope of what can be claimed, but as merely providing description of particular implementations. Particular features described herein can be combined in any suitable manner in individual implementations. Conversely, various features described herein in the context of separate implementations can also be implemented in combination. Additionally, although individual features can be described above as being implemented in particular combinations, one or more features from a combination can in some cases be left out of the combination, and the application can be directed to a sub-combination or variation of a sub-combination. The logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps can be provided, or steps can be eliminated, from the described flows, and other components can be added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.
Claims
1. A method for encoding low-frequency effect (LFE) channels, comprising: Use one or more processors to receive time-domain LFE channel signals; The time-domain LFE channel signal is filtered using a low-pass filter to produce a filtered time-domain LFE channel signal, wherein the low-pass filter has a cutoff frequency. The one or more processors are used to convert the filtered time-domain LFE channel signal into a frequency domain representation of the time-domain LFE channel signal, which includes a certain number of coefficients representing the spectrum of the time-domain LFE channel signal. The coefficients are arranged into two or more subband groups corresponding to different frequency bands of the time-domain LFE channel signal using one or more processors, wherein the different frequency bands include a main LFE band below the cutoff frequency of the LFE loudspeaker and at least one other LFE band above the cutoff frequency of the LFE loudspeaker, wherein each subband group has a width, and the sum of the widths of the subband groups includes the main LFE band and the at least one other LFE band; The one or more processors are used to quantize the coefficients in each frequency band group according to the frequency response curve of the low-pass filter to generate quantized coefficients. The quantized coefficients in the sub-band group are encoded using one or more processors with an entropy decoder for each band group tuning; and The one or more processors are used to generate a bitstream containing the encoded quantized coefficients; and The bitstream is stored on a storage device or streamed to a downstream device using one or more processors.
2. A low-latency, low-frequency-effect (LFE) decoder, comprising: One or more processors; as well as A non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the operation of claim 1 of the method.
3. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operation of claim 1.
Citation Information
Patent Citations
Low bit-rate audio encoding
US20070112560A1
Implementing a High Quality VOIP Device
US20110103377A1
Method and apparatus for encoding and decoding audio signal using layered sinusoidal pulse coding
US20120095754A1
Audio processing system
US20160055855A1
Acoustic signal transform coding method and decoding method having a high efficiency envelope flattening method therein
US5684920A