METHOD FOR ENCODING A LOW-FREQUENCY EFFECTS CHANNEL, LOW-LATENCY LOW-FREQUENCY EFFECTS ENCODER, METHOD FOR DECODING A LOW-FREQUENCY EFFECTS CHANNEL, AND LOW-LATENCY LOW-FREQUENCY EFFECTS DECODER

AR125511B2Active Publication Date: 2026-08-28DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
ARP20220100527
Authority / Receiving Office
AR · AR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-24
Filing Date
2022-03-08
Publication Date
2026-08-28
Estimated Expiration
2040-09-03

AI Technical Summary

Technical Problem

Existing audio codecs struggle to efficiently process low frequency effects (LFE) channels with low latency and high quality, particularly in immersive audio services, lacking a codec that can handle deep bass sounds ranging from 20-120 Hz with low latency and suitable bit rates.

Method used

A low latency LFE codec that filters and converts LFE channel signals into frequency-domain representations, quantizes coefficients based on a low-pass filter frequency response curve, and encodes them using entropy encoding, with configurable quantization schemes and separate encoding of sign bits, ensuring low latency and efficient bit rate management.

Benefits of technology

The codec achieves low latency of 20-33 milliseconds, supports bit rates from 2 kbps to 4 kbps during active frames and 50 bps during silence, maintaining high-quality reconstruction of LFE signals up to 120 Hz.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

In some implementations, a method for encoding a low-frequency effects (LFE) channel comprises: receiving an LFE channel signal in the time domain; filtering the LFE channel signal in the time domain using a low-pass filter; converting the filtered LFE channel signal in the time domain into a frequency-domain representation of the LFE channel signal that includes a number of coefficients representing a frequency spectrum of the LFE channel signal; arranging coefficients into a number of subband groups corresponding to different frequency bands of the LFE channel signal; quantizing coefficients in each subband group according to a frequency response curve of the low-pass filter; encoding the quantized coefficients in each subband group by using an entropy encoder tuned to the subband group;and generate a bitstream that includes the encoded quantized coefficients; and store the bitstream on a storage device or continuously transmit the bitstream to a downstream device.
Need to check novelty before this filing date? Find Prior Art

Description

LOW-LATENCY LOW-FREQUENCY EFFECTS CODEC TECHNICAL FIELD

[0001] This disclosure relates in general to the processing of audio signals and, in particular, to the processing of low-frequency effects (LFE) channels. BACKGROUND

[0002] Standardization efforts for immersive services include the development of an Immersive Voice and Audio Service (IVAS) codec for voice, multistream teleconferencing, virtual reality (VR), and streaming of live and on-demand user-generated content, for example. One objective of the IVAS standard is to develop a single codec with excellent audio quality, low latency, support for spatial audio coding, an appropriate range of bit rates, high fault resilience, and low practical implementation complexity. To achieve this objective, the goal is to develop an IVAS codec capable of handling low-latency LFE operations on IVAS-enabled devices or any other device capable of processing LFE signals.The LFE channel is intended for deep, bass sounds ranging from 20-120 Hz and is usually sent to a speaker designed to reproduce low-frequency audio content. SYNTHESIS

[0003] Implementations for a configurable low-latency LFE codec are revealed.

[0004] In some implementations, a method for encoding a low-frequency effects (LFE) channel comprises: receiving, using one or more processors, an LFE channel signal in the time domain; filtering, using a low-pass filter, the LFE channel signal in the time domain; converting, using one or more processors, the filtered LFE channel signal in the time domain into a frequency-domain representation of the LFE channel signal that includes a number of coefficients representing a frequency spectrum of the LFE channel signal; arranging, using one or more processors, coefficients into a number of subband groups corresponding to different frequency bands of the LFE channel signal; quantizing, using one or more processors, the coefficients into a number of subband groups corresponding to different frequency bands of the LFE channel signal; and quantizing, using one or more processors, the LFE channel signal into a number of subband groups. 238390 1689964 of 24 processors, coefficients in each subband group according to a frequency response curve of the low-pass filter; encoding, using the one or more processors, the quantized coefficients in each subband group by using an entropy encoder tuned for the subband group; and generating, using the one or more processors, a bitstream that includes the encoded quantized coefficients; and storing, using the one or more processors, the bitstream in a storage device or streaming the bitstream to a downstream device.

[0005] In some implementations, quantizing the coefficients in each subband group further comprises: generating a scaling factor based on a maximum number of available quantization points and a sum of the absolute values ​​of the coefficients; and quantizing the coefficients using the scaling factor.

[0006] In some implementations, if a quantized coefficient exceeds the maximum number of quantization points, the scaling factor is reduced and the coefficients are quantized again.

[0007] In some implementations, the quantization points are different for each subband group.

[0008] In some implementations, the coefficients in each subband group are quantized according to either a fine quantization scheme or a coarse quantization scheme, where the fine quantization scheme assigns more quantization points to one or more subband groups than are assigned to the respective subband groups according to the coarse quantization scheme.

[0009] In some implementations, the sign bits for the coefficients are encoded separately from the coefficients.

[0010] In some implementations, there are four subband groups, and a first subband group corresponds to a first frequency range of 0-100 Hz, a second subband group corresponds to a second frequency range of 100-200 Hz, a third subband group corresponds to a third frequency range of 200-300 Hz, and a fourth subband group corresponds to a fourth frequency range of 300-400 Hz.

[0011] In some implementations, the entropy encoder is either an arithmetic entropy encoder or a Huffman entropy encoder.

[0012] In some implementations, converting the filtered time-domain LFE channel signal into a frequency-domain representation of the LFE channel signal that includes a number of coefficients representing a spectrum of 238390 1689964 of 24 frequency of the LFE channel signal, further comprises: determining a first stride length of the LFE channel signal; designating a first window size of a window generation function based on the first stride length; applying the first window size to one or more frames of the LFE channel signal in the time domain; and applying a modified discrete cosine transform (MDCT) to the windowed frames to generate the coefficients.

[0013] In some implementations, the method further comprises: determining a second step length of the LFE channel signal; designating a second window size of the windowing function based on the second step length; and applying the second window size to one or more frames of the LFE channel signal in the time domain.

[0014] In some implementations, the first step length is N milliseconds (ms), N is greater than or equal to 5 ms and less than or equal to 60 ms, the first window size is greater than or equal to 10 ms, the second step length is 5 ms and the second window size is 10 ms.

[0015] In some implementations, the first step length is 20 milliseconds (ms), the first window size is 10 ms or 20 ms or 40 ms, the second step length is 10 ms and the second window size is 10 ms or 20 ms.

[0016] In some implementations, the first step length is 10 milliseconds (ms), the first window size is 10 ms or 20 ms, the second step length is 5 ms, and the second window size is 10 ms.

[0017] In some implementations, the first step length is 20 milliseconds (ms), the first window size is 10 ms, 20 ms or 40 ms, the second step length is 5 ms and the second window size is 10 ms.

[0018] In some implementations, the window generation function is a Kaiser-Bessel Derivative (KBD) function with a configurable fade-out length.

[0019] In some implementations, the low-pass filter is a low-pass filter Fourth-order Butterworth with a cutoff frequency of approximately 130 Hz or less.

[0020] In some implementations, the method further comprises: determining, using one or more processors, whether an energy level of an LFE channel signal frame is below a threshold; according to the energy level that is below 238390 1689964 of 24 a threshold level, generate a silence frame indicator that signals this to the decoder; insert the silence frame indicator into the LFE channel bitstream metadata; and reduce the LFE channel bit rate when the silence frame is detected.

[0021] In some implementations, a method for decoding a low-frequency effect (LFE) comprises: receiving, using one or more processors, an LFE channel bitstream, the LFE channel bitstream including entropically encoded coefficients representing a frequency spectrum of an LFE channel signal in the time domain; decoding, using the one or more processors, the quantized coefficients using an entropy decoder; inversely quantizing, using the one or more processors, the inversely quantized coefficients, wherein the coefficients were quantized into subband groups corresponding to frequency bands according to a frequency response curve of a low-pass filter used to filter the LFE channel signal in the time domain in an encoder;convert, using one or more processors, the inversely quantized coefficients to an LFE channel signal in the time domain; adjust, using one or more processors, a delay to the LFE channel signal in the time domain; and filter, using a low-pass filter, the LFE channel signal with the adjusted delay.

[0022] In some implementations, a low-pass filter order is configured to ensure that a first total algorithmic delay due to encoding and decoding the LFE channel is less than or equal to a second total algorithmic delay due to encoding and decoding other audio channels of a multichannel audio signal that includes the LFE channel signal.

[0023] In some implementations, the method further comprises: determining whether the second total algorithmic delay exceeds a threshold value; and according to the second total algorithmic delay that exceeds the threshold value, setting the low-pass filter as an N-order low-pass filter, where N is an integer greater than or equal to two; and according to the second total algorithmic delay that does not exceed the threshold value, setting the order of the low-pass filter to be less than N.

[0024] Other implementations disclosed herein are directed to a computer-readable system, apparatus, and medium. Details of the disclosed implementation are set forth in the accompanying drawings and in the description below. Other 238390 1689964 of 24 features, objectives and advantages are evident from the description, drawings and claims.

[0025] Particular embodiments disclosed herein provide one or more of the following advantages. The disclosed low-latency LFE codec: 1) primarily targets the LFE channel; 2) primarily targets a frequency range of 20 to 120 Hz, but carries audio at 300 Hz in low / medium bitrate scenarios and 400 Hz in high bitrate scenarios; 3) achieves a low bitrate by applying a quantization scheme according to a frequency response curve to an input low-pass filter; 4) has low algorithmic latency and is designed to operate at a 20-millisecond (ms) step and have a total algorithmic latency (including frame generation) of 33 milliseconds;5) can be configured for smaller steps and lower algorithmic latency to support other scenarios, including lower step settings of 5 milliseconds and total algorithmic latency (including frame generation) of 13 milliseconds; 6) automatically chooses a low-pass filter at the decoder output based on the available latency with the LFE codec; 7) has a silent mode with a low bit rate of 50 bits per second (bps) during silence; and 8) during active frames the bit rate varies between 2 kilobits per second (kbps) and 4 kbps based on the quantization level used, and during silent frames the bit rate is 50 bps. DESCRIPTION OF THE DRAWINGS

[0026] To facilitate description, drawings show specific arrangements or orderings of schematic elements, such as those representing devices, units, instruction blocks, and data elements. However, those skilled in the art will understand that the specific arrangement or ordering of schematic elements in the drawings does not imply that a particular order or sequence of processing or separation of processes is required. Furthermore, the inclusion of a schematic element in a drawing does not imply that such element is required in all embodiments or that the features represented by such element cannot be included in or combined with other elements in some implementations. 238390 1689964 of 24

[0027] Furthermore, in drawings, when connecting elements, such as solid or dashed lines or arrows, are used to illustrate a connection, relationship, or association between two or more other schematic elements, the absence of any such connecting element does not imply that a connection, relationship, or association cannot exist. In other words, some connections, relationships, or associations between elements are not shown in drawings to avoid obscuring the clarity of the disclosure. Additionally, for ease of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, when a connecting element represents the communication of signals, data, or instructions, those skilled in the art will understand that such an element represents one or multiple paths of the signal, as necessary, to effect the communication.

[0028] FIG. 1 illustrates an IVAS codec for encoding and decoding IVAS and LFE bitstreams, according to one or more implementations.

[0029] FIG. 2A is a block diagram illustrating LFE encoding, according to one or more implementations.

[0030] FIG. 2B is a block diagram illustrating LFE decoding, according to one or more implementations.

[0031] FIG. 3 is a graph illustrating a frequency response of a fourth-order Butterworth low-pass filter with a corner or cutoff of 130 Hz, according to one or more implementations.

[0032] FIG. 4 is a graphic illustrating a Fielder window, according to one or more implementations.

[0033] FIG. 5 illustrates the variation of fine quantization points with frequency, according to one or more implementations.

[0034] FIG. 6 illustrates the variation of coarse quantization points with frequency, according to one or more implementations.

[0035] FIG. 7 illustrates a probability distribution of finely quantized MDCT coefficients, according to one or more implementations.

[0036] FIG. 8 illustrates a probability distribution of coarsely quantized MDCT coefficients, according to one or more implementations.

[0037] FIG. 9 is a flowchart of a modified discrete cosine transform (MDCT) coefficient encoding process, according to one or more implementations. 238390 1689964 of 24

[0038] FIG. 10 is a flowchart of a modified discrete cosine transform (MDCT) coefficient decoding process, according to one or more implementations.

[0039] FIG. 11 is a block diagram of a system for implementing the features and processes described with reference to FIGS. 1-10, according to one or more implementations.

[0040] The same reference symbol used in various drawings indicates similar elements. DETAILED DESCRIPTION

[0041] The following detailed description sets forth numerous specific details to provide a thorough understanding of the various embodiments described. It will be evident to those skilled in the art that the various implementations described can be carried out without these specific details. In other instances, known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily detract from the clarity of aspects of the embodiments. From here on, several features are described that can be used either independently of one another or in any combination with other features. Nomenclature

[0042] As used herein, the term “includes” and its variants should be read as open terms meaning “includes, without limitation.” The term “or” should be read as “and” unless the context clearly indicates otherwise. The term “based on” should be read as “based at least in part on.” The term “an exemplary implementation” should be read as “at least one exemplary implementation.” The term “another implementation” should be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” should be read as obtain, receive, compute, calculate, estimate, predict, or derive. Furthermore, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meanings as are commonly understood by a person skilled in the art to which this disclosure pertains. System overview 238390 1689964 of 24

[0043] FIG. 1 illustrates an IVAS 100 codec for encoding and decoding IVAS bitstreams, including an LFE channel bitstream, according to one or more implementations. For encoding, the IVAS 100 codec receives N+1 channels of audio data 101, where N channels of audio data 101 are input into the spatial analysis and downmixing unit 102 and one LFE channel is input into the LFE channel encoding unit 105. The audio data 101 includes, but is not limited to: mono signals, stereo signals, binaural signals, spatial audio signals (e.g., multichannel spatial audio objects), first-order Ambisonics (FoA), higher-order Ambisonics (HoA), and any other audio data.

[0044] In some implementations, the spatial analysis and downmixing unit 102 is configured to implement complex advanced coupling (CACPL) for downmixing stereo audio data and / or spatial reconstruction (SPAR) for downmixing FoA audio data. In other implementations, the spatial analysis and downmixing unit 102 implements other formats. The output of the spatial analysis and downmixing unit 102 includes spatial metadata and 1 to N channels of audio data. The spatial metadata is input to the spatial metadata encoding unit 104, which is configured to quantize and encode the spatial metadata.In some implementations, quantization may include fine, moderate, coarse, and extra coarse quantization strategies, and entropic coding may include Huffman coding or arithmetic.

[0045] Audio data channels 1 through N are fed into the main audio channel encoding unit 103, which is configured to encode the audio data channels 1 through N into one or more Enhanced Voice Services (EVS) bitstreams. In some implementations, the main audio channel encoding unit 103 complies with 3GPP TS 26.445 and provides a wide range of features, including improved encoding quality and efficiency for narrowband (EVS-NB) and wideband (EVS-WB) voice services, improved quality using superwideband voice (EVS-SWB), improved quality for music and mixed content in conversational applications, robustness against packet loss, and delay variation. 238390 1689964 of 24 backward compatibility for the AMR-WB codec (Adaptive Broadband Multi-Rate).

[0046] In some implementations, the main audio channel coding unit 103 includes a mode selection and preprocessing unit that selects between a speech encoder for encoding speech signals and a perceptual encoder for encoding audio signals at a specified bit rate based on bit rate / mode control. In some implementations, the speech encoder is an enhanced variant of algebraically code-excited linear prediction (ACELP), extended with specialized LP-based modes for different speech classes.

[0047] In some implementations, the audio encoder is a modified discrete cosine transform (MDCT) encoder with increased efficiency at low delay / low bit transfer rates and is designed to perform seamless and reliable switching between the voice and audio encoders.

[0048] As described above, the LFE channel signal is intended for bass and deep sounds ranging from 20 to 120 Hz and is typically sent to a loudspeaker designed to reproduce low-frequency audio content (e.g., a very low-frequency loudspeaker). The LFE channel signal is input to the LFE channel signal encoding unit 105, which is configured to encode the LFE channel signal as described with reference to FIG. 2A.

[0049] In some implementations, an IVAS decoder includes the spatial metadata decoding unit 106, which is configured to retrieve the spatial metadata, and the main audio channel decoding unit 107, which is configured to retrieve the audio signals from channels 1 to N. The retrieved spatial metadata and the retrieved audio signals from channels 1 to N are fed into the spatial synthesis / upmixing / rendering unit 109, which is configured to synthesize and render the audio signals from channels 1 to N into output audio signals of N or more channels (e.g., N+1) using the spatial metadata to be played back through speakers of various audio systems, including, without limitation: home theater systems, videoconferencing room systems, virtual reality (VR) equipment, and any other audio system capable of rendering audio.The LFE 108 channel decoding unit receives the LFE bit stream and is. 238390 1689964 of 24 configured to decode the LFE bit stream, as described with reference to FIG. 2B.

[0050] Despite the exemplary implementation of encoding / decoding of The LFE described above is performed by an IVAS codec. The low-latency LFE codec described below can be a standalone LFE codec, or it can be included in any proprietary or standardized audio codec that encodes and decodes low-frequency signals in audio applications where low latency and configurability are required or desired.

[0051] FIG. 2A is a block diagram illustrating the functional components of the LFE 105 channel encoding unit shown in FIG. 1, according to one or more embodiments. FIG. 2B is a block diagram illustrating the functional components of the LFE 108 channel decoder shown in FIG. 1, according to one or more embodiments. The LFE channel decoder 108 includes the inverse quantization and entropic decoding unit 204, the inverse MDCT and windowing unit 205, the delay adjustment unit 206, and the output low-pass filter 207. The delay adjustment unit 206 can be before or after the low-pass filter 207 and performs delay adjustment (e.g., by buffering the decoded LFE channel signal) to match the decoded LFE channel signal and the decoded output of the main codec.From here on, the LFE 105 channel encoding unit and the LFE 108 channel decoding unit described in relation to FIG. 2B are collectively referred to as an LFE codec.

[0052] The LFE 105 channel encoding unit includes an input low-pass filter (LPF) 201, a windowing and MDCT unit 202, and a quantization and entropy encoding unit 203. In one embodiment, the input audio signal is a pulse-code modulated (PCM) audio signal, and the LFE 105 channel encoding unit expects an input audio signal with a step of either 5 milliseconds, 10 milliseconds, or 20 milliseconds. Internally, the LFE 105 channel encoding unit operates on 5-millisecond or 10-millisecond subframes, and windowing and MDCT are performed on a combination of these subframes. In one embodiment, the LFE 105 channel encoding unit operates with a 20-millisecond input step and internally divides this input into two subframes of equal length. The last subframe of the pre-LFE input frame is 238390 1689964 of 24 is concatenated with the first subframe of the current LFE input frame, and windows are generated. The first subframe of the current LFE input frame is concatenated with the second subframe of the current LFE input frame, and windows are generated. The MDCT is performed twice, once in each block enclosed in windows.

[0053] In one embodiment, the algorithmic delay (without frame delay) is equal to 8 milliseconds plus the delay incurred by the input LPF 103 plus the delay incurred by the output LPF 207. With a fourth-order input LPF 201 and a fourth-order output LPF 207, the total system latency is approximately 15 milliseconds. With a fourth-order input LPF 201 and a second-order output LPF 207, the total LFE codec latency is approximately 13 milliseconds.

[0054] FIG. 3 is a graph illustrating the frequency response of an exemplary input LPF 201, according to one or more embodiments. In the example shown, LPF 201 is a fourth-order Butterworth filter with a cutoff frequency of 130 Hz. Other embodiments may use a different type of LPF (e.g., Chebyshev, Bessel) with the same or different order and the same or different cutoff frequency.

[0055] Figure 4 is a graphic illustrating a Fielder window, according to one or more embodiments. In one embodiment, the windowing function applied by the windowing unit and MDCT 202 is a Fielder window function with a fade length of 8 milliseconds. The Fielder window is a Kaiser-Bessel (KBD) derived window with alpha=5, which is a window that by construction satisfies the Princen-Bradley condition for MDCT, and is thus used with the Advanced Audio Coding (AAC) digital audio format. Other windowing functions may also be used. Quantification and entropic coding

[0056] In one embodiment, the entropic quantization and encoding unit 203 implements a quantization strategy that follows the frequency response curve of the input LPF 201 to quantize the MDCT coefficients more effectively. In one embodiment, the frequency range is divided into four subband groups representing four frequency bands: 0–100 Hz, 100–200 Hz, 200–300 Hz, and 300–400 Hz. These bands are examples, and more or fewer bands with the same or different frequency ranges may be used. In particular, the MDCT coefficients are quantized using a scaling factor that is dynamically calculated based on the coefficient values ​​of 238390 1689964 of 24 MDCT in a particular frame and the quantization points are selected according to the frequency response curve of the LPF, as shown in FIGS. 5-8. This quantization strategy helps to reduce the quantization points for MDCT coefficients belonging to the 100-200 Hz, 200-300 Hz and 300-400 Hz bands, while maintaining optimal quantization points for the main LFE band of 0-100 Hz, which is where the energy of most low-frequency effects (e.g., background noise) will be found.

[0057] In one embodiment, a quantization strategy for an input PCM step of Fien milliseconds (ms) (input frame length) to the LFE 105 channel coding unit is described below, wherein the frame length, Fien, can take any value given by 5*f ms, herein 1<=f<=12.

[0058] First, the input PCM step is divided into N subframes of equal lengths, each subframe having a width of (Sw) = Fien / N ms. N must be selected so that each Sw is a multiple of 5 ms (for example, if Fien = 20 ms then N can be 1, 2, or 4; if Fien = 10 ms then N can be 1 or 2; and if Fien = 5 ms then N is equal to 1). Suppose that Si is the i° subframe in any given frame, here i is an integer with range 0 <= i <= N, where S0 corresponds to the last subframe of the input frame prior to the LFE 105 encoding unit and S1 to Sn are the N subframes of the current frame.

[0059] Next, each subframe Si and Si+1 is concatenated and windowed using a Fielder window (see FIG. 4), and then MDCT is performed on these windowed samples. This results in a total of N MDCTs for each frame. The number of MDCT coefficients from each MDCT (num_coeffs) = sampling frequency * Sw / 1000. The frequency resolution of each MDCT (width of each MDCT coefficient) (Wmdct) is approximately 1000 / (2*Sw) Hz. Since very low frequency speakers typically have a low-frequency cutoff (LPF) of around 100-120 Hz, and the post-LPF energy after 400 Hz is generally very low, the MDCT coefficients up to 400 Hz are quantized and sent to the LFE 108 decoding unit, while the remaining MDCT coefficients are quantized to 0. Sending the MDCT coefficients up to 400 Hz ensures high-quality reconstruction up to 120 Hz in the LFE 108 decoding unit.The total number of MDCT coefficients for quantifying and coding (Nquant) is therefore equal to N*400 / Wmdct.

[0060] Next, the MDCT coefficients are arranged into M subband groups where the width of each subband group is a multiple of Wmdct and the sum of the widths of 238390 1689964 of 24 all subband groups equals 400 Hz. Let's assume that the width of each subband is SBWm Hz, where m is an integer with range 1 <= m <= M. With this width, the number of coefficients in the mth subband group = SNquant = N * SBWm / Wmdct (i.e., SBWm / Wmdct coefficients of each MDCT). The MDCT coefficients in each subband group are then scaled by a shift factor, described below, determined by the sum or maximum of the absolute values ​​of all the MDCT coefficients Nquant. The scaled MDCT coefficients in each subband group are then quantized and encoded separately using a quantization scheme that follows the LPF curve at the encoder input. The encoding of the quantized MDCT coefficients is done with an entropy encoder (e.g., an arithmetic or Huffman encoder).Each subband group is encoded with a distinct entropy encoder, and each entropy encoder uses an appropriate probability distribution model to encode the respective subband group efficiently.

[0061] An exemplary quantization strategy for a 20-millisecond (ms) step (Fien = 20 ms), 2 subframes (N = 2), and a sampling rate of 48000 will be described below. With this exemplary input configuration, the subframe width Sw = 10 ms and the number of MDCTs = N = 2. The first MDCT is performed over a 20 ms block. This block is formed by concatenating a 10–20 ms subframe from the previous 20 ms input and a 0–10 ms subframe from the current 20 ms input, and then windowing with the 20 ms long Fielder window (see FIG. 4). With N = 1 and N = 4, the Fielder window is scaled accordingly, and the fade length is changed to 16 / N ms. The second MDCT is performed on a 20 ms block formed by including the current 20 ms input frame in windows with a 20 ms long Fielder window.The number of MDCT coefficients (num_coeffs) with each MDCT = 480, the width of each MDCT coefficient Wmdct = 50 Hz, the total number of coefficients to quantize and encode Nquant = 16 and the total number of coefficients to quantize and encode according to MDCT = 16 / N = 8.

[0062] Next, the MDCT coefficients are arranged into 4 subband groups (M=4), where each subband group corresponds to a 100 Hz band (0-100, 100-200, 200-300, 300-400, SBWm = 100 Hz, number of coefficients in each subband group = SNquant = N*SBWm / Wmdct = 4). Let's assume that a1, a2, as, a4, as, a6, a?, as are the first 8 MDCT coefficients to be quantized starting from the first MDCT and b1, b2, bs, b4, 238390 1689964 of 24 bs, bó, b?, bs are the first 8 MDCT coefficients that must be quantified from the second MDCT. The 4 subband groups are arranged to have the following coefficients: subband group 1 = {ai, a2, bi, b2}, subband group 2 = {as, a4, bs, b4}, subband group S = {as, a6, bs, bó}, subband group 4 = {a?, as, b?, bs}, where each subband group corresponds to a 100 Hz band.

[0063] A frame with a gain of around -50 dB (or less) may have MDCT coefficients with values ​​on the order of 10⁻² or 10⁻¹, or even lower, while a frame with full-scale gain may have MDCT coefficients with values ​​of 20 or more. To accommodate this wide range of values, a shift factor is calculated based on the maximum available quantization points (max_value) and a sum of the absolute values ​​of the MDCT coefficients (lfe_dct_new) as follows: shift = floor(shifts_per_double*log2(max_value / sum(abs(lfe_dct_new)))),

[0064] In one implementation, lfe_dct_new is a set of 16 coefficients of MDCT, shifts_per_double is a constant (e.g., 4), max_value is an integer chosen for fine quantization (e.g., 6s quantization values) and for coarse quantization (e.g., S1 quantization values), and shift is limited to a 5-bit value from 4 to S5 for fine quantization and 2 to SS for coarse quantization.

[0065] The quantified MDCT coefficients are then calculated as follows: vals = round(lfe_dct_new*(2Λ(shift / shifts_per_double))), where the round() operation rounds the result to the nearest integer value.

[0066] If the quantized values ​​(vals) exceed the maximum allowed number of available quantization points (max_val), the shift factor is reduced and the quantized values ​​(vals) are recalculated. In other implementations, instead of the sum function sum(abs(lfe_dct_new))), the maximum function max(abs(lfe_dct_new))) can be used to calculate the shift factor, although the quantization values ​​will be more spread out using the max() function, making it more difficult to design an effective entropy encoder.

[0067] In the quantization steps described above, the quantized values ​​for each subband group are calculated together in a loop, but the points of 2S8S90 The 1,689,964 24-bit quantization values ​​are distinct for each subband group. If the first subband group exceeds the allowed range, then the scaling factor is reduced. If any of the other subband groups exceeds the allowed range, then that subband group is truncated to its maximum value. The sign bits for all MDCT coefficients and the absolute value of the quantized MDCT coefficients are encoded separately for each subband group.

[0068] Figure 5 illustrates the variation of the fine quantization points with frequency, according to one or more implementations. With fine quantization, subband group 1 (0–100 Hz) has 64 quantization points, subband group 2 (100–200 Hz) has 32 quantization points, subband group 3 (200–300 Hz) has 8 quantization points, and subband group 4 (300–400 Hz) has 2 quantization points. In one embodiment, each subband group is entropically encoded with a separate entropy encoder (e.g., an arithmetic or Huffman entropy encoder), where each entropy encoder uses a distinct probability distribution. Consequently, the primary 0–100 Hz range is allocated the majority of the quantization points.

[0069] It should be noted that the assignment of quantization points to subband groups 1-4 follows the shape of the LPF frequency response curve, which has more information at lower frequencies than at higher frequencies and no information outside the cutoff frequency. In order to reconstruct frequencies up to 130 Hz correctly, the MDCT coefficients corresponding to frequencies above 130 Hz are also encoded to avoid or minimize overlap. In some implementations, the MDCT coefficients up to 400 Hz are encoded so that frequencies up to 130 Hz can be properly reconstructed in the decoding unit.

[0070] Figure 6 illustrates the variation of coarse quantization points with frequency, according to one or more implementations. With coarse quantization, subband group 1 (0–100 Hz) has 32 quantization points, subband group 2 (100–200 Hz) has 16 quantization points, subband group 3 (200–300 Hz) has 4 quantization points, and subband group 4 (300–400 Hz) is unquantized and entropically encoded. In one embodiment, each subband group is entropically encoded with a separate entropy encoder using a distinct probability distribution. 238390 1689964 of 24

[0071] Figure 7 illustrates a probability distribution of finely quantized MDCT coefficients, according to one or more implementations. The y-axis is the frequency of occurrence and the x-axis is the number of quantization points. Sg1 is subband group 1 corresponding to MDCT coefficients quantized in the 0-100 Hz band, Sg2 is subband group 2 corresponding to MDCT coefficients quantized in the 100-200 Hz band, Sg3 is subband group 3 corresponding to MDCT coefficients quantized in the 200-300 Hz band, and Sg4 is subband group 4 corresponding to MDCT coefficients quantized in the 300-400 Hz band.

[0072] FIG. 8 illustrates a probability distribution of the coefficients of MDCTs quantized with coarse quantization, according to one or more implementations. The y-axis is the frequency of occurrence and the x-axis is the number of quantization points. Sg1 is subband group 1 corresponding to MDCT coefficients quantized in the 0-100 Hz band, Sg2 is subband group 2 corresponding to MDCT coefficients quantized in the 100-200 Hz band, Sg3 is subband group 3 corresponding to MDCT coefficients quantized in the 200-300 Hz band, and Sg4 is subband group 4 corresponding to MDCT coefficients quantized in the 300-400 Hz band.

[0073] It should be noted that the primary band (0-100 Hz) is where most of the LFE effects are found and, consequently, more quantization points are allocated to it for higher resolution. However, fewer bits are allocated to the primary band in coarse quantization than in fine quantization. In one embodiment, the use of fine or coarse quantization for an MDCT coefficient frame depends on the desired target bit rate set by the main audio channel encoder 103. The main audio channel encoder 103 sets this value either once during initialization or dynamically on a frame-by-frame basis, depending on the bits required or used to encode the main audio channels in each frame. Plots of silence

[0074] In some implementations, a signal is added to the LFE channel bitstream to indicate silent frames. A silent frame is a frame that has energy below a specific threshold. In some implementations, 1 bit is included in the LFE channel bitstream transmitted to the decoder (for example, inserted in the 238390 1689964 of 24 frame header) to indicate a silent frame, and all MDCT coefficients in the LFE channel bit stream are set to 0. This technique can reduce the bit transfer rate to 50 bps during silent frames. Decoder low-pass filter (LPF)

[0075] Two options are provided for implementing LPF 207 (see FIG. 2B) at the output of the LFE channel decoding unit 108. LPF 207 is selected based on the available delay (total delay of other audio channels minus LFE fade-out delay minus input LPF delay). It should be noted that other channels are expected to be encoded / decoded by main audio channel encoding / decoding units 103, 107, and the delays for those channels depend on the algorithmic delay of the main audio channel encoding / decoding units 103, 107.

[0076] In one implementation, if the available delay is less than 3.5 ms, a second-order Butterworth low-pass filter with a cutoff at 130 Hz is used; otherwise, a fourth-order Butterworth low-pass filter with a cutoff at 130 Hz is used. Thus, in the LFE 108 channel decoding unit, there is a trade-off between eliminating overlapping energy beyond the cutoff frequency and algorithmic delay. In some implementations, LPF 207 can be omitted entirely, as very low-frequency loudspeakers typically have an LPF. LPF 207 helps reduce overlapping energy beyond the cutoff at the LFE decoder output itself and can aid in efficient post-processing. Exemplary processes

[0077] FIG. 9 is a flowchart of a 900 process for encoding MDCT coefficients, according to one or more implementations. The 900 process can be implemented using, for example, the 1100 system, which is described with reference to FIG. 11.

[0078] Process 900 includes the steps of: receiving an LFE channel signal in the time domain (901), filtering, using a low-pass filter, the LFE channel signal in the time domain (902), converting the filtered LFE channel signal in the time domain into a frequency-domain representation of the LFE channel signal that includes a number of coefficients representing a frequency spectrum of the signal 238390 1689964 of 24 LFE channel (903); arrange the coefficients into a number of subband groups corresponding to different frequency bands of the LFE channel signal (904); quantize the coefficients in each subband group according to a frequency response curve of the low-pass filter using a scaling factor (905); encode the quantized coefficients in each subband group using an entropy encoder configured for the subband group (906); generate a bitstream that includes the encoded quantized coefficients (907); and store the bitstream in a storage device or transmit the bitstream to a downstream device (908).

[0079] FIG. 10 is a flow diagram of a process 1000 for decoding MDCT coefficients, according to one or more implementations. Process 1000 can be implemented using, for example, system 1100, which is described with reference to FIG. 11.

[0080] Process 1000 includes the steps of: receiving an LFE channel bitstream (1001), wherein the LFE channel bitstream includes entropically encoded coefficients representing a frequency spectrum of an LFE channel signal in the time domain; decoding and inversely quantizing the coefficients (1002), wherein the coefficients were quantized into subband groups corresponding to distinct frequency bands according to a frequency response curve of a low-pass filter using a scaling factor; converting the decoded and inversely quantized coefficients into an LFE channel signal in the time domain (1003); adjusting a delay of the LFE channel signal in the time domain (1004); and filtering, using a low-pass filter, the LFE channel signal with adjusted delay (1005).In one embodiment, the low-pass filter order can be configured based on the total algorithmic delay available from a primary codec used to encode / decode full-bandwidth channels of a multichannel audio signal that includes the LFE channel signal in the time domain. In some implementations, the decoding unit only needs to know whether the MDCT coefficients were encoded with fine or coarse quantization by the encoding unit. The quantization type can be indicated using a bit in the LFE bitstream header or by any other suitable signaling mechanism.

[0081] In some implementations, the decoding of inversely quantized coefficients in pulse-code modulated (PCM) samples in the time domain is performed as follows. The inversely quantized coefficients in each group 238390 The 1,689,964 sub-band values ​​of 24 are rearranged into N groups (N being the number of MDCTs calculated in the encoding unit), where each group has coefficients corresponding to the respective MDCT. According to the exemplary implementation described above, the encoding unit encodes the following 4 sub-band groups: subband group 1 = {a1, a2, b1, b2}, subband group 2 = {a3, a4, b3, b4}, subband group 3 = {a5, a6, b5, be}, subband group 4 = {a7, a8, b?, b8}.

[0082] The decoding unit decodes the four subband groups and reassembles them into {ai, a2, a3, a4, as, ae, a?, a8} and {bi, b2, b3, b4, bs, be, b?, b8}, and then pads the groups with zeros to obtain the desired inverse MDCT input length (iMDCT). N iMDCTs are performed to inversely transform the MDCT coefficients in each group into blocks in the time domain. In this example, each block is 2*Sw ms wide, where Sw is the subframe width defined above. This block is then windowed using the same Fielder window used by the LFE encoding unit shown in FIG. 4. Each subframe Si (i is an integer between 1 <= i <= N) is reconstructed by appropriate overlay, adding the windowed data from the previous iMDCT output and the current iMDCT output. Finally, the output of (1003) is reconstructed by concatenating all N subframes. Exemplary system architecture

[0083] Figure 11 is a block diagram of an 1100 system for implementing the features and processes described with reference to Figures 1-10, according to one or more implementations. The 1100 system includes one or more server computers or any client device, including but not limited to: call servers, user equipment, conference room systems, home theater systems, virtual reality (VR) equipment, and immersive content receiving devices. The 1100 system includes any consumer device, including but not limited to: smartphones, tablets, covert computers, vehicle computers, game consoles, surround sound systems, booths, etc. 238390 1689964 of 24

[0084] As shown, the 1100 system includes a central processing unit (CPU) 1101 that can perform various processes according to a program stored in, for example, a read-only memory (ROM) 1102 or a program loaded from, for example, a storage unit 1108 into random-access memory (RAM) 1103. Data required when the CPU 1101 performs the various processes is also stored in RAM 1103 as needed. The CPU 1101, ROM 1102, and RAM 1103 are connected to each other by a bus 1104. An input / output (I / O) interface 1105 is also connected to bus 1104.

[0085] The following components are connected to the I / O interface 1105: an input unit 1106, which may include a keyboard, mouse, etc.; an output unit 1107, which may include a display such as a liquid crystal display (LCD) and one or more speakers; the storage unit 1108, which includes a hard disk, or other suitable storage device; and a communication unit 1109, which includes a network interface card such as a network card (for example, wired or wireless).

[0086] In some implementations, the 1106 input unit includes one or more microphones in various positions (depending on the host device) that allow the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0087] In some implementations, the 1107 output unit includes systems with varying numbers of speakers. The 1107 output unit (depending on the capabilities of the host device) can represent audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0088] The communication unit 1109 is configured to communicate with other devices (for example, via a network). A disk drive 1110 is also connected to the I / O interface 1105, as required. Removable media 1111, such as a magnetic disk, optical disk, magneto-optical disk, flash drive, or other suitable removable media, is mounted in the unit 1110 so that a computer program read from it is installed on the storage unit 1108, as required. A person skilled in the art will understand that although the 1100 system is described as including the components described above, in actual applications, it is possible to add, remove, and / or 238390 1689964 of 24 replace some of these components, and all such modifications or alterations fall within the scope of this disclosure.

[0089] According to exemplary embodiments of this disclosure, the processes described above can be implemented as computer software programs or on a computer-readable storage medium. For example, the embodiments of this disclosure include a computer program product that includes a computer program tangibly realized on a computer-readable medium, a computer program that includes program code for performing methods. In such embodiments, the computer program can be downloaded and mounted from the network by means of communication unit 1309, and / or installed from removable media 1111.

[0090] In general, various exemplary embodiments of this disclosure may be implemented in special-purpose hardware or circuitry (e.g., control circuitry), software, logic systems, or any combination thereof. For example, the units described above may be executed by control circuitry (e.g., a CPU in combination with other components of FIG. 11), and thus the control circuitry may perform the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device (e.g., control circuitry).While several aspects of the exemplary realizations in this disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it should be noted that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-restrictive examples, in hardware, software, firmware, special-purpose circuits or logic systems, general-purpose hardware or controllers, or other computer devices, or some combination thereof.

[0091] Likewise, several blocks shown in the flowcharts can be viewed as steps in the method, and / or as operations resulting from the functioning of the computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, the realizations of this disclosure include a computer program product that includes a computer program tangibly realized on a human-readable medium 238390 1689964 of 24 computer, computer program containing program codes configured to carry out the methods described above.

[0092] In the context of disclosure, a computer-readable medium may be any tangible medium that can contain or store a program to be used by or in connection with a system, apparatus, or device to execute instructions. A computer-readable medium may be a computer-readable signaling medium or a computer-readable storage medium. A computer-readable medium may be non-transient and may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof.More specific examples of computer-readable storage media would include an electrical connection with one or more wires, a portable floppy disk, a hard disk, RAM, ROM, erasable programmable read-only memory (EPROM, or Flash memory), optical fiber, a portable compact disc with read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0093] The computer program code for carrying out the methods of this disclosure may be written in any combination of one or more programming languages. Such computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data-processing device having control circuitry, such that the program code, when executed by the processor of the computer or other programmable data-processing device, causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented.The program code can run entirely on one computer, partially on the computer as a standalone software package, partly on the computer and partly on a remote computer, or entirely on the remote computer or server, or it can be distributed among one or more remote computers and / or servers.

[0094] While this document contains many specific implementation details, these should not be interpreted as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features described herein in the context of separate embodiments may also be implemented in 238390 1689964 of 24 combinations in a single embodiment. Conversely, several features described in the context of a single embodiment may also be implemented in multiple separate embodiments or in any suitable subcombination. Furthermore, although some features may be described above as acting in certain combinations and even initially claimed as such, one or more features of a claimed combination may, in some cases, be removed from the combination, and the claimed combination may be directed to a subcombination or a variation of a subcombination. The logical flows illustrated in the figures do not require the particular order shown, or sequential order, to achieve desired results. Likewise, other steps may be provided, or steps may be removed, from the described flows, and other components may be added to, or removed from, the described systems.Therefore, other implementations fall within the scope of the following claims.

Claims

1. A method for encoding a low-frequency effects (LFE) channel, characterized in that it comprises: receiving, using one or more processors, a time-domain LFE channel signal; filtering, using a low-pass filter, the time-domain LFE channel signal to produce a filtered time-domain LFE channel signal, wherein the low-pass filter has a cutoff frequency; converting, using the one or more processors, the filtered time-domain LFE channel signal into a frequency-domain representation of the time-domain LFE channel signal that includes a number of coefficients representing a frequency spectrum of the time-domain LFE channel signal;arranging, using one or more processors, the coefficients in two or more subband groups corresponding to different frequency bands of the LFE channel signal in the time domain, wherein the different frequency bands include a main LFE frequency band that is below a cutoff frequency of an LFE speaker and at least one other LFE frequency band that is greater than the cutoff frequency of the LFE speaker, wherein each subband group has a width, and a sum of the widths of the subband groups includes the main LFE frequency band and the at least one other LFE frequency band; generating a scaling factor based on a maximum number of available quantization points and a sum of the absolute values ​​of the coefficients;Quantifying, using one or more processors, the coefficients of each subband group according to a frequency response curve of the low-pass filter and using the scaling factor to produce quantized coefficients; encoding, using one or more processors, the quantized coefficients of each subband group using an entropy encoder tuned for the subband group; and generating, using one or more processors, a bitstream that includes the encoded quantized coefficients. Three claims follow;