METHOD FOR ENCODING A LOW-FREQUENCY EFFECTS CHANNEL AND LOW-LATENCY LOW-FREQUENCY EFFECTS ENCODER

AR125559B2Active Publication Date: 2026-08-28DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
ARP20220100647
Authority / Receiving Office
AR · AR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-24
Filing Date
2022-03-18
Publication Date
2026-08-28
Estimated Expiration
2040-09-03

AI Technical Summary

Technical Problem

Existing audio codecs struggle to efficiently process low frequency effect (LFE) channels with low latency and high quality, particularly in immersive audio services, lacking optimal bit rate management and complexity in practical implementation.

Method used

A low latency LFE codec that encodes LFE channels by filtering, converting to frequency-domain, quantizing coefficients based on a frequency response curve, and encoding with entropy encoders, using subband groups and adjustable quantization schemes to manage bit rates effectively.

Benefits of technology

The codec achieves low latency and efficient bit rate management, supporting LFE channels with low algorithmic latency and adaptable configurations, ensuring high-quality audio reproduction across varying scenarios.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

In some implementations, a method for encoding a low-frequency effects (LFE) channel comprises: receiving an LFE channel signal in the time domain; filtering the LFE channel signal in the time domain using a low-pass filter; converting the filtered LFE channel signal in the time domain into a frequency-domain representation of the LFE channel signal that includes a number of coefficients representing a frequency spectrum of the LFE channel signal; arranging coefficients into a number of subband groups corresponding to different frequency bands of the LFE channel signal; quantizing coefficients in each subband group according to a frequency response curve of the low-pass filter; encoding the quantized coefficients in each subband group by using an entropy encoder tuned to the subband group;and generate a bitstream that includes the encoded quantized coefficients; and store the bitstream on a storage device or continuously transmit the bitstream to a downstream device.
Need to check novelty before this filing date? Find Prior Art

Description

LOW LATENCY LOW FREQUENCY EFFECTS CODEC TECHNICAL FIELD

[0001] This disclosure relates to audio signal processing in general and, in particular, to low frequency effect (LFE) channel processing. BACKGROUND

[0002] Standardization attempts for immersive services include development of an Immersive Voice and Audio Service (IVAS) codec for voice, multi-stream teleconferencing, virtual reality (VR), video streaming, live and delayed content generated by the user, for example. One goal of the IVAS standard is to develop a single codec with excellent audio quality, low latency, support for spatial audio coding, an appropriate range of bit rates, high quality error resilience, and complexity in practical implementation. . To achieve this goal, it is desired to develop an IVAS codec that can handle low latency LFE operations on IVAS-enabled devices or any other device capable of processing LFE signals. The LFE channel is intended for deep, bass sounds ranging from 20-120 Hz and is typically sent to a speaker designed to play low-frequency audio content. SYNTHESIS

[0003] Implementations for a low latency configurable LFE codec are disclosed.

[0004] In some implementations, a method for encoding a low frequency effects (LFE) channel comprises: receiving, using one or more processors, a time-domain LFE channel signal; filtering, using a low pass filter, the LFE channel signal in the time domain; converting, using the one or more processors, the filtered time-domain LFE channel signal into a frequency-domain representation of the LFE channel signal that includes a number of coefficients representing a frequency spectrum of the LFE channel signal; arranging, using the one or more processors, coefficients in a number of subband groups corresponding to different frequency bands of the LFE channel signal; quantify, using the one or more 238481 1723267 from 24 processors, coefficients in each subband group according to a low-pass filter frequency response curve; encoding, using the one or more processors, the quantized coefficients in each subband group by using an entropy encoder tuned for the subband group; and generating, using the one or more processors, a bit stream including the encoded quantized coefficients; and storing, using the one or more processors, the bit stream in a storage device or streaming the bit stream to a downstream device.

[0005] In some implementations, quantizing the coefficients in each subband group further comprises: generating a scaling factor based on a maximum number of available quantization points and a sum of the absolute values ​​of the coefficients; and quantizing the coefficients using the scaling factor.

[0006] In some implementations, if a quantized coefficient exceeds the maximum number of quantization points, the scaling factor is reduced and the coefficients are re-quantized.

[0007] In some implementations, the quantization points are different for each subband group.

[0008] In some implementations, the coefficients in each subband group are quantized according to a fine quantization scheme or a coarse quantization scheme, whereby the fine quantization scheme assigns more quantization points to one or more subband groups. subband than those assigned to the respective subband groups according to the coarse quantization scheme.

[0009] In some implementations, the sign bits for the coefficients are encoded separately from the coefficients.

[0010] In some implementations, there are four subband groups, and a first subband group corresponds to a first frequency range of 0-100 Hz, a second subband group corresponds to a second frequency range of 100-200 Hz, a third subband group corresponds to a third frequency range of 200-300 Hz and a fourth subband group corresponds to a fourth frequency range of 300-400 Hz.

[0011] In some implementations, the entropy encoder is an arithmetic entropy encoder or a Huffman entropy encoder.

[0012] In some implementations, converting the filtered time-domain LFE channel signal into a frequency-domain representation of the LFE channel signal that includes a number of coefficients representing a spectrum of 238481 1723267 of the LFE channel signal frequency, further comprising: determining a first stride length of the LFE channel signal; designating a first window size of a windowing function based on the first step length; applying the first window size to one or more frames of the LFE channel signal in the time domain; and applying a Modified Cosine Discrete Transform (MDCT) to the windowed frames to generate the coefficients.

[0013] In some implementations, the method further comprises: determining a second pitch length of the LFE channel signal; designating a second window size of the windowing function based on the second step length; and applying the second window size to the one or more frames of the LFE channel signal in the time domain.

[0014] In some implementations, the first step length is N milliseconds (ms), N is greater than or equal to 5 ms and less than or equal to 60 ms, the first window size is greater than or equal to 10 ms , the second step length is 5 ms and the second window size is 10 ms.

[0015] In some implementations, the first step length is 20 milliseconds (ms), the first window size is 10 ms or 20 ms or 40 ms, the second step length is 10 ms, and the second window size is 10 more or 20 ms.

[0016] In some implementations, the first step length is 10 milliseconds (ms), the first window size is 10 ms or 20 ms, the second step length is 5 ms, and the second window size is 10 ms.

[0017] In some implementations, the first step length is 20 milliseconds (ms), the first window size is 10 ms, 20 ms, or 40 ms, the second step length is 5 ms, and the second window size is 10 more.

[0018] In some implementations, the windowing function is a Kaiser-Bessel derived (KBD) function with a configurable fade length.

[0019] In some implementations, the low pass filter is a low pass filter Fourth-order Butterworth with a cutoff frequency of approximately 130 Hz or less.

[0020] In some implementations, the method further comprises: determining, using the one or more processors, whether an energy level of an LFE channel signal frame is below a threshold; according to the energy level that is below 238481 1723267 24 a threshold level, generate a silence frame flag indicating this to the decoder; inserting the silence frame indicator into metadata of the LFE channel bitstream; and reducing a bit rate of the LFE channel upon detection of the silence frame.

[0021] In some implementations, a method for decoding a low frequency effect (LFE) comprises: receiving, using one or more processors, an LFE channel bit stream, LFE channel bit stream including entropy coded coefficients representing a frequency spectrum of an LFE channel signal in the time domain; decoding, using the one or more processors, the quantized coefficients using an entropy decoder; reverse quantizing, using the one or more processors, the inverse quantized coefficients, where the coefficients were quantized into subband groups corresponding to frequency bands according to a frequency response curve of a low-pass filter used to filter the time domain LFE channel signal in an encoder; converting, using the one or more processors, the inversely quantized coefficients to a time-domain LFE channel signal; adjusting, using the one or more processors, a delay of the LFE channel signal in the time domain; and filtering, using a low-pass filter, the delay-adjusted LFE channel signal.

[0022] In some implementations, a low pass filter order is configured to ensure that a first total algorithmic delay due to LFE channel encoding and decoding is less than or equal to a second total algorithmic delay due to encoding and decoding of the LFE channel. decoding other audio channels of a multi-channel audio signal including the LFE channel signal.

[0023] In some implementations, the method further comprises: determining if the second total algorithmic delay exceeds a threshold value; and according to the second total algorithmic delay that exceeds the threshold value, configuring the low-pass filter as an N-order low-pass filter, where N is an integer greater than or equal to two; and according to the second total algorithmic delay that does not exceed the threshold value, set the order of the low-pass filter to be less than N.

[0024] Other implementations disclosed herein are directed to a computer-readable system, apparatus, and medium. Details of the disclosed implementation are set forth in the accompanying drawings and in the description below. other 238481 1723267 of 24 Features, objects, and advantages are apparent from the description, drawings, and claims.

[0025] Particular embodiments disclosed herein provide one or more of the following advantages. The disclosed low latency LFE codec: 1) mainly targets the LFE channel; 2) primarily targets the 20 to 120 Hz frequency range, but carries audio at 300 Hz in low / medium bitrate scenarios and 400 Hz in high bitrate scenarios; 3) achieves a low bit rate by applying a quantization scheme according to a frequency response curve to an input low-pass filter; 4) has low algorithmic latency and is designed to operate at a 20 millisecond (ms) step and have a total algorithmic latency (including frame generation) of 33 milliseconds; 5) can be configured for smaller steps and lower algorithmic latency to support other scenarios, including lower step settings of 5 milliseconds and total algorithmic latency (including frame generation) of 13 milliseconds; 6) automatically chooses a low-pass filter at the decoder output based on the latency available with the LFE codec; 7) has a silent mode with a low bit rate of 50 bits per second (bps) during silence; and 8) during active frames the bit rate varies between 2 kilobits per second (kbps) and 4 kbps based on the quantization level used, and during silent frames the bit rate is 50 bps. DESCRIPTION OF THE DRAWINGS

[0026] In the drawings, for ease of description, specific arrangements or arrangements of schematic elements are shown, such as those representing devices, units, instruction blocks, and data elements. However, those skilled in the art will understand that the specific arrangement or ordering of schematic elements in the drawings does not imply that any particular order or sequence of processing or separation of processes is required. Furthermore, the inclusion of a schematic element in a drawing does not mean that said element is required in all implementations or that the features represented by said element cannot be included in or combined with other elements in some implementations. 238481 1723267 of 24

[0027] In addition, in the drawings, when connecting elements, such as solid or broken lines or arrows, are used to illustrate a connection, relationship or association between two or more other schematic elements, the absence of any such connecting element does not assumes that there can be no connection, relationship or association. In other words, some connections, relationships or associations between elements are not shown in the drawings so as not to detract from the disclosure. Also, for ease of illustration, a single connection element is used to represent multiple connections, relationships, or associations between elements. For example, when a connection element represents a communication of signals, data, or instructions, it will be understood by those skilled in the art that such element represents one or multiple signal paths, as necessary, to affect the communication.

[0028] FIG 1 illustrates an IVAS codec for encoding and decoding IVAS and LFE bitstreams, according to one or more implementations.

[0029] FIG 2A is a block diagram illustrating LFE coding, according to one or more implementations.

[0030] FIG 2B is a block diagram illustrating LFE decoding, according to one or more implementations.

[0031] FIG 3 is a graph illustrating a fourth-order Butterworth low-pass filter frequency response with a 130 Hz corner or cutoff, according to one or more implementations.

[0032] FIG 4 is a graph illustrating a Fielder window, according to one or more implementations.

[0033] FIG 5 illustrates the variation of fine quantization points with frequency, according to one or more implementations.

[0034] FIG 6 illustrates the variation of coarse quantization points with frequency, according to one or more implementations.

[0035] FIG 7 illustrates a probability distribution of finely quantized MDCT coefficients, according to one or more implementations.

[0036] FIG 8 illustrates a probability distribution of coarsely quantized MDCT coefficients, according to one or more implementations.

[0037] FIG 9 is a flowchart of a Modified Discrete Cosine Transform (MDCT) coefficient encoding process, according to one or more implementations. 238481 1723267 of 24

[0038] FIG 10 is a flowchart of a Modified Discrete Cosine Transform (MDCT) coefficient decoding process, according to one or more implementations.

[0039] FIG 11 is a block diagram of a system for implementing the features and processes described with reference to FIGS 1-10, according to one or more implementations.

[0040] The same reference symbol used in various drawings indicates similar elements. DETAILED DESCRIPTION

[0041] In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the various described embodiments. It will be apparent to those skilled in the art that the various described implementations can be put into practice without these specific details. In other instances, known methods, procedures, components, and circuitry have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Various features are described hereinafter that can be used each independently of the other or with any combination of other features. Nomenclature

[0042] As used herein, the term "includes" and its variants should be read as open terms meaning "includes, without limitation." The term "or" should be read as "and / or" unless the context clearly indicates otherwise. The term "based on" should be read as "based at least in part on." The term "one exemplary implementation" should be read as "at least one exemplary implementation". The term "other implementation" should be read as "at least one other implementation". The terms "determined," "determines," or "determining" should be read as obtain, receive, compute, calculate, estimate, predict, or derive. Furthermore, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as is normally understood by the person skilled in the art to which this disclosure pertains. System overview 238481 1723267 of 24

[0043] FIG 1 illustrates an IVAS codec 100 for encoding and decoding IVAS bitstreams, including an LFE channel bitstream, according to one or more implementations. For encoding, the IVAS codec 100 receives N+1 channels of audio data 101, where N channels of audio data 101 are input to the spatial analysis and downmix unit 102 and one channel of LFE is input to the LFE channel encoding 105. Audio data 101 includes but is not limited to: mono signals, stereo signals, binaural signals, spatial audio signals (e.g., multi-channel spatial audio objects), First Order Ambisonics (FoA), First Order Ambisonics (FoA), Higher Order (HoA) and any other audio data.

[0044] In some implementations, the spatial analysis and downmix unit 102 is configured to implement complex advanced coupling (CACPL) to analyze / downmix stereo audio data and / or spatial reconstruction (SPAR). ) to analyze / downmix FoA audio data. In other implementations, the spatial analysis and downmix unit 102 implements other formats. The output of the spatial analysis and downmix unit 102 includes spatial metadata and 1 to N channels of audio data. The spatial metadata is input into the spatial metadata encoding unit 104, which is configured to quantize and entropy encode the spatial metadata. In some implementations, quantization may include fine, moderate, coarse, and extra-coarse quantization strategies, and entropy coding may include arithmetic or Huffman coding.

[0045] Audio data channels 1 to N are input into the main audio channel encoding unit 103 which is configured to encode audio data channels 1 to N into one or more voice service bit streams Enhanced EVS (EVS). In some implementations, the main audio channel coding unit 103 is compliant with 3GPP TS 26.445 and provides a wide range of functionality, such as improved coding quality and efficiency for narrowband voice services (EVS-NB). and broadband (EVS-WB), enhanced quality using super-wideband voice (EVS-SWB), enhanced quality for music and content mixed in conversational applications, robustness against packet loss, and delay variation and 238481 1723267 of 24 backwards compatibility for the AMR-WB (Adaptive Multi-Rate Wideband) codec.

[0046] In some implementations, the main audio channel encoding unit 103 includes a preprocessing and mode selection unit that selects between a vocoder for encoding speech signals and a perceptual encoder for encoding audio signals at a rate bit rate specified based on bit rate / mode control. In some implementations, the vocoder is an enhanced variant of Algebraic Code Excited Linear Prediction (ACELP), extended with specialized modes based on LP (Linear Prediction) for different classes of speech. voice.

[0047] In some implementations, the audio encoder is a Modified Cosine Discrete Transform (MDCT) encoder with increased efficiency at low delay / low bit rates and is designed to perform continuous and reliable switching between audio encoders. voice and audio

[0048] As described above, the LFE channel signal is intended for deep bass sounds ranging from 20 to 120 Hz and is typically sent to a speaker designed to reproduce low frequency audio content (for example, a speaker for very low frequencies). The LFE channel signal is input to the LFE channel signal encoding unit 105 which is configured to encode the LFE channel signal as described with reference to FIG 2A.

[0049] In some implementations, an IVAS decoder includes spatial metadata decoding unit 106 that is configured to recover spatial metadata, and main audio channel decoding unit 107 that is configured to recover audio signals from channels 1 to N. The retrieved spatial metadata and the audio signals of channels 1 to N recovered are input to the spatial synthesis / upmix / rendering unit 109, which is configured to synthesize and render the audio signals of the channels 1 to N. channels 1 to N into output audio signals of N or more channels (for example, N+1) using the spatial metadata to be reproduced through the speakers of various audio systems, including, without limitation: home theater systems home, video conferencing room systems, virtual reality (VR) equipment, and any other audio system that is capable of rendering audio. The LFE channel decoding unit 108 receives the LFE bit stream and is 238481 1723267 of 24 configured to decode the LFE bitstream, as described with reference to FIG 2B.

[0050] Although the exemplary implementation of encoding / decoding of LFE described above is realized by an IVAS codec, the low latency LFE codec described below may be a standalone LFE codec, or may be included in any standard or proprietary audio codec that encodes and decodes low frequency signals in audio applications where low latency and configurability are required or desired.

[0051] FIG 2A is a block diagram illustrating the functional components of the LFE channel coding unit 105 shown in FIG 1, according to one or more embodiments. FIG 2B is a block diagram illustrating the functional components of the LFE channel decoder 108 shown in FIG 1, according to one or more embodiments. The LFE channel decoder 108 includes the inverse quantization and entropy decoding unit 204, the inverse MDCT and windowing unit 205, the delay adjustment unit 206, and the output low-pass filter 207. The adjustment unit The delay filter 206 may be before or after the low-pass filter 207, and performs delay adjustment (for example, by buffering the decoded LFE channel signal) to match the decoded LFE channel signal and the decoded output of the main codec. Hereinafter, the LFE channel encoding unit 105 and the LFE channel decoding unit 108 described in connection with FIG. 2B are collectively referred to as an LFE codec.

[0052] The LFE channel coding unit 105 includes an input low pass filter (LPF) 201, a windowing and MDCT generation unit 202, and an entropy coding and quantization unit 203. In one embodiment, the signal The input audio signal is a Pulse Code Modulated (PCM) audio signal, and the LFE channel encoding unit 105 expects an input audio signal with a step of either 5 milliseconds, 10 milliseconds or 20 milliseconds. Internally, the LFE channel coding unit 105 operates in 5-millisecond or 10-millisecond subframes and windowing and MDCT are performed over a combination of these subframes. In one embodiment, the LFE channel encoding unit 105 executes with an input step of 20 milliseconds and internally divides this input into two subframes of equal length. The last subframe of the input frame prior to the LFE is 238481 1723267 of 24 concatenates with the first subframe of the current input frame to the LFE and windows are generated. The first subframe of the current input frame to the LFE is concatenated with the second subframe of the current input frame to the LFE and windows are generated. The MDCT is performed twice, once on each windowed block.

[0053] In one embodiment, the algorithmic delay (no frame delay) is equal to 8 milliseconds plus the delay incurred by input LPF 103 plus the delay incurred by output LPF 207. With a fourth-order input LPF 201 and a fourth-order output LPF 207, the total latency of the system is approximately 15 milliseconds. With a fourth order input LPF 201 and a second order output LPF 207, the total latency of the LFE codec is approximately 13 milliseconds.

[0054] FIG 3 is a graph illustrating a frequency response of an exemplary input LPF 201, according to one or more embodiments. In the example shown, the LPF 201 is a 4th order Butterworth filter with a cutoff frequency of 130 Hz. Other implementations may use a different type of LPF (for example, Chebyshev, Bessel) with the same or different order and the same or different cutoff frequencies.

[0055] FIG 4 is a graph illustrating a Fielder window, according to one or more embodiments. In one embodiment, the windowing function applied by the windowing unit and MDCT 202 is a Fielder windowing function with a fade length of 8 milliseconds. The Fielder window is a Kaiser-Bessel-derived (KBD) window with alpha=5, which is a window that by construction meets the Princen-Bradley condition for MDCT, and is thus used with the audio format digital Advanced Audio Coding (AAC). Other window generation functions can also be used. Quantification and entropy coding

[0056] In one embodiment, the entropy coding and quantization unit 203 implements a quantization strategy that follows the frequency response curve of the input LPF 201 to quantize the MDCT coefficients more efficiently. In one embodiment, the frequency range is divided into 4 subband groups representing 4 frequency bands: 0-100 Hz, 100-200 Hz, 200-300 Hz, and 300-400 Hz. These bands are examples and can be used more or less bands with the same or different frequency ranges. In particular, the MDCT coefficients are quantized using a scaling factor that is dynamically calculated based on the MDCT coefficient values. 238481 1723267 of 24 MDCT on a particular frame and the quantization points are selected according to the LPF frequency response curve, as shown in FIGS 5-8. This quantization strategy helps to reduce the quantization points for the MDCT coefficients belonging to the 100-200 Hz, 200-300 Hz, and 300-400 Hz bands, while maintaining optimal quantization points for the main LFE band. 0-100 Hz, which is where the energy for most low-frequency effects (such as background noise) will be found.

[0057] In one embodiment, a quantization strategy is described below for an input PCM step of Fien milliseconds (ms) (input frame length) to the LFE channel coding unit 105, where the frame length , Well, it can take any value given by 5*f ms, here 1<=f<=12.

[0058] First, the input PCM step is divided into N subframes of equal lengths, each subframe having a width of (Sw) = Fien / N ms. N must be selected such that each Sw is a multiple of 5 ms (for example, if Fien = 20 ms then N can be 1, 2 or 4; if Fien = 10 ms then N can be 1 or 2; and if Fien = 5 plus then N equals 1). Suppose Si is the ith subframe in any given frame, here i is an integer with rank 0 <= i <= N, where S0 corresponds to the last subframe of the input frame prior to LFE coding unit 105 and S1 to Sn are the N subframes of the current frame.

[0059] Each Si and Si+1 subframe is then concatenated and windowed with a Fielder window (see FIG 4) and then MDCT is performed on these windowed samples. This results in a total of N MDCTs for each frame. The number of MDCT coefficients coming from each MDCT (num_coeffs) = sample rate* Sw / 1000. The frequency resolution of each MDCT (width of each MDCT coefficient) (Wmdct) is about 1000 / (2*Sw) Hz. Since woofer drivers typically have an LPF cutoff of about 100-120 Hz, and since the post-LPF energy after 400 Hz is generally very low, the MDCT coefficients up to 400 Hz are quantized and sent to the LFE decoding unit 108 while the rest of the MDCTs are quantized to 0. Sending the MDCT coefficients up to 400 Hz ensures high-quality reconstruction up to 120 Hz in the LFE decoding unit 108. The total number of MDCT coefficients to quantize and encode (Nquant) is , therefore, equal to N*400 / Wmdct.

[0060] The MDCT coefficients are then arranged into M subband groups where the width of each subband group is a multiple of Wmdct and the sum of the widths of 238481 1723267 of 24 all subband groups equals 400 Hz. Suppose the width of each subband is SBWm Hz, where m is an integer with rank 1 <= m <= M. With this width, the number of coefficients in the m° subband group = SNquant = N* SBWm / Wmdct (ie SBWΜ / Wmdct coefficients of each MDCT). The MDCT coefficients in each subband group are then scaled with a scaling factor (shift), described below, determined by the sum or maximum of the absolute values ​​of all MDCT coefficients Nquant. The scaled MDCT coefficients in each subband group are then separately quantized and encoded using a quantization scheme that follows the curve of the LPF at the encoder input. Encoding of the quantized MDCT coefficients is performed with an entropy encoder (eg, an arithmetic or Huffman encoder). Each subband group is coded with a different entropy coder and each entropy coder uses an appropriate probability distribution model to code the respective subband group efficiently.

[0061] An exemplary quantization strategy for a 20 millisecond (ms) step (Fien = 20 ms), 2 subframes (N = 2) and sample rate = 48000 will now be described. With this exemplary input configuration, the subframe width Sw = 10 ms and the number of MDCTs = N = 2. The first MDCT is performed over a 20 ms block. This block is formed by concatenating a 10-20ms subframe of the previous 20ms input and a 0-10ms subframe of the current 20ms input, and then windowing with the 20ms Fielder window long (see FIG 4). With N = 1 and N = 4, the Fielder window is scaled accordingly and the fade length is changed to 16 / N ms. The second MDCT is performed on a 20 ms block formed by windowing the current 20 ms input frame with a 20 ms long Fielder window. The number of MDCT coefficients (num_coeffs) with each MDCT = 480, the width of each MDCT coefficient Wmdct = 50 Hz, the total number of coefficients to quantize and encode Nquant = 16, and the total number of coefficients to quantize and encode according to MDCT = 16 / N = 8.

[0062] Then, the MDCT coefficients are arranged in 4 subband groups (M=4), where each subband group corresponds to a 100 Hz band (0-100, 100-200, 200-300, 300-400 , SBWm =100 Hz, number of coefficients in each subband group = SNquant=N*SBWm / Wmdct = 4). Suppose a1, a2, as, a4, as, a6, a?, as are the first 8 MDCT coefficients to be quantized from the first MDCT and b1, b2, bs, b4, 238481 1723267 of 24 bs, bó, b?, bs are the first 8 MDCT coefficients to be quantized from the second MDCT. The 4 subband groups are arranged to have the following coefficients: subband group 1 = {ai, a2, bi, b2}, subband group 2 = {as, a4, bs, b4}, subband group S = {as , a6, bs, bó}, subband group 4 = {a?, as, b?, bs}, where each subband group corresponds to a 100 Hz band.

[0063] A frame with a gain around -S0 dB (or less) can have MDCT coefficients with values ​​in the order of 10-2 or 10-1, or even less, while a frame with full scale gain may have MDCT coefficients with values ​​of 20 or more. To meet this wide range of values, a scaling factor (shift) is calculated based on the maximum available quantization points (max_value) and an absolute value sum of the MDCT coefficients (lfe_dct_new) as follows : shift = floor(shifts_per_double*log2(max_value / sum(abs(lfe_dct_new)))),

[0064] In one implementation, lfe_dct_new is a set of 16 coefficients of MDCT, shifts_per_double is a constant (for example, 4), max_value is an integer chosen for fine quantization (for example, 6s quantization values) and for coarse quantization (for example, it is S1 quantization values), and shift is limited to a 5-bit value of 4 to S5 for fine quantization and 2 to SS for coarse quantization.

[0065] The quantized MDCT coefficients are then calculated as follows: waltz = round(lfe_dct_new*(2Λ(shift / shifts_per_double))), where the round() operation rounds the result to the nearest integer value.

[0066] If the quantized values ​​(vals) exceed the maximum allowed number of available quantization points (max_val), the scale factor (shift) is reduced and the quantized values ​​(vals) are recalculated. In other implementations, instead of the sum function sum(abs(lfe_dct_new))), the maximum function max(abs(lfe_dct_new))) can be used to calculate the scaling factor (shift), even though the values ​​of quantization will be more sparse using the max() function, which makes designing an efficient entropy encoder more difficult.

[0067] In the quantization steps described above, the quantized values ​​for each subband group are computed together in a loop, but the quantization points 2S8481 1723267 of 24 quantization are different for each subband group. If the first subband group exceeds the allowed range, then the scaling factor is reduced. If any of the other subband groups exceed the allowed range, then that subband group is truncated to the max_value. The sign bits for all MDCT coefficients and the absolute value of the quantized MDCT coefficients are coded separately for each subband group.

[0068] FIG 5 illustrates the variation of fine quantization points with frequency, according to one or more implementations. With fine quantization, subband group 1 (0-100 Hz) has 64 quantization points, subband group 2 (100-200 Hz) has 32 quantization points, subband group 3 (200-300 Hz) has 8 quantization points and subband group 4 (300-400 Hz) has 2 quantization points. In one embodiment, each subband group is entropy coded with a separate entropy coder (eg, a Huffman or arithmetic entropy coder), where each entropy coder uses a different probability distribution. Therefore, the primary 0-100 Hz range is assigned most of the quantization points.

[0069] It should be noted that the assignment of quantization points to subband groups 1-4 follows the shape of the LPF frequency response curve, which has more information at lower frequencies than higher frequencies and no information at all. outside the cutoff frequency. In order to reconstruct frequencies up to 130 Hz correctly, the MDCT coefficients corresponding to frequencies above 130 Hz are also coded to avoid or minimize aliasing. In some implementations, MDCT coefficients up to 400 Hz are encoded so that frequencies up to 130 Hz can be properly reconstructed in the decoding unit.

[0070] FIG 6 illustrates the variation of coarse quantization points with frequency, according to one or more implementations. With coarse quantization, subband group 1 (0-100 Hz) has 32 quantization points, subband group 2 (100-200 Hz) has 16 quantization points, subband group 3 (200-300 Hz) it has 4 quantization points and subband group 4 (300-400 Hz) is unquantized and entropy coded. In one embodiment, each subband group is entropy coded with a separate entropy coder using a different probability distribution. 238481 1723267 of 24

[0071] FIG 7 illustrates a probability distribution of finely quantized MDCT coefficients, according to one or more implementations. The y axis is the frequency of occurrence and the x axis is the number of quantization points. Sg1 is subband group 1 that corresponds to the quantized MDCT coefficients in the 0-100 Hz band, Sg2 is subband group 2 that corresponds to the quantized MDCT coefficients in the 100-200 Hz band. Sg3 ​​is subband group 3 corresponding to the quantized MDCT coefficients in the 200-300 Hz band. Sg4 is subband group 4 corresponding to the quantized MDCT coefficients in the 300-400 Hz band.

[0072] FIG 8 illustrates a probability distribution of the coefficients of MDCTs quantized with coarse quantization, according to one or more implementations. The y axis is the frequency of occurrence and the x axis is the number of quantization points. Sg1 is subband group 1 that corresponds to the quantized MDCT coefficients in the 0-100 Hz band, Sg2 is subband group 2 that corresponds to the quantized MDCT coefficients in the 100-200 Hz band. Sg3 ​​is subband group 3 corresponding to the quantized MDCT coefficients in the 200-300 Hz band. Sg4 is subband group 4 corresponding to the quantized MDCT coefficients in the 300-400 Hz band.

[0073] It should be noted that the primary band (0-100 Hz) is where most of the LFE effects are found and therefore more quantization points are assigned to them for higher resolution. However, there are fewer bits allocated to the primary band in coarse quantization than for fine quantization. In one embodiment, the use of fine quantization or coarse quantization for an MDCT coefficient frame depends on the desired target bit rate set by the main audio channel encoder 103. The main audio channel encoder 103 sets this value once during initialization or dynamically on a frame by frame basis based on the bits required or used to encode the main audio channels in each frame. plots of silence

[0074] In some implementations, a signal is added to the LFE channel bitstream to indicate frames of silence. A silence frame is a frame that has power below a specified threshold. In some implementations, 1 bit is included in the LFE channel bit stream transmitted to the decoder (for example, inserted into the 238481 1723267 of 24 frame header) to indicate a frame of silence, and all MDCT coefficients in the LFE channel bit stream are set to 0. This technique can reduce the bit rate to 50 bps during frames of silence. Decoder Low Pass Filter (LPF)

[0075] Two options are provided to implement the LPF 207 (see FIG 2B) at the output of the LFE channel decoding unit 108. The LPF 207 is selected based on the available delay (total delay of other audio channels minus LFE fade delay minus input LPF delay). It should be noted that other channels are expected to be encoded / decoded by main audio channel encode / decode units 103, 107, and the delays for those channels depend on the algorithmic delay of the main audio channel encode / decode units 103. , 107.

[0076] In one implementation, if the available delay is less than 3.5 ms, a second order Butterworth low pass filter with cutoff at 130 Hz is used; otherwise, a fourth-order Butterworth low-pass filter with cutoff at 130 Hz is used. Thus, in the LFE channel decoding unit 108 there is a trade-off between overlapping energy removal beyond the cutoff frequency and algorithmic delay. In some implementations, the LPF 207 may be eliminated entirely since subwoofers typically have an LPF. The LPF 207 helps reduce aliasing power beyond clipping at the output of the LFE decoder itself and can aid in efficient post-processing. exemplary processes

[0077] FIG 9 is a flowchart of a process 900 for encoding MDCT coefficients, according to one or more implementations. Process 900 can be implemented using, for example, system 1100, which is described with reference to FIG 11.

[0078] Process 900 includes the steps of: receiving a time-domain LFE channel signal (901), filtering, using a low-pass filter, the time-domain LFE channel signal (902) , converting the filtered time-domain LFE channel signal into a frequency-domain representation of the LFE channel signal that includes a number of coefficients representing a frequency spectrum of the time-domain signal. 238481 1723267 24 channel LFE (903); arranging the coefficients into a number of subband groups corresponding to different frequency bands of the LFE channel signal (904); quantizing the coefficients in each subband group according to a frequency response curve of the low pass filter using a scaling factor (905); encoding the quantized coefficients in each subband group using an entropy encoder configured for the subband group (906); generating a bit stream including the encoded quantized coefficients (907); and storing the bit stream in a storage device or transmitting the bit stream to a downstream device (908).

[0079] FIG 10 is a flowchart of a process 1000 for decoding MDCT coefficients, according to one or more implementations. Process 1000 can be implemented using, for example, system 1100, which is described with reference to FIG 11.

[0080] Process 1000 includes the steps of: receiving an LFE channel bit stream (1001), where the LFE channel bit stream includes entropy coded coefficients representing a frequency spectrum of an LFE channel signal in the domain of time; decoding and inverse quantizing the coefficients (1002), where the coefficients were quantized into subband groups corresponding to different frequency bands according to a frequency response curve of a low pass filter using a scaling factor; converting the decoded and inversely quantized coefficients into a time-domain LFE channel signal (1003); setting a delay of the LFE channel signal in the time domain (1004); and filtering, using a low-pass filter, the delay-adjusted LFE channel signal (1005). In one embodiment, the low-pass filter order may be configured based on a total algorithmic delay available from a primary code used to encode / decode full-bandwidth channels of a multi-channel audio signal including the channel signal. of LFE in the time domain. In some implementations, the decoding unit only needs to know whether the MDCT coefficients were encoded with fine or coarse quantization by the encoding unit. The quantization type may be indicated using a bit in the LFE bitstream header or by any other suitable signaling mechanism.

[0081] In some implementations, decoding of inverse quantized coefficients in pulse code modulated (PCM) samples in the time domain is performed as follows. The inversely quantified coefficients in each group 238481 1723267 out of 24 subband are rearranged into N groups (N is the number of MDCTs computed in the coding unit), where each group has coefficients corresponding to the respective MDCT. According to the exemplary implementation described above, the coding unit encodes the following 4 subband groups: subband group 1 = {at, a2, b1, b2}, subband group 2 = {as, a4, bs, bg, subband group S = {as, ae, bs, be}, subband group 4 = {a?, as, b?, bs}.

[0082] The decoding unit decodes the 4 subband groups and rearranges them into {ai, a2, as, a4, as, ae, a?, as} and {bi, b2, bs, b4, bs, be , b?, bs}, and then pad the groups with zeros to get the desired inverse MDCT (iMDCT) input length. N iMDCTs are performed to inversely transform the MDCT coefficients in each block group in the time domain. In this example, each block is 2*Sw ms wide, where Sw is the subframe width defined above. This block is then windowed using the same Fielder window used by the LFE coding unit shown in FIG 4. Each If subframe (i is an integer between 1 <= i <= N) is reconstructed by an appropriate overlay by adding the windowed data from the previous iMDCT output and the current iMDCT output. Finally, the output of (100S) is reconstructed by concatenating all N subframes. Exemplary system architecture

[0083] FIG 11 is a block diagram of a system 1100 for implementing the features and processes described with reference to FIGS 1-10, according to one or more implementations. System 1100 includes one or more server computers or any client device, including but not limited to: call servers, user equipment, conference room systems, home theater systems, virtual reality (VR) equipment, and devices for receiving immersive content. The 1100 system includes any consumer device, including but not limited to: smartphones, tablets, stealth computers, vehicle computers, game consoles, surround systems, booths, etc. 2SS4S1 1723267 of 24

[0084] As shown, system 1100 includes a central processing unit (CPU) 1101 that can perform various processes in accordance with a program stored in, for example, read-only memory (ROM, 1102 or a program loaded from, for example, a storage unit 1108 to a random access memory (RAM) 1103. In the RAM 1103, also stored, as appropriate, necessary, the data required when the CPU 1101 performs the various processes. The CPU 1101, the ROM 1102, and the RAM 1103 are connected to each other by a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0085] The following components are connected to the I / O interface 1105: an input unit 1106, which may include a keyboard, mouse, etc.; an output unit 1107 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 1108 including a hard drive, or other suitable storage device; and a communication unit 1109 including a network interface card as a network card (eg, wired or wireless).

[0086] In some implementations, the input unit 1106 includes one or more microphones in different positions (depending on the host device) that allow the capture of audio signals in various formats (for example, mono, stereo, spatial, immersive, and others). suitable formats).

[0087] In some implementations, output unit 1107 includes systems with various numbers of speakers. Output unit 1107 (depending on the capabilities of the host device) can represent audio signals in various formats (eg, mono, stereo, immersive, binaural, and other suitable formats).

[0088] Communication unit 1109 is configured to communicate with other devices (eg via a network). An 1110 disk drive is also connected to the 1105 I / O interface, as needed. A removable medium 1111, such as a magnetic disk, optical disk, magneto-optical disk, flash drive, or other suitable removable medium is mounted in drive 1110, so that a computer program read from there is installed in storage drive 1108 , as necessary. One skilled in the art will understand that while system 1100 is described as including the components described above, in actual applications, it is possible to add, remove, and / or 238481 1723267 of 24 to replace some of these components, and all such modifications or alterations fall within the scope of this disclosure.

[0089] According to exemplary embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product that includes a computer program tangibly embodied on a computer-readable medium, a computer program that includes program code for performing methods. In such embodiments, the software may be downloaded and mounted from the network via communication unit 1309, and / or installed from removable media 1111.

[0090] In general, various exemplary embodiments of the present disclosure may be implemented in hardware or special purpose circuitry (eg, control circuitry), software, logic, or any combination thereof. For example, the units discussed above can be executed by control circuitry (for example, a CPU in combination with other components of FIG 11) and thus the control circuitry can perform the actions described in the present disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device (eg, control circuitry). While various aspects of the exemplary embodiments of the present disclosure are illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it should be noted that the blocks, apparatus, systems, techniques, or methods described in the present may be implemented, as non-limiting examples, in hardware, software, firmware, special purpose circuitry or logic system, general purpose hardware or controller or other computing devices, or some combination thereof.

[0091] Likewise, various blocks shown in the flowcharts can be viewed as method steps, and / or as operations resulting from the operation of the computer program code, and / or as a plurality of built-in coupled logic circuit elements. to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product that includes a computer program tangibly realized on a machine-readable medium. 238481 1723267 of 24 computer, a computer program containing program codes configured to carry out the methods described above.

[0092] In the context of the disclosure, a computer-readable medium can be any tangible medium that can contain, or store a program for use by or in connection with a system, apparatus, or device to execute instructions. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium may be non-transient and may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of a computer-readable storage medium would include an electrical connection with one or more cables, a portable floppy disk, a hard disk, a RAM memory, a ROM memory, an erasable programmable read-only memory (EPROM, for its or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of they.

[0093] The computer program code for carrying out the methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus having control circuitry such that the program codes, when executed by the processor of the computer or other programmable data processing device, cause the functions / operations specified in the flow charts and / or block diagrams to be implemented. Program code may run entirely on one computer, partly on the computer as a stand-alone software package, partly on the computer and partly on a remote computer, or completely on the remote computer or server, or may be distributed among one or more remote computers. and / or servers.

[0094] While this document contains many implementation specifics, these should not be construed as limitations on the scope of what can be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described herein in the context of separate embodiments may also be implemented in 238481 1723267 of 24 combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable sub-combination. Furthermore, while some features may be described above as operating in certain combinations and even initially claimed as such, one or more features of a claimed combination may, in some cases, be eliminated from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination. The logical flows illustrated in the figures do not require the particular order shown, or sequential order, to achieve desirable results. Likewise, other steps may be provided, or removed, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other implementations fall within the scope of the following claims.

Claims

1. A method for encoding a low-frequency effects (LFE) channel, characterized in that it comprises: receiving, using one or more processors, a time-domain LFE channel signal; filtering, using a low-pass filter, the time-domain LFE channel signal to produce a filtered time-domain LFE channel signal, wherein the low-pass filter has a cutoff frequency; converting, using the one or more processors, the filtered time-domain LFE channel signal into a frequency-domain representation of the time-domain LFE channel signal that includes a number of coefficients representing a frequency spectrum of the time-domain LFE channel signal;arranging, using one or more processors, the coefficients in two or more subband groups corresponding to different frequency bands of the LFE channel signal in the time domain, wherein the different frequency bands include a main LFE frequency band that is below a cutoff frequency of an LFE speaker and at least one other LFE frequency band that is greater than the cutoff frequency of the LFE speaker, wherein each subband group has a width, and a sum of the widths of the subband groups includes the main LFE frequency band and the at least one other LFE frequency band; quantizing, using one or more processors, the coefficients of each subband group according to a frequency response curve of the low-pass filter to produce quantized coefficients;Encoding, using one or more processors, the quantized coefficients of each subband group using an entropy encoder tuned for the subband group; and generating, using one or more processors, a bitstream that includes the quantized coefficients. Claim 1 follows.