Low-latency, bass-enhancing codec
The low-latency LFE codec addresses latency and quality issues in IVAS codecs by filtering and entropy encoding LFE channels, achieving efficient low-latency LFE processing with high-quality audio and adaptable bitrates.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2025-12-17
- Publication Date
- 2026-04-10
AI Technical Summary
Existing IVAS codecs lack efficient methods for handling low-frequency effect (LFE) channels with low latency and high-quality error resilience, particularly in immersive audio applications.
A configurable low-latency LFE codec that processes LFE channels by filtering, converting to frequency-domain representation, quantizing coefficients based on a frequency response curve, and encoding with entropy encoders, while adjusting delay and using low-pass filters to maintain low latency and high audio quality.
The codec achieves low latency and high-quality LFE processing, supporting frequencies up to 400 Hz with varying bitrates, and reduces latency to 13-33 milliseconds, including silent mode operation with minimal bitrate.
Smart Images

Figure 2026062726000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-references to related applications This application claims priority to U.S. Provisional Patent Application No. 62 / 895,049, filed on 3 September 2019, and U.S. Provisional Patent Application No. 63 / 069,420, filed on 24 August 2020, respectively, which are incorporated herein by reference.
[0002] This disclosure generally relates to audio signal processing, and more particularly to processing of low-frequency effect (LFE) channels.
[0003] background Standardization of immersive services includes the development of Immersive Voice and Audio Service (IVAS) codecs for voice, multi-stream teleconferencing, virtual reality (VR), and user-generated live and non-live content streaming. The goal of the IVAS standard is to develop a single codec with excellent sound quality, low latency, support for spatial audio coding, a suitable bitrate range, high-quality error resilience, and practical implementation complexity. To achieve this goal, the development of IVAS codecs capable of handling low-latency LFE operation is desired for IVAS-enabled devices and other devices capable of processing LFE signals. The LFE channel targets deep low frequencies from 20 to 120 Hz and is typically sent to speakers designed to play low-frequency audio content. [Overview of the project]
[0004] summary Embodiments of a configurable low-latency LFE codec are disclosed.
[0005] In some embodiments, a method for encoding a low-frequency effect (LFE) channel includes the steps of: receiving a time-domain LFE channel signal using one or more processors; filtering the time-domain LFE channel signal using a low-pass filter; converting the filtered time-domain LFE channel signal into a frequency-domain representation of the LFE channel signal including a plurality of coefficients representing the frequency spectrum of the LFE channel signal using one or more processors; arranging the coefficients into a plurality of subband groups corresponding to different frequency bands of the LFE channel signal using one or more processors; quantizing the coefficients of each subband group according to the frequency response curve of the low-pass filter using one or more processors; encoding the quantized coefficients of each subband group using an entropy encoder tuned for each subband group using one or more processors; generating a bitstream containing the encoded quantized coefficients using one or more processors; and storing the bitstream in a storage device or streaming the bitstream to a downstream device using one or more processors.
[0006] In some embodiments, the step of quantizing the coefficients in each subband group further includes generating a scaling shift coefficient based on the maximum number of available quantization points and the sum of the absolute values of the coefficients, and quantizing the coefficients using the scaling shift coefficient.
[0007] In some embodiments, if a quantized coefficient exceeds the maximum number of quantization points, the scaling shift coefficient is reduced and the coefficient is quantized again.
[0008] In some embodiments, the quantization point differs for each subband group.
[0009] In some embodiments, the coefficients of each subband group are quantized according to a fine quantization scheme or a coarse quantization scheme, the fine quantization scheme assigns more quantization points to one or more subband groups than would be assigned to each subband group according to the coarse quantization scheme.
[0010] In some embodiments, the sign bit for the coefficient is encoded separately from the coefficient.
[0011] In some embodiments, there are four subband groups, where the first subband group corresponds to a first frequency range of 0 to 100 Hz, the second subband group corresponds to a second frequency range of 100 to 200 Hz, the third subband group corresponds to a third frequency range of 200 to 300 Hz, and the fourth subband group corresponds to a fourth frequency range of 300 to 400 Hz.
[0012] In some embodiments, the entropy encoder is an arithmetic entropy encoder.
[0013] In some embodiments, the step of converting the filtered time-domain LFE channel signal into a frequency-domain representation of the LFE channel signal including a plurality of coefficients representing the frequency spectrum of the LFE channel signal further includes determining a first stride length of the LFE channel signal; specifying a first window size of a window function based on the first stride length; applying the first window size to one or more frames of the time-domain LFE channel signal; and applying a modified discrete cosine transform (MDCT) to the windowed frames to generate the coefficients.
[0014] In some embodiments, the method further includes determining a second stride length of the LFE channel signal, specifying a second window size of the window function based on the second stride length, and applying the second window size to the one or more frames of the time domain LFE channel signal.
[0015] In some embodiments, the first stride length is N milliseconds (ms), where N is greater than or equal to 5 ms and less than or equal to 60 ms, the first window size is greater than or equal to 10 ms, the second stride length is 5 ms, and the second window size is 10 ms.
[0016] In some embodiments, the first stride length is 20 milliseconds (ms), the first window size is 10 ms, 20 ms, or 40 ms, the second stride length is 10 ms, and the second window size is 10 ms or 20 ms.
[0017] In some embodiments, the first stride length is 10 milliseconds (ms), the first window size is 10 ms or 20 ms, and the second stride length is 5 ms, and the second window size is 10 ms.
[0018] In some embodiments, the first stride length is 20 milliseconds (ms), the first window size is 10 ms, 20 ms, or 40 ms, the second stride length is 5 ms, and the second window size is 10 m.
[0019] In some embodiments, the window function is a Kaiser - Bessel - derived (KBD) window function having a configurable fade length.
[0020] In some embodiments, the low-pass filter is a fourth-order Butterworth filter low-pass filter with a cutoff frequency of approximately 130 Hz or less.
[0021] In some embodiments, the method further includes the steps of: using one or more processors to determine whether the energy level of a frame of the LFE channel signal is below a threshold; generating a silent frame indicator for the decoder in response to the energy level being below a threshold; inserting the silent frame indicator into the metadata of the LFE channel bitstream; and reducing the LFE channel bitrate when a silent frame is detected.
[0022] In some embodiments, a method for decoding a low-frequency effect (LFE) includes a method for decoding a low-frequency effect (LFE) channel bitstream, comprising: receiving an LFE channel bitstream using one or more processors, which includes entropy-encoded coefficients representing the frequency spectrum of a time-domain LFE channel signal; decoding the quantized coefficients using an entropy decoder using one or more processors; inversely quantizing the inversely quantized coefficients using one or more processors, wherein the coefficients are quantized in a group of subband groups corresponding to a group of frequency bands according to the frequency response curve of a low-pass filter used to filter the time-domain LFE channel signal in an encoder; converting the inversely quantized coefficients into a time-domain LFE channel signal using one or more processors; adjusting the delay of the time-domain LFE channel signal using one or more processors; and filtering the delay-adjusted LFE channel signal using a low-pass filter.
[0023] In some embodiments, the order of the low-pass filter is configured such that the first total algorithmic delay resulting from encoding and decoding the LFE channel is less than or equal to the second total algorithmic delay resulting from encoding and decoding the other audio channels of the multi-channel audio signal including the LFE channel signal.
[0024] In some embodiments, the method includes the steps of: determining whether the second total algorithm delay exceeds a threshold; configuring the low-pass filter as an N-th order low-pass filter, where N is an integer of 2 or more, depending on whether the second total algorithm delay exceeds the threshold; and setting the order of the low-pass filter to less than N, depending on whether the second total algorithm delay does not exceed the threshold. It further encompasses.
[0025] Other embodiments disclosed herein relate to systems, devices, and computer-readable media. Details of the disclosed embodiments are made clear in the accompanying drawings and the following description. Other features, purposes, and advantages are evident from the following description, drawings, and claims.
[0026] Certain embodiments disclosed herein offer one or more of the following advantages: The low-latency LFE codec of this disclosure 1) primarily targets the LFE channel, 2) primarily targets the frequency range of 20 to 120 Hz, but transmits audio up to 300 Hz in low / medium bitrate situations and up to 400 Hz in high bitrate situations, 3) achieves low bitrate by applying a quantization scheme corresponding to the frequency response curve of the input low-pass filter, 4) is designed to have low algorithmic latency, operating with a stride of 20 milliseconds (ms) and having a total algorithmic latency of 33 msec (including framing), and 5) has smaller strides to support other situations. It can be configured with a lower stride and algorithmic latency, including configurations with a stride of 5 msec and a total algorithmic latency (including framing) of 13 msec, 6) at the decoder output, an automatic low-pass filter is selected based on the latency obtained by the LFE codec, 7) there is a silent mode with a low bitrate of 50 bits / second (bps) when there is no sound, and 8) when there is an active frame, the bitrate varies between 2 kilobits / second (kbps) and 4 kbps depending on the quantization level used, and the bitrate becomes 50 bps when there is no sound. [Brief explanation of the drawing]
[0027] In the drawings, certain arrangements or orderings of graphic elements, such as elements representing devices, units, instruction blocks, and data elements, are shown for the sake of clarity. However, it should be understood by those skilled in the art that the particular ordering or arrangement of these graphic elements in the drawings is not intended to implicitly mean that a particular order or sequence of processing is required, nor that separation of processes is required. Furthermore, the inclusion of graphic elements in the drawings is not intended to implicitly mean that such elements are required in all embodiments, nor is it intended to implicitly mean that feature parts represented by such elements cannot be included in or combined with other elements in some embodiments.
[0028] Furthermore, where connecting elements such as solid or dashed lines or arrows are used in drawings to indicate connections, relationships, or associations between two or more other graphic elements, the absence of any such connecting element is not intended to implicitly mean that there is no possibility of such connections, relationships, or associations existing. In other words, some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the disclosure. In addition, for the sake of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, if a connecting element represents the communication of signals, data, or commands, it should be understood by those skilled in the art that such an element represents one or more signal paths as necessary to carry out the communication.
[0029] [Figure 1] Figure 1 shows an IVAS codec for encoding and decoding IVAS and LFE bitstreams in one or more embodiments.
[0030] [Figure 2A] Figure 2A is a block diagram showing LFE encoding in one or more embodiments.
[0031] [Figure 2B] Figure 2B is a block diagram showing LFE decoding in one or more embodiments.
[0032] [Figure 3] Figure 3 is a plot showing the frequency response of a fourth-order Butterworth low-pass filter with a 130 Hz corner cutoff in one or more embodiments.
[0033] [Figure 4] Figure 4 is a plot showing the Fielder window in one or more embodiments.
[0034] [Figure 5] Figure 5 shows the change in the fine quantization point with respect to frequency in one or more embodiments.
[0035] [Figure 6] Figure 6 shows the change in coarse quantization points with respect to frequency in one or more embodiments.
[0036] [Figure 7] Figure 7 shows the probability distribution of MDCT coefficients quantized by fine quantization in one or more embodiments.
[0037] [Figure 8] Figure 8 shows the probability distribution of MDCT coefficients quantized by coarse quantization in one or more embodiments.
[0038] [Figure 9] Figure 9 is a flowchart of the process for encoding modified discrete cosine transform (MDCT) coefficients in one or more embodiments.
[0039] [Figure 10]Figure 10 is a flowchart of the process for decoding modified discrete cosine transform (MDCT) coefficients in one or more embodiments.
[0040] [Figure 11] Figure 11 is a block diagram of a system 1100 for implementing the features and processes described with reference to Figures 1-10 in one or more embodiments.
[0041] The same reference symbols used in each drawing indicate similar elements. [Modes for carrying out the invention]
[0042] In the following detailed description, a great many specific details are given in order to provide a full understanding of the various embodiments described. It will be apparent to those skilled in the art that the various embodiments described can be carried out without these specific details. In other cases, known methods, procedures, components, and circuits are not described in detail so as not to unnecessarily obscure the aspects of the embodiments. Several features that can be used independently of each other or in any combination of other features are described below.
[0043] nomenclature The terms "include" and their variations as used herein The term "including, but not limited to" should be interpreted as an open-ended term. The term "or" should be interpreted as "and / or" unless the context clearly indicates another meaning. The term "based on" should be interpreted as "at least partially based on." The terms "one exemplary embodiment" and "one exemplary embodiment" should be interpreted as "at least one exemplary embodiment." The term "another embodiment" should be interpreted as "at least one other embodiment." The terms "determined," "determines," and "determining" should be interpreted as "to obtain," "to receive," It should be interpreted as “calculate,” “calculate,” “estimate,” “predict,” or “derive.” In addition, unless otherwise defined, all technical and scientific terms used herein in the following description and claims are in the art to which this disclosure belongs. It has the same meaning as generally understood by those skilled in the art.
[0044] System Overview Figure 1 shows an IVAS codec 100 for encoding and decoding an IVAS bitstream containing an LFE channel bitstream in one or more embodiments. The IVAS codec 100 receives N+1 channels of audio data 101 for encoding. The N channels of audio data 101 are input to a spatial analysis and downmix unit 102, and one LFE channel is input to an LFE channel encoding unit 105. The audio data 101 includes, but is not limited to, monaural signals, stereo signals, binaural signals, spatial audio signals (e.g., multi-channel spatial audio objects), first-order ambisonics (FoA), higher-order ambisonics (HoA), and other arbitrary audio data.
[0045] In some embodiments, the spatial analysis and downmix unit 102 is configured to implement Complex Advanced Coupling (CACPL) for analyzing / downmixing stereo audio data and / or Spatial Reconstruction (SPAR) for analyzing / downmixing FoA audio data. In other embodiments, the spatial analysis and downmix unit 102 implements other formats. The output of the spatial analysis and downmix unit 102 includes spatial metadata and 1 to N channel audio data. The spatial metadata is input to the spatial metadata encoding unit 104. The spatial metadata encoding unit 104 is configured to quantize and entropy encode the spatial metadata. In some embodiments, the quantization may include fine quantization, medium quantization, coarse quantization, and very coarse quantization strategies, and the entropy encoding may include Huffman or arithmetic coding.
[0046] The audio data channels 1 through N are input to the primary audio channel encoding unit 103. The primary audio channel encoding unit 103 encodes the audio data channels 1 through N into one or more enhanced voice It is configured to encode into services (EVS) bitstreams. In some embodiments, the primary audio channel encoding unit 103 conforms to 3GPP TS 26.445 and provides a wide range of functionality, including improved quality and encoding efficiency for narrowband (EVS-NB) and wideband (EVS-WB) voice services, improved quality with superwideband (EVS-SWB) voice, improved quality for mixed content and music in conversational applications, robustness to packet loss and delay jitter, and backward compatibility with the AMR-WB codec.
[0047] In some embodiments, the primary audio channel encoding unit 103 includes a preprocessing / mode selection unit. This preprocessing / mode selection unit selects between an audio encoder for encoding the audio signal and a perceptual encoder for encoding the audio signal at a specified bitrate, based on mode / bitrate control. In some embodiments, the audio encoder is an improved variation of algebraic code-excited linear prediction (ACELP), extended by dedicated LP-type modes for different audio classes.
[0048] In some embodiments, the audio encoder is a modified discrete cosine transform (MDCT) encoder with improved efficiency at low latency and low bitrate, and is designed to perform seamless and reliable switching between the speech encoder and the audio encoder. It is being done.
[0049] As mentioned above, the LFE channel signal targets deep low frequencies in the 20-120 Hz range and is typically sent to a speaker designed to play low-frequency audio content (e.g., a subwoofer). The LFE channel signal is input to an LFE channel signal encoding unit 105, which is configured to encode the LFE channel signal, as described with reference to Figure 2A.
[0050] In some embodiments, the IVAS decoder includes a spatial metadata decoding unit 106 configured to restore spatial metadata and a primary audio channel decoding unit 107 configured to restore 1 to N channel audio signals. The restored spatial metadata and restored 1 to N channel audio signals are input to a spatial synthesis / upmixing / rendering unit 109. This spatial synthesis / upmixing / rendering unit 109 is configured to use the spatial metadata to synthesize and render the 1 to N channel audio signals into N or more channel output audio signals for playback on speakers of various audio systems, including but not limited to home theater systems, video conferencing systems, virtual reality (VR) gear, and any other audio systems capable of rendering audio. The LFE channel decoding unit 108 is configured to receive an LFE bitstream and decode the LFE bitstream as described with reference to Figure 2B.
[0051] The above-mentioned LFE encoding / decoding implementation examples are performed by the IVAS codec, but the low-latency LFE codecs described below may be standalone LFE codecs or may be included in any proprietary or standard audio codecs that encode and decode low-frequency signals in audio applications where low latency and configurability are required or desired.
[0052] Figure 2A is a block diagram showing the functional components of the LFE channel encoding unit 105 shown in Figure 1 in one or more embodiments. Figure 2B is a block diagram showing the functional components of the LFE channel decoder 108 shown in Figure 1 in one or more embodiments. The LFE channel decoder 108 includes an entropy decoding / inverse quantization unit 204, an inverse MDCT / windowing unit 205, a delay adjustment unit 206, and an output LPF 207. The delay adjustment unit 206 may be located before or after the LPF 207 and performs delay adjustment (for example, by buffering the decoded LFE channel signal) to match the decoded LFE channel signal with the primary codec decode output. Hereinafter, the LFE channel encoding unit 105 and the LFE channel decoding unit 108, described with reference to Figure 2B, are collectively referred to as the LFE codec.
[0053] The LFE channel encoding unit 105 includes an input low-pass filter (LPF) 201, a windowing / MDCT unit 202, and a quantization and entropy coding unit 203. In one embodiment, the input audio signal is a pulse code modulated (PCM) audio signal, and the LFE channel encoding unit 105 expects an input audio signal with a stride of either 5 milliseconds, 10 milliseconds, or 20 milliseconds. Internally, the LFE channel encoding unit 105 operates in 5-millisecond or 10-millisecond subframes, and windowing and MDCT are performed in combinations of these subframes. In one embodiment, the LFE channel encoding unit 105 operates with a 20-millisecond input stride and internally divides this input into two subframes of equal length. The last subframe of the previous input frame to the LFE is LF The first subframe of the current input frame to E is concatenated and windowed. The first subframe of the current input frame to LFE is concatenated and windowed with the second subframe of the current input frame to LFE. MDCT is executed twice, once for each windowed block.
[0054] In one embodiment, the algorithmic delay (without framing delay) is equal to 8 milliseconds plus the delay caused by the input LPF103 and the delay caused by the output LPF207. Using a 4th-order input LPF201 and a 4th-order output LPF207, the total system latency is approximately 15 milliseconds. With a 4th-order input LPF201 and a 2nd-order output LPF207, the total LFE codec latency is approximately 13 milliseconds.
[0055] Figure 3 is a plot showing the frequency response of an exemplary input LPF201 in one or more embodiments. In the example shown, LPF201 is a fourth-order Butterworth filter with a cutoff frequency of 130 Hz. In other embodiments, different types of LPFs (e.g., Chebyshev, Bessel) with the same or different orders and the same or different cutoff frequencies may be used.
[0056] Figure 4 is a plot showing the Fielder window in one or more embodiments. In one embodiment, the windowing function applied by the windowing MDCT unit 202 is a Fielder window function with a fade length of 8 milliseconds. The Fielder window is a Kaiser-Bessel-derived (KBD) window with alpha=5, which structurally satisfies the Prince-Bradley condition of MDCT and is therefore used with the Advanced Audio Coding (AAC) digital audio format. Other window functions are also available.
[0057] Quantization and Entropy Coding In one embodiment, the quantization / entropy encoding unit 203 executes a quantization strategy according to the frequency response curve of the input LPF 201 to more efficiently quantize the MDCT coefficients. In one embodiment, the frequency range is divided into four sub-band groups representing four frequency bands, namely 0 to 100 Hz, 100 to 200 Hz, 200 to 300 Hz, and 300 to 400 Hz. These bands are examples, and more or fewer bands with the same or different frequency ranges can be used. More specifically, as shown in FIGS. 5 to 8, the MDCT coefficients are quantized using scaling shift coefficients dynamically calculated based on the MDCT coefficient values in a specific frame, and quantization points are selected as per the LPF frequency response curve. This quantization strategy helps reduce the quantization points of the MDCT coefficients belonging to the 100 to 200 Hz, 200 to 300 Hz, and 300 to 400 Hz bands, while on the other hand, maintaining the optimal quantization points for the 0 to 100 Hz primary LFE band where most of the energy of low bass effects (such as rumbling) is found.
[0058] In one embodiment, the quantization strategy for the F len millisecond (ms) input PCM stride (input frame length) to the LFE channel encoding unit 105 is described below. The frame length F len can take any value given by 5*f ms, where 1 <= f <= 12.
[0059] First, the input PCM stride is divided into N sub-frames of equal length, and each sub-frame width (S w ) = F len / N ms. N needs to be selected such that each S w is a multiple of 5 ms (for example, when F len = 20 ms, N is 1, 2, or 4; when F len = 10 ms, N is 1 or 2; when F len = 5 ms, N is equal to 1). S i Let S0 be the i-th subframe in a given frame, where i is an integer in the range 0 <= i <= N, and S0 corresponds to the last subframe of the previous input frame to the LFE encoding unit 105, and S1 through S N These are the N subframes of the current frame.
[0060] Next, each S i and S i+1 Subframes are concatenated and windowed using a Fielder window (see Figure 4), and MDCT is performed on these windowed samples. As a result, a total of N MDCTs are obtained for each frame. The number of MDCT coefficients (num_coeffs) for each MDCT is equal to the sampling frequency × S. w This becomes / 1000. Frequency resolution of each MDCT (width of each MDCT coefficient) (W mdct ) is approximately 1000 / (2×S w The total number of MDCT coefficients to be quantized and encoded is (N)Hz. Subwoofers typically have an LPF cutoff of around 100-120Hz, and the energy after the LPF above 400Hz is typically very small. Therefore, the MDCT coefficients up to 400Hz are quantized and sent to the LFE decoding unit 108, and the remaining MDCT coefficients are quantized to 0. Sending the MDCT coefficients up to 400Hz ensures high-quality reconstruction up to 120Hz in the LFE decoding unit 108. Thus, the total number of MDCT coefficients to be quantized and encoded is (N)Hz. quant ) is N×400 / W mdct It becomes equal to.
[0061] Next, the MDCT coefficient is calculated based on the width of each subband group, W. mdct The number of subbands is a multiple of SBW, and the M subbands are arranged such that the sum of the widths of all subbands equals 400Hz. The width of each subband is SBW m Let Hz be the frequency, and m be an integer within the range 1 <= m <= M. In this bandwidth, the number of coefficients in the m-th subband group = S / N quant =N×SBW m / W mdct (i.e., from each MDCT to SBW) m / W mdctThe coefficients are (number of coefficients). And the MDCT coefficient for each subband group is N quant The MDCT coefficients are scaled by a shift scaling coefficient (shift), which is determined by the sum or maximum of the absolute values of all the MDCT coefficients, as described below. Next, the scaled MDCT coefficients of each subband group are quantized and encoded separately using a quantization scheme that follows the LPF curve of the encoder input. The encoding of the quantized MDCT coefficients is performed using an entropy encoder (e.g., an arithmetic encoder or a Huffman encoder). Each subband group is encoded with a different entropy encoder, and each entropy encoder efficiently encodes its respective subband group using an appropriate probability distribution model.
[0062] 20 milliseconds (ms) stride (F len An example of a quantization strategy with a subframe width S = 20ms, 2 subframes (N=2), and a sampling frequency of 48000 is described. In this example's input configuration, the subframe width S w =10ms, number of MDCTs = N=2. The first MDCT is performed on a 20ms block. This block is formed by concatenating the 10-20ms subframes of the previous 20ms input and the 0-10ms subframes of the current 20ms input, and windowing them in a 20ms-long Fielder window (see Figure 4). For N=1 and N=4, the Fielder window is scaled appropriately, and the fade length is changed to 16 / Nms. The second MDCT is performed on a 20ms block formed by windowing the current 20ms input frame in a 20ms-long Fielder window. The number of MDCT coefficients (num_coeffs) from each MDCT is 480, and the width of each MDCT coefficient is W. mdct = 50Hz, total number of coefficients N to quantize and encode. quant =16, and the total number of coefficients to be quantized and coded for each MDCT was set to 16 / N = 8.
[0063] Next, the MDCT coefficients are placed into four subband groups (M=4). Each subband group corresponds to a 100Hz bandwidth (0-100, 100-200, 200-300, 300-400, SBW). m =100Hz, number of coefficients for each subband group =S / N quant =N×SBW m / W mdct =4). The first eight MDCT coefficients, b1, b2, b3, b4, quantize a1, a2, a3, a4, a5, a6, a7, a8 from the first MDCT. Let b5, b6, b7, and b8 be the first eight MDCTs quantized from the second MDCT. The four subband groups are arranged to have the following coefficients: Subband group 1 = {a1, a2, b1, b2} Subband group 2 = {a3, a4, b3, b4} Subband group 3 = {a5, a6, b5, b6} Subband group 4 = {a7, a8, b7, b8} Here, each subband group corresponds to a bandwidth of 100Hz.
[0064] In frames with a gain of approximately -30dB (or less), 10 -2 Mokushiha 10 -1 MDCT coefficients can have values of a certain degree or less, but frames with full-scale gain can have MDCT coefficients of 20 or more. To accommodate this wide range of values, the scaling shift coefficient (shift) is calculated based on the maximum number of available quantization points (max_value) and the sum of the absolute values of the MDCT coefficients (lfe_dct_new), as follows: shift=floor(shifts_per_double×log 2 (max_value / sum(abs(lfe_dct_new))))
[0065] In one embodiment, lfe_dct_new is an array of 16 MDCT coefficients, shifts_per_double is a constant (e.g., 4), max_value is an integer selected for fine quantization (e.g., 63 quantization values) and coarse quantization (e.g., 31 quantization values), and shift is limited to a 5-bit value from 4 to 35 for fine quantization and from 2 to 33 for coarse quantization.
[0066] Next, the quantized MDCT coefficients are calculated as follows. vals=round(lfe_dct_new×(2^(shift / shifts_per_double))) The round() operation here rounds the result to the nearest integer value.
[0067] If the quantized values (vals) exceed the maximum allowable number of quantization points (max_val), the scale shift coefficient (shift) is reduced and the quantized values (vals) are recalculated. In other embodiments, the scaling shift coefficient (shift) can be calculated using the max function max(abs(lfe_dct_new))) instead of the sum function sum(abs(lfe_dct_new))), but using the max() function results in more scattered quantized values, making it difficult to design an efficient entropy encoder.
[0068] In the quantization step described above, the quantized values of each subband group are calculated together in one loop, but the quantization point is different for each subband group. If the first subband group exceeds the acceptable range, the scaling shift coefficient is reduced. If any of the other subband groups exceeds the acceptable range, that subband group is truncated to max_value. The sign bit for all MDCT coefficients and the absolute value of the quantized MDCT coefficients are encoded separately for each subband group.
[0069] Figure 5 shows the change in the number of fine quantization points with frequency in one or more embodiments. In fine quantization, subband group 1 (0-100 Hz) has 64 quantization points, subband group 2 (100-200 Hz) has 32 quantization points, subband group 3 (200-300 Hz) has 8 quantization points, and subband group 4 (300-400 Hz) has 2 quantization points. In one embodiment, each subband group is entropy coded with a separate entropy encoder (e.g., an arithmetic encoder or a Huffman entropy encoder), and each entropy encoder uses a different probability distribution. Therefore, 0-1 The 00Hz primary range is assigned the most quantization points.
[0070] The assignment of quantization points to subband groups 1-4 follows the shape of the LPF frequency response curve, where there is more low-frequency information than high-frequency information and no information outside the cutoff frequency. To correctly reconstruct frequencies up to 130Hz, MDCT coefficients corresponding to frequencies above 130Hz are also encoded to avoid or minimize aliasing. In some embodiments, MDCT coefficients up to 400Hz are encoded so that frequencies up to 130Hz can be properly reconstructed by the decoding unit.
[0071] Figure 6 shows the change in coarse quantization points with frequency in one or more embodiments. In coarse quantization, subband group 1 (0-100 Hz) has 32 quantization points, subband group 2 (100-200 Hz) has 16 quantization points, subband group 3 (200-300 Hz) has 4 quantization points, and subband group 4 (300-400 Hz) is not quantized or entropy coded. In one embodiment, each subband group is entropy coded using separate entropy encoders with different probability distributions.
[0072] Figure 7 shows the probability distribution of quantized MDCT coefficients by fine quantization in one or more embodiments. The y-axis represents the frequency of occurrence, and the x-axis represents the number of quantization points. Sg1 is subband group 1 corresponding to quantized MDCT coefficients in the 0-100 Hz band, Sg2 is subband group 2 corresponding to quantized MDCT coefficients in the 100-200 Hz band, Sg3 is subband group 3 corresponding to quantized MDCT coefficients in the 200-300 Hz band, and Sg4 is subband group 4 corresponding to quantized MDCT coefficients in the 300-400 Hz band.
[0073] Figure 8 shows the probability distribution of quantized MDCT coefficients by coarse quantization in one or more embodiments. The y-axis represents the frequency of occurrence, and the x-axis represents the number of quantization points. Sg1 is subband group 1 corresponding to the quantized MDCT coefficients in the 0-100 Hz band, and Sg2 is subband group 2 corresponding to the quantized MDCT coefficients in the 100-200 Hz band. Sg3 is subband group 3 corresponding to the quantized MDCT coefficients in the 200-300 Hz band. Sg4 is subband group 4 corresponding to the quantized MDCT coefficients in the 300-400 Hz band.
[0074] Furthermore, since the LFE effect is most prevalent in the primary bandwidth (0-100Hz), more quantization points are allocated to increase resolution. However, in coarse quantization, fewer bits are allocated to the primary bandwidth than in fine quantization. In one embodiment, whether fine or coarse quantization is used for the MDCT coefficients for one frame depends on the desired target bitrate set by the primary audio channel encoder 103. The primary audio channel encoder 103 sets this value once during initialization, or dynamically on a frame-by-frame basis based on the bits required or used to encode the primary audio channel in each frame.
[0075] Silent frame In some embodiments, a signal is added to the LFE channel bitstream to indicate silent frames. A silent frame is a frame with energy below a specified threshold. In some embodiments, to indicate a silent frame, one bit is included in the LFE channel bitstream sent to the decoder (for example, inserted in the frame header), and all MDCT coefficients in the LFE channel bitstream are set to 0. This technique can reduce the bitrate to 50 bps during silent frames.
[0076] Decoder LPF Two options for implementing LPF207 (see Figure 2B) are provided at the output of LFE channel decoding unit 108. LPF207 is selected based on the available delay (total delay of other audio channels minus LFE phasing delay minus input LPF delay). Note that the other channels are expected to be encoded / decoded by primary audio channel encoding / decoding units 103 and 107, and the delays of those channels depend on the algorithmic delays of the primary audio channel encoding / decoding units 103 and 107.
[0077] In one embodiment, if the available delay is less than 3.5 ms, a second-order Butterworth LPF with a cutoff at 130 Hz is used; otherwise, a fourth-order Butterworth LPF with a cutoff at 130 Hz is used. Thus, in the LFE channel decoding unit 108, there is a trade-off between the removal of aliasing energy above the cutoff frequency and the algorithmic delay. In some embodiments, the subwoofer typically has an LPF, so the LPF 207 can be completely removed. The LPF 207 can help reduce aliasing energy above the cutoff frequency at the LFE decoder output itself, which can contribute to efficient post-processing.
[0078] Process example Figure 9 is a flowchart of the process 900 for encoding MDCT coefficients in one or more embodiments. The process 900 can be implemented, for example, using the system 1100 described with reference to Figure 11.
[0079] Process 900 includes the following steps: receiving a time-domain LFE channel signal (901); filtering the time-domain LFE channel signal using a low-pass filter (902); converting the filtered time-domain LFE channel signal into a frequency-domain representation of the LFE channel signal, which includes a plurality of coefficients representing the frequency spectrum of the LFE channel signal (903); arranging the coefficients into a plurality of subband groups corresponding to different frequency bands of the LFE channel signal (904); quantizing the coefficients of each subband group according to the frequency response curve of the low-pass filter using a scaling-shift coefficient (905); encoding the quantized coefficients of each subband group using an entropy encoder configured for the subband group (906); generating a bitstream containing the encoded quantized coefficients (907); and storing the bitstream in a memory device or streaming the bitstream to a downstream device (908).
[0080] Figure 10 is a flowchart of a process 1000 for decoding MDCT coefficients in one or more embodiments. Process 1000 can be implemented, for example, using system 1100 as described with reference to Figure 11.
[0081] Process 1000 includes the following steps: receiving an LFE channel bitstream, the LFE channel bitstream including entropy-encoded coefficients representing the frequency spectrum of a time-domain LFE channel signal (1001); decoding and inversely quantizing the coefficients, the coefficients being quantized into subband groups corresponding to different frequency bands according to the frequency response curve of a low-pass filter using scaling-shift coefficients (1002); converting the decoded and inversely quantized coefficients into a time-domain LFE channel signal (1003); adjusting the delay of the time-domain LFE channel signal (1004); and filtering the delay-adjusted LFE channel signal using a low-pass filter (10 05). In one embodiment, the order of the low-pass filter may be set based on the total algorithmic delay obtained from the primary codec used to encode / decode the full-bandwidth channels of a multi-channel audio signal, including the time-domain LFE channel signal. In some embodiments, the decoding unit only needs to know whether the MDCT coefficients were encoded with fine quantization or coarse quantization by the encoding unit. The type of quantization can be indicated by a bit in the LFE bitstream header or by another suitable signaling mechanism.
[0082] In some embodiments, decoding the inversely quantized coefficients into time-domain PCM samples is performed as follows: The inversely quantized coefficients of each subband group are rearranged into N groups (where N is the number of MDCTs calculated in the encoding unit), and each group has coefficients corresponding to its respective MDCT. As in the implementation example described above, the encoding unit encodes the following four subband groups. Subband group 1 = {a1, a2, b1, b2} Subband group 2 = {a3, a4, b3, b4} Subband group 3 = {a5, a6, b5, b6} Subband group 4 = {a7, a8, b7, b8}
[0083] The decoding unit decodes four subband groups and rearranges them into {a1, a2, a3, a4, a5, a6, a7, a8} and {b1, b2, b3, b4, b5, b6, b7, b8}, and pads these groups with zeros to obtain the desired inverse MDCT (iMDCT) input length. N iMDCT operations are performed to inversely transform the MDCT coefficients of each group into time-domain blocks. In this example, each block is 2 × Swms wide, where S w This is the subframe width defined above. Next, this block is windowed using the same Fielder window used in the LFE encoding unit shown in Figure 4. Each subframe S i (where i is an integer between 1 and N) is reconstructed by appropriately overlapping and adding the windowed data of the previous iMDCT output and the current iMDCT output. Finally, the output of (1003) is reconstructed by concatenating all N subframes.
[0084] System Architecture Example Figure 11 is a block diagram of System 1100 for implementing the features and processes described with reference to Figures 1-10 in one or more embodiments. System 1100 includes one or more server computers or any client devices, including but not limited to: call servers, user equipment, conference room systems, home theater systems, virtual reality (VR) gear, and content ingestion devices. System 1100 also includes, but is not limited to: any consumer equipment, including: smartphones, tablet computers, wearable computers, vehicle computers, game consoles, surround sound systems, kiosks, etc.
[0085] As shown in the figure, the system 1100 includes a central processing unit (CPU) 1101 capable of executing various processes according to, for example, a program stored in read-only memory (ROM) 1102, or a program loaded from a memory unit 1108 into random access memory (RAM) 1103. The RAM 1103 also stores data required when the CPU 1101 executes various processes, as needed. The CPU 1101, ROM 1102, and RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0086] The following components, namely an input unit 806 which may include a keyboard, mouse, etc., an output unit 807 which may include a display such as a liquid crystal display (LCD) and one or more speakers, a storage unit 1108 which includes a hard disk or another suitable storage device, and a communication unit 1109 which includes a network interface card such as a network card (e.g., wired or wireless), are connected to the I / O interface 1105.
[0087] In some embodiments, the input unit 1106 includes one or more microphones located at different positions (depending on the host device) that enable the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0088] In some embodiments, the output unit 1107 includes a system having a varying number of speakers. The output unit 1107 can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats) (depending on the capabilities of the host device).
[0089] The communication unit 1109 is configured to communicate with other devices (for example, via a network). The drive 810 is also connected to the I / O interface 1105 as needed. Removable media 1111, such as a magnetic disk, optical disk, magneto-optical disk, flash drive, or other suitable removable media, is mounted on the drive 1110 so that computer programs read from there are installed in the storage unit 1108 as needed. Those skilled in the art will understand that although the system 1100 is described as including the components described above, in actual use it is possible to add, remove, and / or replace some of these components, and all such changes or modifications are all within the scope of this disclosure.
[0090] According to exemplary embodiments of the present disclosure, the processes described above can be implemented as a computer software program or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product which includes a computer program tangibly embodied on a machine-readable medium, the computer program which includes program code that performs the method. In such embodiments, the computer program can be downloaded and implemented from a network via a communication unit 1309 and / or installed from removable media 1111.
[0091] In general, various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry (e.g., control circuits), software, logic, or any combination thereof. For example, the unit described above can be executed by control circuits (e.g., a CPU in combination with the other components in Figure 11), and thus these control circuits can perform the operations described in this disclosure. Some embodiments can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device (e.g., control circuits). Various embodiments of the exemplary embodiments of this disclosure are illustrated and described using block diagrams, flowcharts, or other graphic representations, but it will be understood that any blocks, apparatus, systems, techniques, or methods described herein can be implemented, in non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or any combination thereof.
[0092] In addition, the various blocks shown in the flowchart can be considered as a plurality of coupled logic circuit elements configured to perform method steps and / or operations and / or associated functions (which may be multiple) resulting from the operation of the computer program code. For example, embodiments of the present disclosure include a computer program product which includes a computer program tangibly embodied on a machine-readable medium, the computer program which includes program code configured to perform the methods described above.
[0093] In the context of this disclosure, a machine / computer-readable medium can be any tangible medium capable of containing or storing programs used by or in connection with an instruction execution system, instruction execution unit, or instruction execution device. A machine / computer-readable medium may be a machine / computer-readable signal medium or a machine / computer-readable storage medium. A machine / computer-readable medium may be non-temporary and may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. More specific examples of machine / computer-readable storage media include electrical connections with one or more wires, portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0094] Computer program code for performing the methods of this disclosure can be written in any combination of one or more programming languages. This computer program code can be provided to a general-purpose computer, a dedicated computer, or a processor of another programmable data processing device having control circuits, such that when the program code is executed by the processor of the computer or other programmable data processing device, it causes the execution of functions / operations specified in flowcharts and / or block diagrams. The program code can be executed as a standalone software package, either entirely or partially on a computer, partially on a computer and partially on a remote computer, entirely on a remote computer or remote server, or distributed across one or more remote computers and / or remote servers.
[0095] This specification includes many specific details of implementation, which should not be construed as limitations on the scope of what can be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments and in any suitable subcombinations. Furthermore, features described above as operating in a particular combination may even be initially claimed as such, but in some cases one or more features from a claimed combination may be removed from that combination, and the claimed combination may cover subcombinations or variations of subcombinations. The logical flow shown in the figures does not require any specific order or sequential order to achieve the desired result. In addition, other steps may be added to the described flow, steps may be removed, and other components may be added to or removed from the described system. Thus, other embodiments are within the scope of the appended claims.
Claims
[Claim 1] A method for encoding a low-frequency effect (LFE) channel, The steps include receiving a time-domain LFE channel signal using one or more processors, A step of filtering the time-domain LFE channel signal using a low-pass filter to generate a filtered time-domain LFE channel signal, wherein the low-pass filter has a cutoff frequency. The steps include: using one or more processors, converting the filtered time-domain LFE channel signal into a frequency-domain representation of the LFE channel signal that includes a plurality of coefficients representing the frequency spectrum of the LFE channel signal; A step of using one or more processors to arrange the coefficients into two or more subband groups corresponding to different frequency bands of the LFE channel signal, wherein the different frequency bands include a primary LFE frequency band lower than the cutoff frequency of the low-pass filter and at least one other LFE frequency band higher than the cutoff frequency of the low-pass filter, where each subband group has a width, and the sum of the widths of the subband groups includes the primary LFE frequency band and the at least one other LFE frequency band. A step of generating quantized coefficients by quantizing the coefficients of each subband group according to the frequency response curve of the low-pass filter using one or more processors, A scaling shift coefficient is generated based on the maximum number of available quantization points and the sum of the absolute values of the coefficients, Quantizing the coefficient using the scaling shift coefficient, Steps including, The steps include using one or more processors to encode the quantized coefficients of each subband group using an entropy encoder adjusted for each subband group, The steps include generating a bitstream containing the encoded quantized coefficients using one or more of the above-mentioned processors, It includes, A method in which the allocation of quantization points follows the shape of the frequency response curve of the low-pass filter, with more quantization points allocated to lower frequency subband groups than to higher frequency subband groups.