Low Bitrate Scene-Based Audio Coding

The integration of Spatial Reconstruction (SPAR) and Directional Audio Coding (DirAC) into a single codec, utilizing quantization, covariance smoothing and decoder-side decorrelation, achieves high audio quality at low bit rates.

JP2026500454APending Publication Date: 2026-01-07DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025522504
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-15
Filing Date
2023-09-29
Publication Date
2026-01-07

AI Technical Summary

Technical Problem

Existing technologies have not effectively addressed the need to combine the complementary aspects of Spatial Reconstruction (SPAR) and Directional Audio Coding (DirAC) for low bit rate spatial audio coding, particularly in the context of Ambisonics, to provide high audio quality at low bit rates.

Method used

The integration of Spatial Reconstruction (SPAR) and Directional Audio Coding (DirAC) into a single codec, utilizing quantization, covariance smoothing and decoder-side decorrelation, to achieve high audio quality at low bit rates.

Benefits of technology

The integration of Spatial Reconstruction (SPAR) and Directional Audio Coding (DirAC) into a single codec, utilizing quantization, covariance smoothing and decoder-side decorrelation, to achieve high audio quality at low bit rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026500454000001_ABST
    Figure 2026500454000001_ABST
Patent Text Reader

Abstract

Embodiments are included for very low bitrate scene-based audio (LBRSBA) coding that combines SPAR and DIRAC. In some embodiments, a method includes receiving scene-based audio metadata, creating spatial reconstruction (SPAR) metadata and directional audio coding (DirAC) metadata from the scene-based audio metadata, forming groups of SPAR metadata bands and groups of DirAC metadata bands, quantizing the groups of SPAR metadata bands and the groups of DirAC metadata bands, and transmitting to a decoder a first data frame that includes the quantized groups of DirAC metadata bands and a first portion of the quantized groups of SPAR metadata bands, and a second data frame that follows the first data frame and includes a second portion of the quantized DirAC metadata bands and the quantized groups of SPAR metadata bands.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Application No. 63 / 421,045, filed October 31, 2023, and U.S. Provisional Application No. 63 / 582,950, filed September 15, 2022, all of which are incorporated herein by reference in their entireties.

[0002] [Technical field] The present disclosure relates generally to audio processing. [Background technology]

[0003] Spatial Reconstruction (SPAR) and Directional Audio Coding (DirAC) are distinct spatial audio coding techniques that each seek to represent an input spatial audio scene in a compact way to enable transmission with a good tradeoff between audio quality and bit rate. One such input format for a spatial audio scene is a scene-based audio representation (e.g., first-order Ambisonics (FOA) or higher-order Ambisonics (HOA)).

[0004] SPAR seeks to maximize perceived audio quality while minimizing bitrate by reducing the energy of the transmitted audio data, while still allowing the second-order statistics (i.e., covariance) of the Ambisonics audio scene to be reconstructed at the decoder side using the transmitted metadata. SPAR seeks to faithfully reconstruct the input Ambisonics scene at the decoder's output.

[0005] DirAC is a technique for representing a spatial audio scene as a set of directions of arrival (DOA) in time-frequency tiles. From this representation, similar sound scenes can be reproduced in different output formats (e.g., binaural). In particular, in the context of Ambisonics, the DirAC representation allows a decoder to generate a high-order output from a low-order input. DirAC seeks to preserve the direction and diffuseness of the dominant sounds in the input scene.

[0006] DirAC and SPAR have different strengths and characteristics, and therefore it is desirable to combine the complementary aspects of DirAC and SPAR (e.g., higher audio quality, reduced bit rate, input / output format flexibility, and / or reduced computational complexity) into a coder / decoder ("codec") such as an Ambisonic codec. Summary of the Invention

[0007] An embodiment is included for low bitrate scene-based audio (LBRSBA) coding using SPAR and DirAC.

[0008] In some embodiments, a method for audio metadata encoding includes receiving, with at least one processor, scene-based audio metadata; generating, with the at least one processor, Spatial Reconstruction (SPAR) metadata and Directional Audio Coding (DirAC) metadata from the scene-based audio metadata; forming, with the at least one processor, groups of SPAR metadata bands and groups of DirAC metadata bands; quantizing, with the at least one processor, the groups of SPAR metadata bands and the groups of DirAC metadata bands; and transmitting to a decoder a first data frame including a first portion of the quantized groups of DirAC metadata bands and the quantized groups of SPAR metadata bands; and a second data frame subsequent to the first data frame, the second data frame including a second portion of the quantized groups of DirAC metadata bands and the quantized groups of SPAR metadata bands.

[0009] In some embodiments, the method further comprises transmitting a signal indicative of the first data frame or the second data frame to the decoder.

[0010] In some embodiments, the group of SPAR metadata bands includes four SPAR metadata bands, the group of DirAC bands includes two DirAC metadata bands, and the group of SPAR bands is lower in frequency than the group of DirAC bands.

[0011] In some embodiments, the group of DirAC metadata bands is transmitted to the decoder at a first temporal resolution, and the first and second portions of the group of SPAR metadata bands are transmitted to the decoder at a second temporal resolution, the second temporal resolution being lower than the first temporal resolution.

[0012] In some embodiments, when the first data frame is an initial data frame or when the group of DirAC metadata bands is encoded within the metadata bitrate budget, the group of DirAC metadata bands is transmitted to the decoder at a first temporal resolution.

[0013] In some embodiments, when a group of SPAR metadata bands is not encoded within the metadata bitrate budget, the group of SPAR metadata bands is transmitted to the decoder at a second temporal resolution.

[0014] In some embodiments, the method further comprises, prior to receiving the scene-based audio metadata, applying, with at least one processor, smoothing to the covariance matrix from which the scene-based audio metadata is formed.

[0015] In some embodiments, the covariance smoothing uses a smoothing factor that increases smoothing in low frequency bands and avoids modifying the amount of smoothing in high frequency bands.

[0016] In some embodiments, the smoothing factor is given by the function smoothing_factor(b)=update_factor(b) / min_pool_size*k*(b+1), where update_factor(b) is the number of frequency bins in frequency band b, min_pool_size is the minimum number of frequency bins desired, and k is a factor that increases or decreases the smoothing.

[0017] In some embodiments, a method of audio metadata decoding includes receiving, using at least one processor, quantized scene-based audio data and corresponding metadata, the metadata including decorrelator coefficients; dequantizing, using the at least one processor, the quantized scene-based audio data and corresponding metadata; decoding, using the at least one processor, the scene-based audio data and corresponding metadata, the decoding including recovering the decorrelator coefficients; smoothing, using the at least one processor, the decorrelator coefficients; and reconstructing, using the at least one processor, a multi-channel audio signal based on at least the decoded scene-based audio data and the smoothed decorrelator coefficients.

[0018] Other embodiments disclosed herein are directed to systems, devices, and computer-readable media. Details of the disclosed embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the detailed description, drawings, and claims.

[0019] Certain embodiments disclosed herein combine complementary aspects of DirAC and SPAR technologies into a single codec that provides high audio quality at low bit rates for scene-based audio (e.g., Ambisonics). [Brief explanation of the drawings]

[0020] [Figure 1] FIG. 1 is a block diagram of an IVAS codec framework according to one or more embodiments. [Figure 2] FIG. 1 is a flow diagram of a covariance smoothing process in accordance with one or more embodiments. [Figure 3] 3 is a flow diagram of an example modification to the process shown in FIG. 2 for a maximum allowable forgetting factor, according to one or more embodiments. [Figure 4] FIG. 10 is a flow diagram of an example modification to a transient detection process flow, according to one or more embodiments. [Figure 5] 1 is a plot of decorrelation coefficients over n frames in accordance with one or more embodiments. [Figure 6] 1 is a flow diagram of LBRSBA (e.g., Ambisonics) processing in accordance with one or more embodiments. [Figure 7] FIG. 7 is a block diagram of an exemplary hardware architecture suitable for implementing the systems and methods described with reference to FIGS. 1-6.

[0021] In the drawings, a particular arrangement or order of schematic elements, such as those representing devices, units, instruction blocks, and data elements, is shown for ease of explanation. However, it should be understood by those skilled in the art that the particular order or arrangement of schematic elements in the drawings is not intended to imply that a particular order or sequence of processing, or separation of processes, is required. Furthermore, the inclusion of a schematic element in a drawing does not mean to imply that such element is required in all embodiments, or that features represented by such element may not be included in or combined with other elements in some implementations.

[0022] Furthermore, when a connecting element, such as a solid or dashed line or arrow, is used in the drawings to indicate a connection, relationship, or association between two or more other schematic elements, the absence of such a connecting element does not imply that the connection, relationship, or association may not exist. In other words, some connections, relationships, or associations between elements may not be shown in the drawings so as not to obscure the present disclosure. Furthermore, for ease of explanation, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, when a connecting element represents communication of signals, data, or instructions, it should be understood by those skilled in the art that such element represents one or more signal paths for affecting the communication, as appropriate.

[0023] The use of the same reference numbers in the various drawings indicates similar elements. DETAILED DESCRIPTION OF THE INVENTION

[0024] In the following detailed description, numerous specific details are set forth to provide a thorough understanding of various described embodiments. It will be apparent to those skilled in the art that various described implementations may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. The following describes several features that can be used independently of each other or in any combination with other features.

[0025] [term] As used herein, the term "comprises" and variations thereof shall be read as an open-ended term meaning "including, but not limited to." The term "or" shall be read as "and / or" unless the context clearly dictates otherwise. The term "based on" shall be read as "based at least in part on." The terms "one exemplary implementation" and "exemplary implementation" shall be read as "at least one exemplary implementation." The terms "determined," "determine," or "determining" shall be read as obtaining, receiving, calculating, computing, estimating, predicting, or deriving. Furthermore, in the following detailed description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0026] [Exemplary IVAS Codec Framework] 1 is a block diagram of an immersive voice and audio services (IVAS) coder / decoder ("codec") framework 100 for encoding and decoding IVAS bitstreams, according to one or more embodiments. IVAS is expected to support a range of audio service capabilities, including, but not limited to, mono-to-stereo upmixing and fully immersive audio encoding, decoding, and rendering. IVAS is also intended to be supported by a wide range of devices, endpoints, and network nodes, including, but not limited to, mobile and smartphones, electronic tablets, personal computers, conference phones, conference rooms, virtual reality (VR) and augmented reality (AR) devices, home theater devices, and other suitable devices.

[0027] The IVAS codec 100 includes an IVAS encoder 101 and an IVAS decoder 104. The IVAS encoder 101 includes a spatial encoder 102 that receives N channels of input spatial audio (e.g., FOA, HOA). In some implementations, the spatial encoder 102 implements SPAR and DirAC to analyze / downmix the N_dmx spatial audio channels, as described in more detail below. The output of the spatial encoder 102 includes a spatial metadata (MD) bitstream (BS) and the N_dmx channels of the spatial downmix. The spatial MD is quantized and entropy coded. In some implementations, the quantization can include fine, medium, coarse, and ultra-coarse quantization strategies, and the entropy coding can include Huffman or arithmetic coding. In some implementations, the spatial encoder allows for three or fewer levels of quantization in a given operating mode, but as the bitrate decreases, the three levels become increasingly coarse overall to meet the bitrate requirements. The core audio encoder 103 (e.g., a Single Channel Element (SCE) coding unit) encodes the N_dmx channels (N_dmx=1 to 16 channels) of the spatial downmix into an audio bitstream, which is then combined with the spatial MD bitstream to form an IVAS-encoded bitstream that is transmitted to the IVAS decoder 104. For LBRSBA, the number of spatial downmix channels is limited to one due to bitrate constraints.

[0028] The IVAS decoder 104 includes a core audio decoder 105 (e.g., a Single Channel Element (SCE)) that decodes the audio bitstream extracted from the IVAS bitstream to recover the N_dmx audio channels. The spatial decoder / renderer 106 (e.g., SPAR / DirAC) decodes the spatial MD bitstream extracted from the IVAS bitstream to recover the spatial MD and synthesizes / renders output audio channels using the spatial MD and spatial upmix for playback on various audio systems with different speaker configurations and capabilities.

[0029] [Low Bitrate SBA (LBRSBA, Low Bitrate SBA)] In some embodiments, it is desirable to implement LBRSBA (e.g., Ambisonics) using the SPAR-DIRAC codec, which can be achieved using one or more of: 1) reduced MD bitrate and band interleaving, 2) additional covariance smoothing to facilitate reduced MD bitrate, and 3) decoder-side decorrelator coefficient smoothing.

[0030] Background information for the above techniques can be found in one or more of the following documents: PCT Application No. 2023 / 063769 "DirAC-SPAR Audio Processing" · U.S. Patent No. 9,978,385 "Parametric reconstruction of audio signals" International Application No. WO20212252748A1 "Encoding of multi-channel audio signals comprising downmixing of a primary and two or more scaled non-primary input channels" International Application No. WO2022120093A1 "Immersive Voice and Audio Services (IVAS) with Adaptive Downmix Strategies" International Application No. WO2021252811A2 "Quantization and Entropy Coding of Parameters for a Low Latency Immersive Audio Codec" U.S. Patent Application Publication No. 20220406318A1, "Bitrate Distribution in Immersive Voice and Audio Services," and · U.S. Patent Application Publication No. 2022 / 0277757 "Systems and Methods for Covariance Smoothing"

[0031] I. Reduced Metadata Bitrate and Bandwidth Interleaving When implementing LBRSBA (e.g., below 24.4 kbps), a trade-off must be made between spatial metadata bitrate and quality limitations and core codec bitrate and audio quality limitations. In some embodiments, the lower bitrate is achieved by operating with fewer frequency bands (e.g., 6 bands instead of 12 bands) to reduce the amount of spatial metadata conveyed from the encoder to the decoder. In one particular embodiment, the bottom four bands (the four lower frequency bands) are assigned to SPAR, and the top two bands (the two higher frequency bands) are assigned to DirAC.

[0032] Table I below shows an example band allocation when going from 12 bands to 6 bands in a particular embodiment. [Table 1]

[0033] As shown in Table I, the SPAR band (Bands 0-7) is reduced to an LBRSBA band group with four bands (Bands 0-3), and the DirAC band (Bands 8-11) is reduced to an LBRSBA band group with two bands (Bands 4 and 5), for a total of six LBRSBA bands. The temporal resolution of the metadata is also reduced. At higher bit rates, DirAC metadata is often calculated at 5 ms resolution, while SPAR metadata is calculated at 20 ms resolution. Therefore, a slower metadata update rate is used in LBRSBA. In particular, SPAR metadata moves to a 40 ms update rate, with occasional 20 ms updates, if allowed by bit rate limitations. The DirAC metadata bit rate either remains at 20 ms resolution (compared to the DirAC baseline) or is reduced from 5 ms to 20 ms compared to the higher bit rates used in non-LBRSBA operation. In some embodiments, only the SPAR MD is reduced to band groups. In some embodiments, more or fewer SPAR or DirAC bands can be grouped together and / or there can be more than two groups of bands.

[0034] In some embodiments, for an initial frame or a frame whose metadata is coded within the metadata bitrate budget, all SPAR and DirAC bands are transmitted to the decoder in the frame. If the metadata bitrate budget cannot be met, a first portion (e.g., the first half) of a group of SPAR metadata bands is transmitted in a first data frame, followed by a second data frame containing a second portion (e.g., the second half) of the group of SPAR metadata bands, and so on. In some embodiments, the selection of which bands to transmit or omit in each frame is interleaved. For example, when band metadata is omitted for a frame, the band metadata is assumed to be the same metadata as that band for the previous frame. The advantage of this interleaving approach is that more frames are generated with (relatively) finely quantized metadata, at the expense of temporal resolution. This significantly reduces the metadata bitrate, leaving more bits for the core coder.

[0035] In some embodiments, the indication of what type of frame is coded (full band, A data frame or B data frame) is achieved by reusing existing SPAR metadata bitstream signaling for staggered coding of metadata.

[0036] Table II below lists exemplary coding schemes for non-LBRSBA and LBRSBA coding. [Table 2]

[0037] Referring to Table II above, BASE indicates entropy coding using an arithmetic coder. FOUR_X indicates temporally staggered coding of several bands using original arithmetic and a temporally staggered arithmetic coder. In some embodiments, a Huffman coder is used.

[0038] In an A or B frame, the metadata for an untransmitted band is kept at the value for that band from the previous frame in which it was transmitted. In the case of packet loss, the best case recovery is one frame if a BASE or BASE_NOEC frame is used in the subsequent frame, or two frames (consecutive A and B frames) otherwise.

[0039] Table III below lists example IVAS SBA (Ambisonics) bit rates, including LBRSBA bit rates. [Table 3]

[0040] [II. Covariance smoothing] In some embodiments, additional covariance smoothing is applied to the LBRSBA covariance matrix to further reduce the SPAR metadata bitrate and improve the single channel element (SCE) core decisions (e.g., ACELP / TCX) used to code the spatial downmix channels. Covariance smoothing is described in U.S. Patent Application Publication No. 2022 / 0277757, "Systems and methods for covariance smoothing," but the smoothing coefficients are modified for LBRSBA as described below. Smoothing is applied to the covariance matrix before the MD is received or calculated. In some embodiments, a frequency-domain representation of the audio is used to generate a covariance matrix that is smoothed using the covariance smoothing techniques described below. After covariance smoothing, the smoothed covariance matrix is ​​used to form SPAR and DirAC metadata, which are grouped into LBRSBA bands, quantized, and transmitted to the decoder, as shown in Table I.

[0041] [i. Smoothing function and forgetting factor] Covariance smoothing utilizes a smoothed matrix. Generally, the smoothed matrix can be calculated using a low-pass filter designed to meet specific smoothing requirements. In some embodiments, the smoothing requirements are such that previous estimates are used to artificially increase the number of frequency samples (bins) used to generate the current estimate of the covariance matrix. In some embodiments, a smoothed matrix σ is calculated from the input covariance matrix A over a sequence of frames.

number

number

[0042] Equation [1] is an example of a smoothing function that is a first-order low-pass filter. Other smoothing functions, such as higher-order filters, can also be used. Key elements of a smoothing function are the lookback aspect, which uses previously smoothed results, and the forgetting factor, which weights the influence of these results. The effect of the forgetting factor is that as smoothing is applied across successive frames, the effect of previous frames becomes less and less influential on the smoothing of the frame being smoothed (adjusted). When the forgetting factor in Equation [1] is 1 (λ=1), no smoothing is performed, which effectively functions as an all-pass filter. When 0<λ<1, the equation functions as a low-pass filter. Lower λ places more emphasis on older covariance data, while higher λ takes newer covariances into account more. A forgetting factor greater than 1 (e.g., 1<λ<2) is implemented as a high-pass filter.

[0043] In some embodiments, the maximum allowable forgetting factor λ max is implemented. This maximum value determines the behavior of the algorithm as the bin / band value gets larger. In some embodiments, λ max < 1 will always perform some smoothing in all bands, regardless of the calculated forgetting factor, and λ max = 1 is the desired N min The smoothing function is applied only to bands with fewer bins than {tilde over (x)}, leaving larger bands unsmoothed.

[0044] In some of these embodiments, the forgetting factor λ for a particular band b is the maximum allowable forgetting factor λ max and the minimum number of bins N that is determined to give good statistical estimates based on the window size. min and the effective number of bins in the band, N b It is calculated as the minimum of the ratio of

number

[0045] In some embodiments, Nb is the actual count of the bins in the frequency band. In some embodiments, N b can be calculated from the sum of the frequency responses of a particular band, e.g., if the band response is r=[0.5,1,1,0.5,0,...,0], then the effective number of bins is N b = sum(r) = 0.5 + 1 + 1 + 0.5 = 3. In some embodiments, λ max = 1, and as a result, λ b is within a reasonable range, e.g., 0≦λ b ≦1, which means that smoothing is applied proportionally to the small sample estimates and no smoothing at all to the large sample estimates. max <1, which smooths larger bands to some extent, regardless of their size (e.g., λ max =0.9). In some embodiments, N min can be selected based on the data at hand that produces the best subjective results. min can be selected based on how early (first subsequent frame after the initial frame of a given window) smoothing is desired.

[0046] In one example, an analysis filter bank is used that has a narrower (i.e., fewer bins and more frames required for good statistical analysis) low frequency band and a wider (i.e., more bins and fewer frames required for good statistical analysis) high frequency band, which increases the amount of smoothing in the lower frequency band and decreases the amount in the higher frequency band (or λ max = 1, no smoothing at all).

[0047] Figure 2 is a flow diagram of a covariance smoothing process according to one or more embodiments. An input frequency-domain signal (e.g., a Fast Fourier Transform (FFT)) 201 provides, for a given band in the input signal, a corresponding covariance matrix over a window. The effective bin count for that band is obtained 202. This can be calculated, for example, by the band's filter bank response value. The desired bin count is determined 203, for example, by a subjective analysis of how many bins are needed to provide a good statistical analysis for that window. A forgetting factor is calculated 204 by taking the ratio of the number of calculated bins to the desired bin count. For a given frame (other than the first frame), new covariance matrix values ​​are calculated 205 based on the new covariance values ​​calculated for the previous frame, the original values ​​for the current frame, and the forgetting factor. The new (smoothed) matrix formed by these new values ​​is used in further signal processing 206.

[0048] 3 illustrates an exemplary modification to process 200 for a maximum allowed forgetting factor, according to one or more embodiments. As shown in FIG. 2, a forgetting factor is calculated for a band 301. Furthermore, a maximum allowed forgetting factor is determined 302. These values ​​are compared 303, and depending on whether the calculated factor is smaller than the maximum allowed factor, the calculated factor is used in smoothing 305 (hereinafter "smoothing_factor"). If the calculated factor is larger than the maximum allowed factor, the maximum allowed factor is used in smoothing 305 304. While this example shows that the calculated factor is used if the factors are equal (not larger), an equivalent flow can be envisioned in which the minimum value is used if they are equal.

[0049] In some embodiments, the smoothing factor is different depending on whether the codec is operating in non-LBRSBA or LBRSBA. The non-LBRSBA smoothing factor is given by the following equation: smoothing_factor(b) = update_factor(b) / min_pool_size = number of bins in frequency band b / minimum number of desired bins [3]

[0050] The LBRSBA smoothing coefficient is given by the following formula: smoothing_factor(b)=update_factor(b) / min_pool_size*k*(b+1) [4] Here, for non-LBRSBA, an exemplary factor is k=0.75 to increase / decrease smoothing in the low frequency band while avoiding smoothing in the higher bands.

[0051] Smoothing of the covariance matrix is ​​then performed according to principles related to equation [1]. For LBRSBA, k is set to 0.5 to further increase smoothing in the lowest bands (e.g., SPAR bands) while avoiding modifying smoothing in higher bands (e.g., DirAC bands). Other embodiments may use other values ​​for k.

[0052] [ii. Smoothing Reset] In some embodiments, it may be desirable to avoid smoothing over transients (sudden changes in signal level) as this can produce undesirable signal distortions / artifacts in the output. In these embodiments, the smoothing can be "reset" at the point when a transient is detected in the signal. The estimated smoothing matrix of the previous time frame can be stored to facilitate calculation of the smoothed value for the current frame. If a transient is detected in the input signal during that frame, the smoothing function can be set to re-initialize itself. When a transient is detected, the past matrix estimate is reset to the current estimate, so that the output of the smoothing filter after the transient is the estimate itself (no changes applied). In other words, for the reset frame,

number

[0053] [iii. Transient detection] 4 is a flow diagram of a process for modifying the transient detection process flow, according to one or more embodiments. A determination is made 401 whether a transient has been detected for a given frame. If so 403, the new matrix values ​​remain the same as the input values. If not 402, the normal smoothing algorithm is used for that frame. The combination (matrix) of the smoothed and unsmoothed (transient) frame values ​​is used for signal processing 404.

[0054] In some embodiments, smoothing is reset when a transient is detected on any channel. For example, if there are N channels, N transient detectors can be used (one per channel), and if any of them detect a transient, smoothing is reset, or the signal ends or smoothing ends (smoothing is turned off). In the example of a stereo input, it may be determined that the channels are sufficiently separate (or possibly separate) because considering only transients in the left channel may mean that important transients in the right channel may be inadequately smoothed (and vice versa). Therefore, two transient detectors are used (left and right), and any one of these can trigger a smoothing reset of the entire 2x2 matrix.

[0055] In some embodiments, the smoothing is reset only for transients in a particular channel. For example, if there are N channels, only M detectors (<N, in some cases 1) are used. In an example of a First Order Ambisonics (FOA) input, it can be determined that the first (W) channel is the most important compared to the other three (X, Y, Z), and considering the spatial relationship between the FOA signals, transients in the latter three channels are likely to be reflected in the W channel anyway. Thus, the system can be set up using a transient detector only on the W channel, and detecting a transient on W triggers the reset of the overall 4×4 covariance smoothing matrix.

[0056] In some embodiments, the reset resets only the covariance elements that have experienced a transient. This means that a transient in the nth channel resets only the values in the nth row and nth column of the covariance matrix (the entire row and entire column). This can be implemented by having separate transient monitoring on each channel, and a detected transient on any given channel triggers the reset of the matrix positions corresponding to the covariance of that channel to other channels (and vice versa, trivially to itself).

[0057] In some embodiments, the reset occurs only when a majority / threshold number of channels detect a transient. For example, in a 4-channel system, the threshold may be set to trigger a reset only if at least two of the channels report a transient in the same frame. In some embodiments, a band-selective covariance smoothing reset is implemented. The covariance smoothing reset function helps to allow the covariance to move quickly when a transient occurs, but if some bands are severely smoothed, for example at the lowest frequencies, the rapidly repeated detected transients and subsequent resets of the covariance smoothing can, in some cases, create an audible stuttering effect. This effect can be minimized / avoided by selectively resetting the bands with less smoothing.

[0058] [III. Decorrelator Smoothing] Due to the relatively coarse quantization of the decorrelator parameters required to meet metadata bitrate targets, quantization errors often manifest as too much decorrelation, or a flickering between significant and insignificant amounts of decorrelation (e.g., large and small amounts, or even zero). In some embodiments, decoder-side decorrelator coefficient smoothing can help prevent this effect from becoming audible. While many forms of smoothing are possible, in this embodiment, the smoothing is mathematically equivalent to that used in the covariance smoothing described above, except that there is no ability to reset the smoothing for transients. In some embodiments, a forgetting factor of 0.5 can be used for all bands, although other values ​​are possible.

[0059] 5 is a plot of smoothed and unsmoothed quantized decorrelation coefficients over n frames in accordance with one or more embodiments. A first plot 501 shows an example of unsmoothed quantized decorrelation coefficients at a decoder with three possible levels (0.0, 0.4, 0.8). A second plot 502 shows an example of smoothed quantized decorrelation coefficients.

[0060] [Example Process] Figure 6 is a flow diagram of LBRSBA Ambisonics processing according to one or more embodiments. Process 600 can be implemented using the electronic device architecture described with reference to Figure 7. In some embodiments, process 600 includes receiving scene-based audio metadata (601), creating Spatial Reconstruction (SPAR) metadata and Directional Audio Coding (DirAC) metadata from the scene-based audio metadata (602), forming groups of SPAR metadata bands and groups of DirAC metadata bands (603), quantizing the groups of SPAR metadata bands and the groups of DirAC metadata bands (604), and transmitting to a decoder a first data frame including a first portion of the quantized DirAC metadata bands and a second data frame subsequent to the first data frame, the second data frame including a second portion of the quantized DirAC metadata bands and the quantized SPAR metadata bands.

[0061] Each of these steps is described more fully above.

[0062] [Example System Architecture] FIG. 7 illustrates a block diagram of an exemplary electronic device architecture 700 suitable for implementing exemplary embodiments of the present disclosure. The architecture 700 may include, but is not limited to, a server and a client device, as described above with reference to FIGS. 1-6. As illustrated, the architecture 700 includes a central processing unit (CPU) 701 that can execute various processes according to a program stored in, for example, a read-only memory (ROM) 702 or loaded from, for example, a storage unit 708 into a random access memory (RAM) 703. The RAM 703 also stores data needed by the CPU 701 to execute various processes, as needed. The CPU 701, ROM 702, and RAM 703 are interconnected via a bus 804. An input / output (I / O) interface 705 is also connected to the bus 704.

[0063] The following components are connected to the I / O interface 705: an input unit 706 which may include a keyboard, a mouse, etc.; an output unit 707 which may include a display such as a liquid crystal display (LCD) and one or more speakers; a storage unit 708 which may include a hard disk or another suitable storage device; and a communication unit 709 which may include a network interface card such as a network card (e.g., wired or wireless).

[0064] In some implementations, the input unit 706 includes one or more microphones in different positions (depending on the host device) that enable capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

[0065] In some implementations, the output unit 707 includes a system with a varying number of speakers, and (depending on the capabilities of the host device) the output unit 707 can render audio signals in a variety of formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

[0066] In some embodiments, the communication unit 709 is configured to communicate with other devices (e.g., via a network). Optionally, a drive 710 is also connected to the I / O interface 705. A removable medium 711, such as a magnetic disk, optical disk, magneto-optical disk, flash drive, or other suitable removable medium, is attached to the drive 710 so that computer programs read therefrom can be installed in the storage unit 708, as needed. While the system 700 has been described as including the above components, those skilled in the art will understand that in actual applications, some of these components can be added, removed, and / or substituted, and all such modifications or variations are within the scope of the present disclosure.

[0067] According to exemplary embodiments of the present disclosure, the above processes may be implemented as a computer software program or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the method. In such embodiments, the computer program may be downloaded and loaded from a network via a communication unit 709 and / or installed from a removable medium 711, as shown in FIG. 7.

[0068] In general, various exemplary embodiments of the present disclosure may be implemented in hardware or dedicated circuitry (e.g., control circuitry), software, logic, or any combination thereof. For example, the above-described units may be executed by a control circuit (e.g., CPU 701 in combination with other components of FIG. 7 ), which may then perform the operations described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device (e.g., control circuitry). While various aspects of the exemplary embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it will be recognized that the blocks, apparatus, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controller or other computing device, or any combination thereof, as non-limiting examples.

[0069] Furthermore, the various blocks illustrated in the flowcharts may be viewed as method steps, and / or as operations resulting from the operations of computer program code, and / or as multiple coupled logic circuit elements configured to perform the associated functions. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code configured to perform the methods described above.

[0070] In the context of this disclosure, a machine-readable medium may be any tangible medium that contains or can store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may be non-transitory and may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of machine-readable storage media include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0071] Computer program code for carrying out the methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus having control circuitry, so that when executed by the processor of the computer or other programmable data processing apparatus, the program code causes the functions / acts specified in the flowcharts and / or block diagrams to be performed. The program code may run entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.

[0072] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in a particular combination and initially claimed as such, one or more features from the claimed combination may, in some cases, be deleted from the combination, and the claimed combination may be directed to subcombinations or variations of the subcombination. The logic flow depicted in the figures does not require the particular order or sequential order depicted to achieve desired results. Furthermore, other steps may be provided or steps may be deleted from the described flow, and other components may be added to or removed from the described system. Accordingly, other implementations are within the scope of the following claims.

Claims

1. 1. A method for audio metadata encoding, comprising: receiving, with at least one processor, scene-based audio metadata; generating, with the at least one processor, spatial reconstruction (SPAR) metadata and directional audio coding (DirAC) metadata from the scene-based audio metadata; forming, with the at least one processor, a group of SPAR metadata bands and a group of DirAC metadata bands; quantizing, with the at least one processor, the group of SPAR metadata bands and the group of DirAC metadata bands; transmitting to a decoder a first data frame including the group of quantized DirAC metadata bands and a first portion of the group of quantized SPAR metadata bands, and a second data frame following the first data frame, the second data frame including the second portion of the group of quantized DirAC metadata bands and the group of quantized SPAR metadata bands; A method comprising:

2. The method of claim 1 , further comprising the step of transmitting a signal indicative of the first data frame or the second data frame to the decoder.

3. 2. The method of claim 1, wherein the group of SPAR metadata bands includes four SPAR metadata bands and the group of DirAC bands includes two DirAC metadata bands, the group of SPAR bands being at a lower frequency than the group of DirAC bands.

4. 2. The method of claim 1 , wherein the group of DirAC metadata bands is transmitted to the decoder at a first temporal resolution, and the first and second portions of the group of SPAR metadata bands are transmitted to the decoder at a second temporal resolution, the second temporal resolution being lower than the first temporal resolution.

5. 5. The method of claim 4, wherein the group of DirAC metadata bands is transmitted to the decoder at the first temporal resolution when the first data frame is an initial data frame or when the group of DirAC metadata bands is encoded within a metadata bitrate budget.

6. The method of claim 4 , wherein when the group of SPAR metadata bands is not encoded within a metadata bitrate budget, the group of SPAR metadata bands is transmitted to the decoder at the second temporal resolution.

7. 2. The method of claim 1 , further comprising, prior to receiving the scene-based audio metadata, applying smoothing, with the at least one processor, to a covariance matrix from which the scene-based audio metadata is formed.

8. The method of claim 7 , wherein the covariance smoothing uses a smoothing factor that increases smoothing in low frequency bands and avoids modifying the amount of smoothing in high frequency bands.

9. 8. The method of claim 7, wherein the smoothing factor is given by the function smoothing_factor(b)=update_factor(b) / min_pool_size*k*(b+1), where update_factor(b) is the number of frequency bins in frequency band b, min_pool_size is the minimum number of frequency bins desired, and k is a factor that increases or decreases the smoothing.

10. 1. A method for audio metadata decoding, comprising: receiving, with at least one processor, quantized scene-based audio data and corresponding metadata, the metadata including decorrelator coefficients; using the at least one processor, dequantizing the quantized scene-based audio data and corresponding metadata; decoding, with the at least one processor, the scene-based audio data and corresponding metadata, the decoding including recovering the decorrelator coefficients; smoothing the decorrelator coefficients with the at least one processor; reconstructing, with the at least one processor, a multi-channel audio signal based at least on the decoded scene-based audio data and the smoothed decorrelator coefficients; A method comprising:

11. 1. A computing device comprising: at least one processor; a memory storing instructions which, when executed by said at least one processor, cause said computing device to perform the method of any one of claims 1 to 10; 1. A computing device comprising: