Method, apparatus and system for directional audio coding-spatial reconstruction audio processing
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-06
- Publication Date
- 2026-03-16
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] This disclosure relates generally to audio processing. [Background technology]
[0002] 1.0 Background Spatial Reconstruction (SPAR) and Directional Audio Coding (DirAC) are independent spatial audio coding techniques that aim to represent an input spatial audio scene in a compact way, allowing its transmission with a good trade-off between audio quality and bitrate. One input format for such a spatial audio scene is an Ambisonics representation (e.g., first-order Ambisonics (FOA) or higher-order Ambisonics (HOA)).
[0003] SPAR seeks to maximize the perceived audio quality while minimizing the bitrate by reducing the energy of the transmitted audio data, while still allowing the second order statistics (i.e., covariance) of the Ambisonics audio scene to be reconstructed at the decoder side by using the transmitted metadata. SPAR seeks to faithfully reconstruct the input Ambisonics scene at the decoder's output.
[0004] DirAC is a technique that represents a spatial audio scene as a collection of directions of arrival (DOAs) in time-frequency tiles. From this representation, it is possible to reproduce similar sounding scenes in different output formats (e.g. binaural). In particular, for Ambisonics, the DirAC representation allows the decoder to generate a higher order output from a lower order input (blind upmix). DirAC tries to preserve the direction and diffuseness of sounds that are dominant in the input scene.
[0005] Both DirAC and SPAR have different strengths and properties. Therefore, it is desirable to combine the complementary aspects of DirAC and SPAR (e.g., higher audio quality, reduced bit rate, input / output format flexibility, and / or reduced computational complexity) into a coder / decoder ("codec") such as the Ambisonics codec.
[0006] 1.1 IVAS Codec Framework Example 1 is a block diagram of an IVAS coder / decoder ("codec") framework 100 for encoding and decoding immersive voice and audio services (IVAS) bitstreams in one or more implementations. IVAS is envisioned to support a variety of audio service capabilities, such as, but not limited to, mono-to-stereo upmix, and fully immersive audio encoding, decoding, and rendering. IVAS is also intended to be supported by a wide range of devices, endpoints, and network nodes, such as, but not limited to, mobile phones and smartphones, electronic tablets, personal computers, conference phones, conference rooms, virtual reality (VR) and augmented reality (AR) devices, home theater equipment, and other suitable devices.
[0007] The IVAS codec 100 includes an IVAS encoder 101 and an IVAS decoder 104. The IVAS encoder 101 includes a spatial encoder 102 that receives N channels of input spatial audio (e.g., FOA, HOA). In some implementations, the spatial encoder 102 implements SPAR and DirAC to analyze / downmix the N_dmx spatial audio channels, as described in more detail below. The output of the spatial encoder 102 includes a spatial metadata (MD) bis stream (BS) and N_dmx channels of spatial downmix. The spatial MD is quantized and entropy coded. In some implementations, the quantization can include fine, medium, coarse and ultra-coarse quantization strategies, and the entropy coding can include Huffman coding or arithmetic coding. The core audio encoder 103 (e.g., an Enhanced Voice Service (EVS) encoding unit) encodes the N_dmx channels (N=1 to 16 channels) of the spatial downmix into an audio bitstream. The audio bitstream is combined with the spatial MD bitstream into an IVAS encoded bitstream. The IVAS encoded bitstream is sent to the IVAS decoder 104.
[0008] The IVAS decoder 104 includes a core audio decoder 105 (e.g., an EVS decoder) that decodes an audio bitstream derived from the IVAS bitstream to recover the N_dmx audio channels. The spatial decoder / renderer 106 (e.g., SPAR / DirAC) decodes a spatial MD bitstream derived from the IVAS bitstream to recover the spatial MD and synthesizes / renders output audio channels using the spatial MD and spatial upmix for playback on various audio systems with different speaker configurations and capabilities. Summary of the Invention
[0009] An embodiment for DirAC-SPAR audio processing is disclosed.
[0010] In some embodiments, a method includes: receiving, by at least one processor, a multi-channel audio signal including a first set of channels; for a first set of frequency bands: computing, by the at least one processor, directional audio coding (DirAC) metadata from the first set of channels; quantizing, by the at least one processor, the DirAC metadata; encoding, by the at least one processor, the quantized DirAC metadata; converting, by the at least one processor, the quantized DirAC metadata into two or more parameters of first spatial reconstruction (SPAR) metadata; for a second set of frequency bands lower than the first set of frequency bands: computing, by the at least one processor, second SPAR metadata from the first set of channels; The method includes: quantizing the second SPAR metadata by the at least one processor; encoding the quantized second SPAR metadata by the at least one processor; generating a downmix based on the first SPAR metadata and the second SPAR metadata by the at least one processor; calculating frequency coefficients from the first set of channels by the at least one processor; downmixing from the coefficients and the downmix to a second set of channels by the at least one processor; encoding the second set of channels by the at least one processor; and outputting a bitstream including the encoded second set of channels, the quantized and encoded second SPAR metadata, and the quantized and encoded DirAC metadata.
[0011] In some embodiments, the first set of channels are primary ambisonic (FOA) channels.
[0012] In some embodiments, one or more parameters in the first SPAR metadata for the first set of frequency bands are encoded into the bitstream rather than being converted from DirAC metadata.
[0013] In some embodiments, the first SPAR metadata parameters encoded in the bitstream are calculated from a combination of DirAC metadata and input covariances of the first set of channels.
[0014] In some embodiments, the second set of channels includes a primary downmix channel, the primary downmix channel being obtained by applying a gain to the first set of channels and adding together the gain adjusted first set of channels, the gain being calculated from the DirAC metadata, and the primary downmix channel being a representation of a dominant specific signal for the first set of channels.
[0015] In some embodiments, a method includes: receiving, by at least one processor, a multi-channel audio signal including a first set of channels and a second set of channels different from the first set of channels; for a first set of frequency bands: computing, by the at least one processor, directional audio coding (DirAC) metadata from the first set of channels; quantizing, by the at least one processor, the DirAC metadata; encoding, by the at least one processor, the quantized DirAC metadata; converting, by the at least one processor, the quantized DirAC metadata into two or more parameters of a first spatial reconstruction (SPAR) metadata; and for a second set of frequency bands lower than the first set of frequency bands: computing, by the at least one processor, a second SPAR metadata from the first set of channels and the second set of channels. the at least one processor calculating R metadata; quantizing the second SPAR metadata; encoding the quantized second SPAR metadata; generating a downmix based on the first SPAR metadata and the second SPAR metadata by the at least one processor; calculating frequency coefficients from the first set of channels and the second set of channels by the at least one processor; downmixing from the coefficients and the downmix to a third set of channels by the at least one processor; encoding the third set of channels by the at least one processor; and outputting a bitstream including the encoded third set of channels, the quantized and encoded second SPAR metadata, and the quantized and encoded DirAC metadata.
[0016] In some embodiments, two or more parameters in the first SPAR metadata are converted from DirAC metadata and the second SPAR data is calculated using input covariances.
[0017] In some embodiments, one or more parameters in the first SPAR metadata for the first set of frequency bands are encoded into the bitstream rather than being converted from DirAC metadata.
[0018] In some embodiments, the first SPAR metadata parameters encoded in the bitstream are calculated from a combination of DirAC metadata and covariances of the second set of channels.
[0019] In some embodiments, the first SPAR metadata parameters encoded in the bitstream include prediction coefficients, cross-prediction coefficients and decorrelation coefficients for the second set of channels.
[0020] In some embodiments, the first set of channels are first order Ambisonic (FOA) channels and the second set of channels include at least one of planar or non-planar higher order Ambisonic (HOA) channels.
[0021] In some embodiments, the two or more parameters of the first SPAR metadata are converted from DirAC metadata, and the second SPAR metadata is calculated and encoded for all frequency bands.
[0022] In some embodiments, the second SPAR metadata is calculated from the first and second sets of channels and the first SPAR metadata.
[0023] In some embodiments, the DirAC metadata is estimated based on the input covariance matrix.
[0024] In some embodiments, generating the SPAR metadata from DirAC metadata includes approximating a second input covariance from the DirAC metadata and a spherical harmonic response, and calculating the two or more parameters in the SPAR metadata from the second input covariance.
[0025] In some embodiments, one or more elements of the second input covariance are generated using decorrelation coefficients in the DirAC metadata and the second SPAR metadata.
[0026] In some embodiments, one or more elements of the second input covariance are generated from DirAC metadata such that the decorrelation coefficients in the SPAR metadata depend only on a diffuseness parameter and an Ambisonics input normalization and one or more constants in the DirAC metadata.
[0027] In some embodiments, the third set of channels includes a primary downmix channel, the primary downmix channel being obtained by applying a gain to the first set of channels and adding together the gain adjusted first set of channels, the gain being calculated from the DirAC metadata, and the primary downmix channel being a representation of a dominant specific signal for the first set of channels.
[0028] In some embodiments, the DirAC metadata includes a diffuseness parameter calculated based on a reference power (E) and intensity (I) of the multi-channel audio signal, where E and I are calculated based on the input covariance.
[0029] In some embodiments, the first set of channels includes a primary Ambisonic (FOA) channel, and the calculation of the reference power in the DirAC metadata ensures that the reference power is always greater than or equal to the variance of the W channel of the FOA channel.
[0030] In some embodiments, the downmix is energy compensated in the first set of frequency bands based on a ratio of the total variance of the first set of channels to the total variance according to the second input covariance generated using the DirAC metadata.
[0031] In some embodiments, a method includes: receiving, by at least one processor, an encoded bitstream including encoded audio channels and metadata, the metadata including first Directional Audio Coding (DirAC) metadata associated with a first frequency band and first Spatial Reconstruction (SPAR) metadata associated with a second frequency band lower than the first frequency band; decoding, by the at least one processor, the first DirAC metadata and the first SPAR metadata; dequantizing, by the at least one processor, the decoded first DirAC metadata and the first SPAR metadata; for the first frequency band: converting, by the at least one processor, the dequantized first DirAC metadata into two or more parameters of a second SPAR metadata; mixing, by the at least one processor, the first and second SPAR metadata into a composite SPAR metadata; decoding, by at least one processor, the encoded audio channels; reconstructing, by the at least one processor, downmix channels from the decoded audio channels; transforming, by the at least one processor, the downmix channels into a frequency banding domain; generating, by the at least one processor, a SPAR upmix based on the composite SPAR metadata; upmixing, by the at least one processor, the downmix channels in the frequency banding domain to a first set of channels based on the SPAR upmix; estimating, by the at least one processor, second DirAC metadata in the second frequency band from the first set of channels and zero or more parameters in the first SPAR metadata; upmixing, by the at least one processor, the first set of channels to a second set of channels in the frequency banding domain based on the first and second DirAC metadata;and transforming, by the at least one processor, the second set of channels from the frequency banded domain to a time domain;
[0032] In some embodiments, the downmix is transformed into the frequency banded domain using a filterbank (complex low delay filterbank).
[0033] In some embodiments, the first set of channels includes a First Order Ambisonics (FOA) channel and zero or more Higher Order Ambisonics (HOA) channels.
[0034] In some embodiments, the HOA channels in the first set of channels include at least one of a planar HOA channel or a non-planar HOA channel.
[0035] In some embodiments, the bitstream includes third SPAR metadata corresponding to an HOA channel of the first set of channels and the first frequency band.
[0036] In some embodiments, the DirAC metadata is estimated from first order Ambisonics (FOA) channels in the frequency band domain for a third set of frequency bands that includes the first set of frequency bands and the second set of frequency bands.
[0037] In some embodiments, the DirAC metadata is estimated for a fourth set of frequency bands, which is a subset of the second set of frequency bands, from SPAR metadata and zero or more elements of covariance generated using the downmix and the upmix in the fourth set of frequency bands.
[0038] In some embodiments, calculating the DirAC metadata from SPAR metadata for the fourth set of frequency bands includes calculating a direction of arrival angle in the DirAC metadata from only prediction coefficients in the SPAR metadata, and calculating a diffuseness parameter in the DirAC metadata from the prediction coefficients and zero or more decorrelation coefficients, and a scale factor in the SPAR metadata.
[0039] In some embodiments, the encoded channels include a first order Ambisonic channel, and upmixing the downmix channel to a first set of channels in the first frequency band includes calculating an upmix scaling gain from the first DirAC metadata and applying the upmix scaling gain to the primary downmix channel to obtain a W channel of the first set of channels in the first frequency band, where the primary downmix channel is a representation of a dominant intrinsic signal for the first set of channels.
[0040] In some embodiments, a non-transitory computer readable storage medium storing instructions that, when executed by a computing device, cause the computing device to perform any of the above methods.
[0041] In some embodiments, a computing device comprises at least one processor and a memory storing instructions that, when executed by the at least one processor, cause the computing device to perform any of the methods described above.
[0042] Other embodiments disclosed herein are directed to systems, devices, and computer-readable media. Details of the disclosed embodiments are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
[0043] Certain embodiments disclosed herein combine complementary aspects of DirAC and SPAR technologies, including higher audio quality, reduced bit rate, input / output format flexibility, and / or reduced computational complexity, to produce a codec (e.g., an Ambisonics codec) that has better overall performance than the DirAC and SPAR codecs. [Brief description of the drawings]
[0044] [Figure 1] FIG. 1 is a block diagram of an IVAS codec framework in one or more embodiments.
[0045] [Diagram 2] FIG. 2 is a block diagram of an encoder implementation with frequency-based and channel-based split between SPAR and DirAC in accordance with one or more embodiments.
[0046] [Diagram 3] FIG. 3 is a block diagram of a decoder implementation with frequency-based and channel-based split between SPAR and DirAC in accordance with one or more embodiments.
[0047] [Figure 4] FIG. 4 is a block diagram of an alternative encoder implementation with frequency-based and channel-based split between SPAR and DirAC in accordance with one or more embodiments.
[0048] [Diagram 5] FIG. 5 is a block diagram of an alternative decoder implementation with frequency-based and channel-based split between SPAR and DirAC in one or more embodiments.
[0049] [Figure 6]FIG. 6 is a flow diagram of a process for encoding FOA input using a codec as described with reference to FIGS. 2 and 4 in some embodiments.
[0050] [Figure 7] FIG. 7 is a flow diagram of a process for encoding FOA+HOA input using the codec described with reference to FIGS. 2 and 4 in accordance with some embodiments.
[0051] [Figure 8] FIG. 8 is a flow diagram of a process for decoding using the codec described with reference to FIGS. 3 and 5 in accordance with some embodiments.
[0052] [Figure 9] FIG. 9 is a block diagram of an example hardware architecture suitable for implementing the systems and methods described with reference to FIGS. 1-8. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0053] In the drawings, a particular arrangement or ordering of schematic elements, such as elements representing devices, units, instruction blocks, and data elements, is illustrated for ease of description. However, it should be understood by those skilled in the art that the particular ordering or arrangement of these schematic elements in the drawings does not imply that a particular order or sequence of processing, or separation of processing, is required. Furthermore, the inclusion of a schematic element in a drawing does not imply that such an element is required in all embodiments, nor does it imply that features represented by such an element may not be included in or combined with other elements in some implementations.
[0054] Furthermore, in the drawings, when a connecting element such as a solid or dashed line or an arrow is used to indicate a connection, relationship, or correspondence between two or more other schematic elements, the absence of such connecting element does not imply that the connection, relationship, or correspondence may not exist. In other words, some connections, relationships, or correspondences between elements are not illustrated so as not to obscure the present disclosure. In addition, for ease of illustration, a single connecting element is used to represent multiple connections, relationships, or correspondences between elements. For example, when a connecting element represents communication of signals, data, or instructions, it should be understood by those skilled in the art that such an element represents one or more signal paths for implementing the communication, as appropriate.
[0055] The use of the same reference numbers in the various drawings indicates similar elements.
[0056] Detailed Description In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments described. It will be apparent to those skilled in the art that the various implementations described may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features are described below that can each be used independently or in some combination with other features.
[0057] nomenclature As used herein, the terms "include" and variations thereof should be interpreted as open-ended terms meaning "including, but not limited to." The term "or" should be interpreted as "and / or" unless the context clearly indicates otherwise. The term "based on" should be interpreted as "based at least in part on." The terms "one implementation" and "an implementation" should be interpreted as "at least one implementation." The term "another implementation" should be interpreted as "at least one other implementation." The terms "determined," "determines," and "determining" should be interpreted as "obtain," "receive," "calculate," "calculate," "estimate," "predict," or "derive." Additionally, in the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0058] 2.0 Algorithm Analysis As mentioned above, SPAR seeks to maximize the perceived audio quality while minimizing the bitrate by reducing the energy of the transmitted audio data, while still allowing the second-order statistics (i.e., covariance) of the Ambisonics audio scene to be reconstructed at the decoder side using the transmitted metadata. DirAC seeks to preserve the direction and diffuseness of the sounds that dominate in the input scene. An overview of the DirAC and SPAR techniques is given later in Sections 2.1 and 2.3, respectively.
[0059] 2.1 DirAC technology 2.1.1 References
[0060] The DirAC technology is described in V. Pulkki, “Directional Audio Coding in Spatial Sound Reproduction and Stereo Upmixing”, in Laboratory of Acoustics and Audio Signal Processing, Helsinki University of Technology, Finland, 2006.
[0061] 2.2 Example of implementation of DirAC analysis in the MDFT domain 2.2.1 DirAC analysis in the MDFT domain In one implementation, the DirAC analysis block takes the Ambisonics time-domain FOA channels as input and transforms them into the frequency domain using a modified discrete Fourier transform (MDFT). It then calculates the magnitude and reference power in the MDFT domain. r , w i , x r , x i , y r , y i , z r , z i Let be the real and imaginary bin samples of the W, X, Y and Z channels of the FOA component of the Ambisonics input in the MDFT domain. The intensity corresponding to frequency bin f of channel X is calculated as follows:
number
[0062] The reference power calculation E at frequency bin f is calculated as follows:
number
[0063] The direction vector dv corresponding to the X channel (or forward / backward direction) and frequency bin f is calculated as follows:
number
number
[0064] Similarly, a direction vector dv corresponding to the Y channel (or left-right direction) and the Z channel (or up-down direction) is calculated.
[0065] 2.2.2 DirAC parameter estimation in banded domain The intensity, reference power and direction vector per bin are then transformed into the banded domain by applying the absolute response of the filter bank to the above calculations in [1], [2] and [3]. The banded intensity, reference power and direction vector in a particular frequency band are denoted by I s , E, dv s where s can be x, y or z.
[0066] 2.2.2.1 Calculating DoA Angles (Azimuth and Elevation) The azimuth and elevation angles in degrees of the dominant sound source in the scene for a particular time-frequency tile are calculated as follows:
number
[0067] 2.2.2.2 Calculation of Diffusivity and Energy Ratio For diffuseness, the long-term averages of E and I are calculated over N frames or M subframes. In an example implementation, a frame represents 20 ms of audio data, a subframe represents 5 ms of audio data, and the long-term averages of E and I are taken over 160 ms of audio data, i.e., 8 frames or 32 subframes. The long-term averages are taken over I slow,s , E slow Then, the diffusivity is given as follows:
number
[0068] The DirAC metadata parameters, i.e., DoA angle and diffuseness parameters, are quantized and coded by the Metadata Quantization and Coding block. Based on the available bitrate, DirAC selects N_dmx audio channels (also called N_dmx downmix channels) from the N channel input, where N_dmx<=N and one of the channels in the N_dmx downmix channels is the W channel of the Ambisonics input to be coded by the core coder. The core coder bits and the DirAC metadata bits are multiplexed into a bitstream and sent to the decoder. The decoder decodes the bitstream and reconstructs the N_dmx downmix channels using the core decoder and the DirAC metadata parameters using the Metadata Unquantization and Decoding block. The N_dmx downmix channels and the DirAC metadata parameters are fed to the DirAC Synthesis and Rendering block, which calculates the directional components of the output spatial audio scene using the W channel and spherical harmonics functions according to the DoA angle. The DirAC synthesis and rendering block also computes the diffuse component of the output spatial audio scene using a decorrelated version of the W channel, which is generated using a decorrelator block, and the diffuseness parameter in the DirAC metadata, and then uses the N_dmx downmix channels and the directional and diffuse components to output the desired audio output format.
[0069] 2.3 Example of SPAR (spatial reconstruction) implementation using FOA input SPAR is a technique for efficiently encoding spatial audio input. SPAR takes a multi-channel input and generates spatial metadata and a downmix signal such that the combination of spatial metadata and downmix signal can be encoded with higher coding efficiency than encoding each channel of the multi-channel input separately. SPAR aims to recreate the covariance of an N-channel multi-channel input and calculates spatial metadata and an N_dmx channel downmix signal, where N_dmx<=N, based on a parameterized input covariance. The spatial metadata and downmix are quantized, encoded, and sent to a decoder. The decoder decodes the bitstream, dequantizes the spatial metadata, and reconstructs the downmix signal. The decoder then utilizes the spatial metadata and downmix, as well as zero or more decorrelators, to reconstruct the multi-channel input audio scene. An example implementation of SPAR is further described in PCT Patent Application No. PCT / US2023 / 010415, filed January 9, 2023, for “Spatial Coding Of Higher Order Ambisonics For A Low Latency Immersive Audio CODEC.”
[0070] 2.3.11th Order Ambisonics (FOA) Input For an FOA input consisting of channels W, Y, Z, X (according to the ACN channel ordering rule), the SPAR downmix signal can vary from 1 to 4 channels, and the spatial metadata parameters include a prediction parameter PR, a cross-prediction parameter C, and a decorrelation parameter P. These parameters are calculated from the covariance matrix of the windowed input audio signal and are calculated in a specified number of frequency bands (e.g., 12 frequency bands). An example representation of the SPAR parameters retrieval is given below:
[0071] 2.3.1.1 Side signal prediction Predict all side signals (Y, Z, X) from the primary audio signal W and calculate the prediction coefficients for the residual channels using equation
[11] .
number
[11] .
number
[0072] The above downmix is also called passive W downmix, where W is not changed during the downmix process. Another way of downmixing is active W downmix, which allows some mix of Y, X, Z channels with the W channel as follows:
number
[0073] 2.3.1.2 Remixed W Channel and Predicted Channels (Y′,Z′,X′) The W channels and the predicted channels (Y', Z', X') are remixed from most acoustically related to least acoustically related, where remixing involves reordering or shuffling the channels based on some methodology, as shown in Equation
[13] .
number
[0074] Note that one embodiment of the remix may reorder the input channels as W, Y', X', Z', assuming that audio cues from left and right are more important than front and back, with top and bottom cues being the least important.
[0075] 2.3.1.3 Post-prediction covariance calculation The covariance of the 4-channel post-prediction and remix-downmix is calculated as shown in Equation
[14] and Equation
[15] .
number
[0076] For the example of WABC downmix ith 1-4 downmix channels, d and u represent the following channels: (where the placeholder variables A, B, C can be any combination of X, Y, Z channels in FOA). [Table 1]
[0077] 2.3.1.4 Extra C coefficient From these calculations, it is determined whether it is possible to cross-predict any remaining portion of the full parametric channel from the transmitted residual channel. The required extra-C coefficients are:
number
[0078] Thus, C has a shape of (1x2) for a 3-channel downmix and (2x1) for a 2-channel downmix. One embodiment of spatial noise filling does not require these C parameters, which can be set to 0. Alternative embodiments of spatial noise filling may also include a C parameter.
[0079] 2.3.1.5 Residual energy in parameterized channels The residual energy in the parameterized channels that must be filled by the decorrelator is calculated. uu is the actual energy R uu (Post-Prediction) and Regenerated Cross-Prediction Energy Reg uu This is the difference between...
number
[0080] 2.4 Merging DirAC and SPAR As mentioned above, both DirAC and SPAR have different strengths and properties, and it is desirable to combine the complementary aspects of each technology to generate an integrated system that is advantageous in one or more of the following aspects: higher audio quality, reduced bit rate, flexibility of input / output formats, and / or reduced computational complexity. Some of the embodiments for efficiently integrating these two technologies are given below.
[0081] 2.4.1 Frequency-Based SPAR-DirAC Decomposition It has been shown that encoding lower frequency bands with SPAR and higher frequency bands with DirAC improves coding efficiency and quality at the decoder while still reconstructing the spatial audio scene. It may also be desirable to encode the lower frequency bands with SPAR in order to reconstruct the input covariance at the output, or to encode the higher frequency bands with DirAC at the same or finer time resolution in order to efficiently upmix the SPAR reconstructed FoA signal to the HOA.
[0082] In the encoder, the first embodiment 1) converts the time domain wideband Ambisonics input to the frequency banded domain using a filter bank, 2) performs DirAC analysis in the high frequency band to obtain DirAC MD parameters in the high frequency band, 3) performs SPAR analysis in the low frequency band to obtain SPAR MD parameters in the low frequency band, 4) obtains SPAR MD parameters in the high frequency band by converting the DirAC MD parameters to SPAR MD using an MD conversion routine (D2S) (described in Sections 2.5 to 3.4), 5) generates a downmix matrix from the SPAR MD and applies the downmix matrix to the input channels to obtain a downmix channel as described in Section 2.3, 6) quantizes and codes the SPAR MD parameters in the low frequency band and the DirAC MD parameters in the high frequency band, 7) codes the downmix channel using a core audio coder, and 8) multiplexes the MD bits and the core coder bits into a bitstream and sends the bitstream to a decoder.
[0083] In the decoder, the second embodiment includes the steps of: 1) obtaining MD bits and core coder bits from the bitstream, 2) using a core audio decoder to decode the downmix channels, 3) decoding and dequantizing low-frequency SPAR MD parameters and high-frequency DirAC MD parameters from the MD bits, 4) obtaining SPAR MD for high-frequency bands from DirAC MD using a D2S conversion routine, 5) performing filter bank analysis on the decoded downmix channels, 6) using SPAR MD in all frequency bands to generate SPAR upmix in the filter bank domain, and 7) generating spatial audio output in the decoder. In some embodiments, as part of step 7), filter bank synthesis is performed on the SPAR upmixed channels to reconstruct the Ambisonics channels in the decoder. In some other embodiments, as part of step 7), DirAC analysis is performed on the upmixed channels generated by SPAR to obtain DirAC MD parameters in all frequency bands and perform DirAC upmix to a desired output format, including but not limited to HOA2 / HOA3.
[0084] At the decoder, the third embodiment includes the steps of: 1) obtaining MD bits and core coder bits from the bitstream, 2) decoding the downmix channels using a core audio decoder, 3) decoding and dequantizing low frequency SPAR MD parameters and high frequency DirAC MD parameters from the MD bits, 4) obtaining high frequency band SPAR Metadata (MD) from DirAC MD using a D2S conversion routine, obtaining low frequency band DirAC MD from SPAR MD and / or downmix covariance using a SPAR-DiRAC (S2D) MD conversion routine (described in section 3.4), 5) performing filter bank analysis on the decoded downmix channels, 6) generating a SPAR upmix in the filter bank domain using SPAR MD in all frequency bands, and 7) generating spatial audio output at the decoder. In some embodiments, as part of step 7), filter bank synthesis is performed on the SPAR upmixed channels to reconstruct the Ambisonics channels at the decoder. In some other embodiments, as part of step 7), DirAC MD parameters in all frequency bands including the low frequency DirAC MD obtained in step 4) are applied to the SPAR upmix to perform DirAC upmix to the desired output format including but not limited to HOA2 / HOA3.
[0085] 2.4.2 Channel-based SPAR-DirAC decomposition (same processing for all bands) In some embodiments, a subset of Ambisonics input channels are reconstructed (residually or parametrically) via SPAR, and some channels may be reconstructed by DirAC. Any further upmix to higher orders is also handled by DirAC. SPAR reconstructs at least enough channels for DirAC analysis to be performed at the decoder, where in general the DirAC analysis requires FOA channels (or planar FOA channels in the planar case). In this specification, residual coding is a direct audio coding of the residual where the output channels are reconstructed with the prediction components from W, and parametric coding is a coding of cross-prediction and decorrelation parameters where the output is reconstructed with the prediction components from W, the cross-prediction components of the residual, and a decorrelated version of W.
[0086] SPAR generally operates with a B-format representation of the input and output Ambisonics audio. DirAC may reconstruct audio signals in A-format or equivalent spatial domain (ESD) or in B-format. The following description mainly focuses on the latter B-format case. However, similar implementations are possible for DirAC synthesis in A-format or ESD. For this purpose, the SPAR-reconstructed B-format channels may be used to generate a set of relatively sparse DirAC prototype signals in B-format, A-format or ESD, from which a dense set of upmix signals may be generated by DirAC synthesis. Here, each of the upmix signals may drive a speaker of a multi-speaker system. Such a multi-speaker system may correspond to a real speaker setup, e.g. 7.1.4 or 5.1, or a virtual speaker system that is an intermediate step to an immersive binaural rendering of the synthesized audio signal.
[0087] The various embodiments described above are listed in Table II below. [Table 2]
[0088] For HOA3 input, the channels are reconstructed according to the following options: FOA, or HOA2, or FOA+2nd order planar channel, or FOA+2nd order+3rd order planar channel are reconstructed using SPAR, while HOA2 and HOA3, or HOA3, or 2nd order height channel and HOA3, or 2nd order and 3rd order height channels are reconstructed using DirAC to reduce computational complexity without compromising quality.
[0089] For FOA input,bitrates where the number of downmix channels in SPAR is less than 4, 1) FOA with SPAR is implemented with DirAC blind upmix to HOA2 / HOA3, or 2) planar FOA with SPAR is implemented with DirAC upmix to full FOA,with possible blind upmix HOA2 / HOA3 with DirAC.
[0090] For FOA input, the bitrate where the number of SPAR downmix channels is 4.,1) FOA with SPAR is implemented using blind upmix to HOA2 / HOA3 using DirAC.
[0091] For FOA input, bitrates where the number of downmix channels in SPAR is less than 3, WY reconstruction with SPAR is implemented with an upmix to planar FOA, along with a possible blind upmix to planar HOA2 / HOA3 using DirAC.
[0092] For FOA input, the bitrate where the number of SPAR downmix channels is 3, WY reconstruction using SPAR is implemented using a blind upmix to the planar HOA2 / HOA3 using DirAC.
[0093] 2.4.3 Reconstruction of individual channels partially with SPAR and DirAC In conjunction with the channel-based SPAR-DirAC partitioning technique disclosed in Section 2.4.2, a further category of channels may be introduced that are partially parametrically reconstructed from both SPAR and DirAC methods. The motivation is to reduce the reliance on a large number of decorrelator outputs in the decoder, which may reduce the mix complexity figure. This approach uses SPAR prediction and cross prediction to reconstruct most of a particular parametrically reconstructed signal, and relies on the diffusive nature of DirAC to restore any missing covariances.
[0094] 2.4.4 Alternative methods to reduce the amount of decorrelator used in the decoder Instead of applying decorrelation proportional to the decorrelation coefficients, in this embodiment, cross-prediction parametrically constructed channel energy matching is achieved by applying gains derived from the SPAR coefficients. A particular Ambisonics signal S can be parametrically reconstructed as follows:
number
number
[0095] 2.4.5 Combining Frequency-Based and Channel-Based Division In some embodiments, sections 2.4.1 and 2.4.2 are combined to achieve the benefits of integrating SPAR and DirAC by performing a combination of frequency-based and channel-based splitting. In one implementation, the input to the integrated SPAR-DirAC system is an N-channel Ambisonics signal. Of these N channels, M channels are fed to the SPAR subsystem, where M<=N. In one embodiment, these M channels include FoA channels. In some other embodiments, these M channels include FOA and planar HOA channels. SPAR can then operate in any downmix configuration based on the operating bitrate, where the number of downmix channels N dmx , 1<=N dmx <=M. For low frequencies, SPAR calculates SPAR parameters including prediction parameters, cross prediction parameters and decorrelation parameters based on the method described in Section 2.3, while for high frequencies, DirAC parameters are calculated as described in Section 2.2 and SPAR parameters are estimated from the DirAC parameters as described in Sections 2.5-3.4 below. In some embodiments, SPAR also calculates SPAR parameters for high frequencies for a subset of input channels based on the method described in Section 2.3.
[0096] At the decoder side, the M channels reconstructed by SPAR are then used by DirAC to reconstruct a representation of the original N-channel input scene.
[0097] An example implementation of frequency-based and channel-based combined partitioning using HOA3 input to the SPAR-DirAC integrated system is given below.
[0098] 2.4.5.1 Encoder embodiment example 1 2 is a block diagram of an encoder 200 with frequency-based and channel-based split between SPAR and DirAC in accordance with one or more embodiments. In this embodiment, SPAR is operating in a 4-channel downmix mode. The input to the encoder 200 is a HOA3 (3rd order Ambisonics) signal. The DirAC parameter estimator 201 estimates the DirAC parameters, calculated according to section 2.2, limited to high frequencies and based on the FOA channel in the Ambisonics input. The estimated DirAC parameters are quantized and coded (202), and the quantized DirAC MD is converted to a SPAR MD (203).
[0099] The SPAR analysis and metadata calculation (204) is based on the FOA, HOA2 and HOA3 channels at low frequencies according to section 2.3. The SPAR metadata is quantized and encoded (205), and the SPAR MD obtained from the quantized SPAR metadata at low frequencies and the DirAC MD at high frequencies is transformed into a downmix matrix 206. An MDFT transform (207) is applied to the FOA, HOA2 and HOA3 signals. The MDFT coefficients and the downmix matrix are cross-faded and frequency band mixed using a filter bank mixer 208 to generate a four-channel downmix. The four-channel downmix is encoded by one or more core codecs 209 (e.g., Enhanced Voice Services (EVS) encoders). The SPAR metadata encoded at low frequencies and the DirAC metadata encoded at high frequencies are packed together with the core codec encoded bits to form the final bitstream 210 output by the encoder 200.
[0100] Encoder 200 is an example embodiment of a combined DirAC and SPAR encoder. In other embodiments, SPAR and DirAC are combined with frequency division only or with channel division only.
[0101] 2.4.5.1.2 Decoder embodiment example 1 3 is a block diagram of a decoder 300 with frequency-based and channel-based split between SPAR and DirAC in one or more embodiments. In this embodiment, the decoder 300 receives a bitstream 301 (210) and provides core codec encoded bits to one or more core codec decoders 307 (e.g., EVS decoders). The DirAC MD 302 at high frequencies is decoded and then converted to SPAR MD 303 at high frequencies using a DirAC MD to SPAR MD conversion (313). In an embodiment, 313 at the decoder is the same as 203 at the encoder. The SPAR MD in the bitstream is decoded to reconstruct SPAR metadata 304 at low frequencies. A SPAR upmix matrix 305 is generated using the low-frequency SPAR metadata 304 retrieved from the bitstream 310 and the high-frequency SPAR metadata 303 converted from the high-frequency DirAC metadata. The downmix channels are reconstructed by one or more instances of a core decoder 307 and transformed into the frequency banded domain by a filter bank 308 (eg, a CLDFB filter bank, a quadrature mirror filter bank (QMF), etc.).
[0102] In some embodiments, the primary downmix channel is input to a decorrelator 309, and the output of the decorrelator 309 is input to a SPAR upmix unit 306 together with an upmix matrix to reconstruct the FOA, planar HOA2 and planar HOA3 channels. The decorrelation can be implemented in the time domain or in the frequency banding domain (e.g., CLDFB domain). The decorrelator may generate a time domain decorrelated output and then transform the output to the frequency banding domain, or transform the input to the frequency banding domain and generate a frequency banding domain decorrelated output. The output channels of 306 are provided to a DirAC parameter estimator 310. The DirAC parameter estimator 310 estimates DirAC metadata at low frequencies based on the reconstructed FOA signal in the frequency banding domain. DirAC upmixer 311 uses low-frequency DirAC metadata and high-frequency DirAC metadata to upmix the FOA, planar HOA2, and planar HOA3 channels into 16 HOA3 channels that are frequency-banded domain representations of the original 16-channel HOA3 input to encoder 200. Synthesizer 312 (e.g., a CLDFB synthesizer) synthesizes / renders the 16-channel HOA3 frequency-banded domain representation into a time-domain representation for playback on various audio systems with different speaker configurations and capabilities.
[0103] 2.4.5.2.1 Encoder embodiment example 2 4 is a block diagram of an alternative encoder 400 with frequency-based and channel-based split between SPAR and DirAC in one or more embodiments. In this embodiment, the input to the encoder 400 is the HOA3 signal. DirAC parameters are estimated (401) and then quantized and coded (402). DirAC parameter estimation is limited to high frequencies and is performed according to section 2.2 based on the FOA channel. SPAR analysis and metadata calculation (404), as well as quantization and coding (405), are performed at low frequencies based on the FOA channel, the planar HOA2 channel, and the planar HOA3 channel, and zero or more non-planar channels (e.g., height channel) according to section 2.3.
[0104] For high frequencies, SPAR analysis and parameter estimation is performed for the non-FoA channels according to Section 3.2.7.2 (which is not done in system 200). In this embodiment, SPAR is operating in a four-channel downmix mode, and SPAR FoA metadata at high frequencies is estimated based on DirAC metadata using the method described in Section 3.2 to obtain SPAR downmix matrices for all frequencies.
[0105] The quantized and encoded SPAR metadata is used to generate a downmix matrix 407. An MDFT transform (406) is applied to the FOA, Plane HOA2, and Plane HOA3 signals. The MDFT coefficients and downmix matrix are cross-faded and frequency band mixed (408) to generate a four-channel downmix. The four-channel downmix is encoded by one or more core codecs 409. The SPAR metadata encoded at low frequencies for the FOA channels, the SPAR metadata encoded at all frequencies for the HOA channels, and the DirAC metadata encoded at high frequencies are packed with the core codec encoded bits to form the final bitstream 410 output by the encoder 400.
[0106] The downmixed channels are encoded (409) by one or more core codecs (e.g., EVS). For FOA channels, SPAR metadata is encoded for low frequencies but DirAC metadata is encoded for high frequencies, while for non-FOA channels, SPAR metadata is encoded for the entire frequency range and packed with core codec encoded bits to form the final bitstream 410 output by the encoder 400. In this embodiment, the SPAR metadata calculation for HOA2 and HOA3 channels at high frequencies is performed according to the method described in section 3.2.7.2. Furthermore, in this embodiment, the SPAR metadata calculation for HOA2 and HOA3 channels at high frequencies (404) depends on the SPAR MD for FOA channel at high frequencies estimated from DirAC MD at high frequencies (303), according to the method described in section 3.2.7.2.
[0107] Note that in embodiment 2, the conversion from DirAC MD to SPAR MD occurs only for the FOA channel, and full-band SPAR MD is used for any HOA channel handled by SPAR. In general, any number of non-planar HOA channels can be handled by SPAR. In embodiment 2, only one non-planar HOA channel was added. Also, in these embodiments, we focus on four downmix channels (Ndmx=4), but any number of transport channels (e.g., 1-16) are possible.
[0108] 2.4.5.2.2 Decoder embodiment example 2 FIG. 5 is a block diagram of an alternative decoder 500 using frequency-based and channel-based split between SPAR and DirAC in accordance with one or more embodiments.
[0109] In this embodiment, the decoder 500 receives an encoded bitstream 504 and provides core codec encoded bits to one or more core decoders 505. The DirAC MD 502 at high frequencies is decoded and then converted to SPAR MD 503 at high frequencies using a DirAC MD to SPAR MD conversion (513). In one embodiment, the DirAC MD to SPAR MD conversion (513) at the decoder is the same as the DirAC MD to SPAR MD conversion (403) at the encoder. The SPAR MDs 504 corresponding to the FOA and planar HOA channels, as well as zero or more non-planar HOA channels, are decoded and provided to a SPAR mix matrix 506. The missing SPAR MDs 503 for the FOA channels at high frequencies are estimated from the DirAC MD in the same manner as the encoder 400. The SPAR upmix matrix 506 is generated using the SPAR MDs 504 retrieved from the bitstream 510 and the high frequency SPAR MDs 503 converted from the high frequency DirAC MDs. The downmix channels reconstructed by one or more instances of the core decoder 505 are transformed into the frequency banded domain with the aid of a filter bank analysis 507 and an upmix matrix 506 is applied to reconstruct the FOA channel, the planar HOA2 channel, the planar HOA3 channel and zero or more non-planar (height) channels.
[0110] The decoded downmix channels output from one or more core decoders 505 are input to a decorrelator 509, and the output of the decorrelator 509 is input to a SPAR upmix unit 508 together with an upmix matrix to reconstruct the FOA channel, the planar HOA2 channel, and the planar HOA3 channel. The decorrelation can be implemented in the time domain or the frequency banding domain (e.g., CLDFB domain). The decorrelator may generate a time domain decorrelated output and then transform the output to the frequency banding domain, or transform the input to the frequency banding domain and generate a decorrelated output in the frequency banding domain. The output channels of 508 are provided to a DirAC parameter estimator 510. The DirAC parameter estimator 510 estimates DirAC metadata in low frequencies based on the reconstructed FOA signal in the frequency banding domain and uses DirAC parameters in high frequencies retrieved from the bitstream 501. Alternatively, the DirAC upmixer 508 may estimate DirAC parameters in the entire frequency range based on the FOA signal in the frequency band domain (eg, the CLDFB domain) and ignore the DirAC parameters in high frequencies from the bitstream 501.
[0111] DirAC upmixer 511 uses DirAC metadata from 510 and 502 to convert the FOA channels, the planar HOA2 channels, the planar HOA3 channels, and zero or more non-planar channels into a HOA3 output that is a frequency band domain (e.g., CLDFB domain) representation of the original 16-channel HOA3 input to encoder 400. Synthesizer 512 (e.g., CLDFB synthesizer) synthesizes / renders the 16-channel HOA3 frequency band domain representation into a time domain representation for playback on various audio systems with different speaker configurations and capabilities. Note that the output of decorrelator 509 is in the CLDFB domain to cover embodiments in which a time domain decorrelator is followed by CLDFB analysis and CLDFB analysis with CLDFB domain decorrelation.
[0112] 2.4.6 Diffusion in DirAC Upmix Channels When estimating higher-order channels from the first-order channel using the DirAC approach, directional panning in the higher-order channels that are upmixed using the DirAC approach can be done by using the DOA angle and the spherical harmonic response. However, adding diffuseness and decorrelation to these higher-order channels should be handled with caution, since it is known that too much decorrelation can impair sound quality, and too little decorrelation can cause spatial collapse.
[0113] Below we describe an embodiment for adding diffusivity to higher order channels upmixed by the DirAC approach.
[0114] 2.4.6.1 Applying uniform decorrelation to all HOA upmixed channels In some embodiments, N d N uncorrelated channels are calculated, where N d is the number of HOA channels to be upmixed by DirAC from the FOA channels. ψ (diffusivity) is calculated using any of the methods described herein, where
number
number
[0115] The above approaches can result in too much decorrelation, making the reconstructed scene more diffuse than desired, and generating too many decorrelator outputs and scaling them to obtain the desired level of diffuseness can be computationally expensive.
[0116] 2.4.6.2 Applying directional decorrelation to all upmixed channels In this embodiment, directional diffuseness information is sent from the encoder to the decoder. The decoder uses this directional diffuseness information to add only the desired amount of decorrelation to the upmixed HOA channels. This method is applicable when the input to the encoder is HOA, and due to bitrate and complexity limitations, only a selected few channels are reconstructed using SPAR, while the remaining channels are upmixed using DirAC. In one implementation example, the encoder can calculate the directional diffuseness using the P (decorrelation) coefficient calculated by SPAR in section 2.3. This method uses the side information to be sent from the encoder to the decoder.
[0117] 2.4.6.3 Adding decorrelation to selected upmix channels In this embodiment, the addition of diffuseness is restricted to a small number of selected channels to keep the overall diffuseness within a desired range. This method also reduces computational complexity. The selection of channels for adding diffuseness can be done statically or dynamically based on signal characteristics.
[0118] 2.4.6.3.1 Static Channel Selection
[0119] In this embodiment, decorrelation is applied to a selected few HOA channels. These channels are selected based on perceptual importance. In one implementation example, if the FOA and planar HOA channels are reconstructed by SPAR and only the non-planar HOA channels are to be upmixed using DirAC to obtain HOA3 output in ACN-SN3D format, channel indexes 6, 10, 12, 14 (channel index ranges from 0 to 15) may be selected to apply decorrelation. This method does not require sending any side information to the decoder.
[0120] 2.4.6.3.2 Dynamic Channel Selection In this embodiment, directional diffuseness information is calculated in the encoder and sent to the decoder to select the channels to which diffuseness should be added during the upmix. This embodiment is only applicable when the input to the encoder is HOA. In the DirAC decoder, only channels that require a higher amount of decorrelation than a first threshold are selected and decorrelation is added. In one implementation example, the encoder calculates directional diffuseness using the P (decorrelation) coefficient calculated by SPAR in section 2.3, compares the P coefficient value with a first threshold, and encodes the channel indexes that have a P coefficient higher than the first threshold. These indexes are read by the decoder. If the number of channel indexes exceeds a second threshold, limited indexes may be selected based on the P coefficient value and the perceptual importance of a given channel. This embodiment requires additional information to be transmitted from the encoder to the decoder.
[0121] To perform the frequency-based split described in section 2.4.1, an efficient mechanism is desired to convert DirAC metadata to SPAR in the DirAC frequency bands and SPAR metadata to DirAC in the SPAR frequency bands, so that the DirAC and SPAR metadata can be reconstructed in all bands when an upmix or downmix needs to be performed. Below are example embodiments for converting DirAC metadata to SPAR and SPAR metadata to DirAC.
[0122] 2.5 Conversion from DirAC to SPAR In some embodiments, an approximation of the input covariance matrix is calculated based on the quantized DirAC MD parameters (azimuth angle (Az), elevation angle (El), diffusivity). Herein, Az and El are expressed as the DOA angle θ D It is also called.
[0123] 2.5.1 Formula In some embodiments, the model covariance block calculates the covariance matrix and prediction coefficients from the DirAC DOA and diffusivity as follows:
[0124]
number
[0125] 2.5.2 Example of covariance calculation In some embodiments, the covariance is calculated as follows:
number
number
[0126] In the above formula, E is an approximation of the total signal energy (given in
[33] below). It is obtained by adding rough estimates of the directional and diffuse energies. w r Let be the real bin samples of the W channel in the MDFT domain, then the energy corresponding to each bin is calculated as follows:
number
[0127] The energy is then converted to frequency banded power by applying a filter bank response for each band. The frequency banded energy for each band is extrapolated to calculate the total signal energy as follows:
number
[0128] The diffuse energy components are added as follows:
number
[0129] In some embodiments, if covariance smoothing is turned off, the above calculated covariance is used to calculate the SPAR coefficients as usual.
[0130] 3.0 Other Embodiments 3.1 DirAC MD calculation 3.1.1 Improved DirAC Diffusivity Calculation DirAC requires time smoothing to compute the diffusivity parameters. In some embodiments, a simple parameter averaging is performed over 160 ms (Equation 12 in Section 2.2.2.2).
[0131] In some embodiments, SPAR's covariance smoothing and / or transient detector-ducker algorithms can be used to improve the calculation of the DirAC diffuseness parameters. For example, SPAR's covariance smoothing algorithm, described in PCT Application No. PCT / 2020 / 044670, "Systems and Methods for Covariance Smoothing," filed July 31, 2020, can be configured to weight recent audio events more heavily than more past events, and can do so differently in each frequency band. This can be more advantageous than a simple averaging operation. Using transient detection and ducking, diffuseness can be reduced instantaneously during short transients without interfering with the long-term smoothing process.
[0132] Since temporal smoothing introduces long-term time dependencies in addition to smoothness over time, in another embodiment differential coding can be used to reduce the MD bitrate and improve frame loss robustness.
[0133] 3.1.2 Calculation of DirAC Metadata in the Frequency Banded Covariance Domain Based on the DirAC analysis described in Section 2.0, in some embodiments, DirAC MD can be calculated based on the input frequency-banded covariance matrix instead of calculating DirAC MD in the FFT (Fast Fourier Transform) domain or MDFT domain and then transforming it to the frequency-banded domain.
[0134] In some embodiments, the calculation of the SPAR metadata can be based on the input frequency banded covariances, as described in Section 2.0.
[0135] In some embodiments, by computing both SPAR and DirAC metadata from the input covariances, the conversion from SPAR to DirAC and from DirAC to SPAR MD in the desired bands can be made better and more computationally efficient. Below is an example of how DirAC MD can be calculated from the input covariances: 1. Compute an N*N frequency banded covariance matrix, where N is the number of input channels. 2. Smooth the covariance matrix as described in Section 3.1.1. 3. Calculate the criterion power as the trace of the covariance matrix. 4. R, which is strength wx , R wy , R wz Calculate where R wx , R wy , R wz is the covariance of the W channel and the X, Y, and Z channels. 5. It is the strength norm Calculate JPEG2025507160000028.jpg1046. 6. It is a directional vector JPEG2025507160000029.jpg1035, where s can be x, y, z. Then calculate the azimuth and elevation angles according to equations [5] and [6] in Section 2.2.2.1.
[0136] Similarly, the calculation of diffuseness can be done based on the frequency banded covariance matrix as follows:
[0137] For diffuseness, first the reference power E and intensity I of the input signal are calculated in a given frequency band.
number
[0138] Since the covariance has already been smoothed as described in section 3.1.1, the diffuseness can be calculated as follows:
number
[0139] In some embodiments, before calculating the diffusivity, E and I are further averaged using a long-term averaging filter as given below.
number
[0140] Here, E a and I a are the long-term average values of energy and intensity, respectively, and these values are then used in place of E and I in the calculation of the diffusivity formula
[36] . The factor f in
[39] and
[40] e and f i is an example of a smoothing factor.
[0141] 3.1.3 Improvement of reference power (E) calculation In some embodiments, an alternative method for calculating the reference power can be used that provides a better estimate of the diffuseness and, when derived from the DirAC coefficients, provides a better estimate of the SPAR coefficients.
[0142] For diffuseness, first the reference power E and intensity I of the input signal are calculated in a given frequency band.
number
[0143] Here, R ij is the covariance between the i-th and j-th channels. The reference power is calculated as follows:
number
[0144] E, calculated in
[43] , provides better estimates for the diffuseness and SPAR coefficients when the energy of the W channel is higher than 0.5*E. The diffuseness is calculated as follows:
number
[0145] Here, E a , I ax , I ay , I az E, I x , I y , I z Alternatively, E a , I ax , I ay , I az can also be calculated based on the covariance matrix, where the diffusivity can be bounded as follows:
number
[0146] The SPAR coefficients can be calculated from the DirAC coefficients using any of the methods described herein.
[0147] 3.2 Improved conversion from DirAC to SPARMD 3.2.1 Alternative method to compute covariance / spar MD from DirAC In some embodiments, the passive prediction coefficients are also i *Resp j where i and j can be w, x, y, z and should be similar to the direction vector dv for a given side channel. Thus, the prediction coefficients are calculated based on the variance of the W channel in the frequency banded domain. norm In some embodiments, the variance of the W channel is smaller than I norm For better estimates of the predictive coefficients, the additional parameter JPEG2025507160000037.jpg87 may be sent to the decoder. In some embodiments, the prediction coefficients may also be calculated directly from the DirAC metadata.
[0148] 3.2.2 DirAC Metadata Quantization In some embodiments, the SPAR MD is calculated based on the quantized DirAC MD.
[0149] 3.2.3 Generic reconstruction of SPAR coefficients from DirAC metadata for any downmix configuration In some embodiments, the input covariance R is a 4x4 matrix calculated based on the DirAC parameters as follows:
number
[0150] The SPAR coefficients are normalized with respect to the covariance, so that the SPAR coefficients obtained from the input covariance R are equal to the SPAR coefficients obtained from E*R, where E can be either the variance of the W channels or the total signal energy or any constant.
[0151] In some embodiments, a normalized covariance matrix R_norm is derived based only on the DirAC parameters. R_norm is a 4x4 covariance matrix for the FOA channel and is an approximation of the actual normalized input covariance matrix, where the actual input covariance matrix is given by:
number
number
number
[0152] The SPAR coefficients, including prediction coefficients, cross-prediction coefficients and decorrelation coefficients, are calculated using the normalized covariance R_norm as disclosed in Section 2.3. ij It is calculated from.
[0153] 3.2.3.1 Example of direct reconstruction of prediction and decorrelation coefficients from DirAC metadata for a one-channel downmix From the above normalized covariance matrix, the SPAR coefficients can be calculated based on the calculations in Section 2.3 as follows:
[0154] The prediction coefficients are calculated as follows:
number
[0155] For a one-channel downmix, the decorrelation coefficients are calculated as follows:
number
[0156] Here the decorrelation coefficients depend on the spherical harmonic response. To avoid this dependency, 3.2.4 can be used.
[0157] 3.2.4 Another variant of the global reconstruction of SPAR coefficients from DirAC metadata for any downmix configuration In this embodiment, based on the DirAC parameters, the actual input covariance R in A 4x4 covariance matrix R, which is an approximation of , is calculated as follows: where the elements of the matrix are approximated as follows:
number
[0158] Since the SPAR coefficients are normalized, they are derived from R in the same way that the SPAR coefficients are derived from E*R, where E can be the variance of just the W channels or the total signal energy or any constant.
[0159] The elements of the normalized 4x4 covariance matrix for the FoA channels are derived based only on the DirAC parameters.
number
[0160] The SPAR coefficients, including prediction coefficients, cross-prediction coefficients and decorrelation coefficients, are calculated from the R_norm as disclosed in Section 2.3.
[0161] 3.2.4.1 Example of direct reconstruction of prediction and decorrelation coefficients from DirAC metadata for a one-channel downmix From the above normalized covariance, the SPAR coefficients can be calculated based on the calculations in Section 2.3 as follows:
[0162] The prediction coefficients can be calculated as follows:
number
[0163] Then, for a one-channel downmix, the decorrelation coefficient can be calculated as follows:
number
[0164] Here, the decorrelation coefficients do not depend on the spherical harmonic response, but only on the diffusivity and some constants.
[0165] Calculating the constant "c" - Solution 1 In one implementation, to further improve the prediction coefficients, (1-cψ) can be set such that the passive W prediction coefficients are:
number
[59]
[0166] Based on equation
[53] , this results in a value of c that is:
number
[0167] In one embodiment, to improve the SPAR coefficients calculated from DirAC MD, the actual normalized input covariance R_norm in A 4x4 covariance matrix R_norm, which is an approximation of , is calculated based on the DirAC parameters as follows: where the elements of the matrix are approximated according to
[54] and
[61] as given below:
number
[0168] For a one-channel downmix, based on equations
[62] -
[64] , the decorrelation coefficients are calculated as follows:
number
[0169] In some embodiments, Q x , Q y , Q z The value can be set to 1 / 3.
[0170] Calculating the constant "c" - Solution 2 In another example embodiment, c can be calculated as follows:
number
[0171] This value of c is substituted into the prediction coefficient calculation formula.
number
[66]
[0172] The prediction coefficients
[66] are similar to the passive prediction coefficients calculation disclosed in Section 2.3.1.1. For this solution, the value of c can be transmitted to the decoder.
[0173] 3.2.5 Energy Compensation for DirAC-Based Downmix The covariance calculation from DirAC metadata (MD) in the solution disclosed in Section 2.5.2 and described in Sections 3.2.3 and 3.2.4 assumes that the signals are fully SN3D normalized as follows:
number
[0174] This assumption is incorrect in real FoA capture, e.g., in overtalk situations, capturing diffuse background noise, etc. The above method results in spatial collapse, especially when the number of downmix channels is limited to one.
[0175] Energy compensation can be applied to prevent spatial collapse by scaling the downmix signal such that the upmixed signal is energy matched with the input. Below is an example implementation of energy compensation with a 1-channel downmix:
[0176] The actual input covariance matrix JPEG2025507160000057.jpg511, where N is the number of input channels, JPEG2025507160000058.jpg88 is calculated to be the frequency banded or wideband covariance of the i-th input channel and the j-th input channel. For FOA inputs, N=4, and i and j can be W, X, Y, Z.
[0177] The normalized actual input covariance matrix R_norm_in NxN is calculated as follows:
number
[0178] Normalized covariance estimate R_norm based on DirAC metadata NxN is calculated according to any of the techniques described in Sections 2.5.2, 3.2.3, and 3.2.4.
[0179] The scaling factor is obtained as follows:
number
[0180] Here, thresh low and thresh high are the lower and upper bounds on the scale factor. low =1 and thresh high =2.
[0181] The SPAR downmix matrix and the SPAR coefficients (including prediction, cross-prediction and decorrelation coefficients) are calculated as disclosed in section 2.3 using the DirAC estimated normalized input covariance matrix.
[0182] Downmix matrix 1xN The downmix matrix is scaled by the scale calculated in Equation
[70] in Section 3.2.5. The actual downmix matrix is Downmix_act 1xN Then,
number
[0183] In one embodiment, for a one-channel downmix, 1xN is given according to equation
[72] as follows:
number
[0184] Here, F W , F Y , F z , F x are the gains used to mix the Y, Z, and X channels respectively into the W channel to form the downmix channel. After being scaled with the "scale" value, the downmix channel is calculated as follows:
number
[0185] In another implementation, F W =1, F Y =F z =F x = 0 and W′ = scale*W. F W , F Y , F z , F x Another implementation example using the calculation is described in Section 3.3.
[0186] The metadata parameters are not changed with this scaling: the encoder encodes the metadata parameters and the scaled downmix and the bitstream are sent to the decoder.
[0187] The decoder decodes the scaled downmix channel W′ and the spatial parameters, including the prediction and decorrelation parameters, and applies the prediction and decorrelation parameters to reconstruct the original input scene as follows:
number
[0188] Here, pr x , pr y , pr z are the prediction parameters, and p x , p y , p z is the decorrelation parameter, D1(W′), D2(W′), D3(W′) are the three decorrelated channels decorrelated with respect to W′, and f s is the active scaling described in Section 3.3. In one implementation, 0≦f s <= 1.
[0189] In this approach, we energy match the reconstructed scene with the input by scaling the reconstructed signal by the scale factor calculated in equation
[70] of this section, without transmitting any additional parameters in the bitstream.
[0190] 3.2.6 Extrapolation of directional diffusivity in the DirAC band The covariance estimate based on DirAC assumes uniform diffusivity in all directions, which may not be correct for real signals such as overtalk scenarios. Adding directional information onto the diffusivity parameters calculated in equation [7] in Section 2.2 can result in additional metadata to be coded in the bitstream. SPAR provides directional diffusivity information in its metadata, and the directional information in the high bands can be extrapolated using the directional information in the low bands.
[0191] In an example embodiment for an FOA input with 1 channel downmix, if SPAR encodes up to a 6 kHz frequency band and DirAC parameters are transmitted for the frequency range 6-24 kHz, the directional information in the SPAR frequency band is extracted as follows:
number
[0192] Here, p x , p y and p z is the SPAR decorrelation parameter in the last SPAR band.
[0193] This directional information can be used in the high frequency bands while computing the downmix using DirAC parameters. An example of estimating the normalized covariance matrix from DirAC metadata with directional spread is shown below. R_norm is a 4x4 matrix for the FOA channels and is calculated as follows:
number
[0194] The downmix matrix and the SPAR coefficients, including prediction coefficients, cross-prediction coefficients and decorrelation coefficients, are calculated from the R_norm as disclosed in Section 2.3. Examples of calculation of prediction coefficients and decorrelation coefficients for a one-channel downmix are given in
[55] –
[58] . The downmix matrix can be further scaled according to
[70] so that the reconstructed Ambisonics signal at the decoder has a good energy match to the Ambisonics signal at the encoder input.
[0195] 3.2.7 DirAC to SPAR Metadata Conversion for HoA Channels 3.2.7.1 Estimating the HoA input covariance matrix from DirAC parameters In this method, the DirAC parameters are used to estimate the input covariance matrix.
[0196] An NxN covariance R is calculated based on the DirAC parameters, where N is the number of input channels in the HOA signal, and where R is an approximation of the actual input covariance matrix. In some embodiments, the covariance R may be calculated as follows:
number
[0197] Here, Resp i is a spherical harmonic function, and Q i and c is a constant in the range of 0 to 1, e.g., c=1 and Q i = 1 / 3, i.e., 0<=i<=3 for the secondary channel, Q i = 1 / 5, i.e., 4<=i<=8 for the third channel, Q i = 1 / 7, i.e., 9 <= i <= 15. In this case, both the encoder and the decoder have prior knowledge of these constants. i and c are dynamically calculated based on the actual input covariance matrix and the above approximation of the input matrix from the DirAC parameters.
[0198] The SPAR coefficients are normalized so that a SPAR coefficient derived from R is equal to a SPAR coefficient derived from E*R, where E can be the variance of just the W channels or the total signal energy or any constant.
[0199] The covariance R in equation
[81] is normalized and the elements of this NxN normalized covariance matrix R_norm are derived based only on the DirAC parameters as follows:
number
[0200] The SPAR coefficients, including prediction coefficients, cross-prediction coefficients and decorrelation coefficients, are calculated from the R_norm as disclosed in Section 2.3.
[0201] 3.2.7.2 Improving spatial resolution of DirAC to SPAR conversion by restricting DirAC covariance estimation to FOA channels only It has been observed that covariance estimation for HOA channels from DirAC parameters is not optimal when there is significant information in the HOA channel. A loss of ambience has been observed when estimating the entire NxN covariance matrix (or all HOA SPAR parameters) from DirAC parameters. For such HOA signals, a different approach is desired. In the following, several embodiments for DirAC to SPAR conversion with improved spatial resolution are described.
[0202] 3.2.7.2.1 SPAR by independently computing and encoding HoA parameters In this method, the DirAC parameters are used to estimate the input covariance matrix for only the FOA channels, and then the SPAR parameters corresponding to the FOA channels are calculated from the estimates, by the methods described in Sections 3.2 and 3.2.4.
[0203] The SPAR parameters, including prediction coefficients, cross-prediction coefficients and decorrelation coefficients for the HoA channel, are calculated independently based on the actual covariance matrix of the input signal according to the method described in Section 2.3.
[0204] This method requires that the SPAR HoA parameters be encoded into the bitstream for all frequencies.
[0205] 3.2.7.2.2 Alternative Calculation of SPARHOA Parameters Based on DirAC-Estimated FOA This method is applicable to SPAR modes where the number of downmix channels is less than the number of input channels to SPAR, i.e., SPAR has cross-prediction and / or decorrelation coefficients for coding HOA channels. In this method, DirAC parameters are used to estimate the input covariance matrix for FOA channels only, and then the SPAR parameters corresponding to FOA channels are calculated from the input covariance matrix. This is done by the methods described in Sections 3.2. and 3.2.4.
[0206] Calculating HOA Prediction Coefficients The SPAR prediction coefficients for the HOA channel are calculated independently based on the actual covariance matrix of the input signal according to the method described in Section 2.3.
[0207] Calculating HOA cross-prediction coefficients Section 2.3 shows that the cross-prediction coefficients in SPAR MD depend on the predicted side channel or residual channel in the downmix. Furthermore, the residual channel in the FOA component of the Ambisonics input depends on the SPAR MD derived from DirAC MD in a set of frequency bands. Therefore, it is confirmed that the cross-prediction coefficients in the HOA channel can depend on the DirAC MD in the FOA channel, and calculating the cross-prediction coefficients in the HOA channel based on the DirAC MD in the FOA channel and the SPAR MD of the FOA and HOA channels can lead to a better estimation of these coefficients. In one implementation example, the HOA channel (4 to N) prediction coefficients are calculated from the actual input covariance matrix as described in Section 2.3. These prediction coefficients are quantized based on a quantization strategy. The FoA prediction coefficients estimated by DirAC, together with the HOA quantized prediction coefficients estimated by SPAR, are used to generate the downmix matrix as described in Section 2.3. The post-prediction covariance matrix is calculated from the actual input covariance matrix and the downmix matrix calculated above. The cross-prediction coefficients are then calculated from the post-prediction matrix as described in Section 2.3.
[0208] Calculating the HOA decorrelation coefficient It has been confirmed that calculating the HOA decorrelation coefficients directly from the Ambisonics input covariances, as described in Section 2.3, without relying on the DirAC MD in the FOA channels, leads to a better estimation of the decorrelation coefficients and the desired decorrelation amount in the reconstructed HOA channels at the decoder. This helps to reduce audio artifacts that may occur due to too much decorrelation, and also avoids spatial collapse due to too little decorrelation. In one implementation example, first, prediction coefficients corresponding to all side channels are calculated from the actual input covariance matrix, as described in Section 2.3. Then, the calculation of the decorrelation coefficients from the prediction coefficients and covariance matrix is the same as in Section 2.3. This method encodes the parameters of the SPAR HoA in the bitstream for all frequencies.
[0209] 3.3 Active W Downmix Based on DirAC Metadata 3.3.1 Based on DirAC-based covariance estimation From the DirAC metadata, the input covariance can be estimated as the DirAC metadata-based input signal (4x4) covariance matrix estimate, as given in Section 3.2.3 or Section 3.2.4.
number
number
[0210] Alternatively, S can be calculated as given in Section 3.2.4 as follows:
number
[0211] One possible approach to do an active downmix based on the above covariance matrix is by having the following prediction matrix:
number
[87]
[0212] Then the post-prediction matrix may be given as:
number
[0213] The other elements of the matrix in
[89] are not shown as they are not relevant for the active downmix gain calculation.
[0214] By setting JPEG2025507160000077.jpg59 Minimizing JPEG2025507160000078.jpg52 results in a linear expression given by:
number
[0215] Where: JPEG2025507160000080.jpg575. Substituting the values in
[90] , E is cancelled in the numerator and denominator, and g can be calculated directly from the DirAC metadata on both the encoder and decoder sides.
[0216] The actual downmix matrix for a 1-channel downmix after scaling is given by:
number
[91]
[0217] The post-predictive scaling factor 'r' is calculated at the decoder by matching the reconstructed W variance with the variance of the W encoder input.
number
[0218] The scaled prediction coefficients are calculated as follows:
number
[0219] Where: JPEG2025507160000085.jpg530 are the active prediction coefficients. The calculation of the decorrelation coefficients is as follows:
number
[91] and the decorrelation coefficients are Post_prediction [4x4] It is calculated from the following:
number
[0220] Here, Res uuis a 3x3 matrix, equal to Post_prediction[2:4,2:4], and P=[p x ;p y ,p z ] is the decorrelation coefficient.
[0221] The calculation of the active W downmix channel from the FOA input [W, Y, Z, X] is given by:
number
[70] , and the computation of another scale factor r is given in
[92] . W′ is coded using the core coder, DirAC MD is coded, and these coded bits are sent jointly to the decoder.
[0222] The inverse prediction matrix at the decoder is given by:
number
[0223] The reconstruction of the FOA channels at the decoder is as follows.
number
[0224] Here, pr x , pr y and pr z are prediction parameters calculated from DirAC MD as given in
[90] , and p x , p y and p z are the decorrelation parameters calculated from DirAC MD as given in
[94] , D1(W′), D2(W′), D3(W′) are the three decorrelated channels decorrelated with respect to W′, and f s is the scaling constant used in
[92] .
[0225] 3.4 Conversion of SPAR to DirAC Metadata It may be desirable to convert SPAR MD to DirAC MD in a set of frequency bands so that DirAC MD is available in all required frequency bands for upmixing to the desired output format at the decoder. Also, a direct conversion from SPAR MD to DirAC MD saves complexity. In one implementation, it is possible to derive the direction vector dv from the prediction coefficients.
number
[0226] The azimuth and elevation angles can then be calculated based on equations [5] and [6].
[0227] 3.4.1 Diffusivity calculations from SPAR metadata Assuming that SPAR perfectly reconstructs the covariance (COV) matrix, the output covariance matrix can be calculated at the decoder from the input (DMX+decorrelator) covariance matrix and the upmix matrix. From the output COV, the reference power and intensity are calculated and averaged over N frames (e.g., 8 frames). From there, the diffuseness is calculated according to equation [7].
[0228] As disclosed below, there are other embodiments that compute DirAC diffusivities directly from the SPAR metadata, without computing the output covariance matrix.
[0229] 3.4.1.1 Alternative for diffuseness to 1-channel downmix using passive W downmix (where the W channel in the downmix is the same as the W channel in the input, or simply a delayed version) Let w, x, y, z be the variances of W, X, Y, and Z. In a one-channel downmix, y can be approximated as follows:
number
number
[0230] The strength can be calculated as follows:
number
[0231] Referring to equation [7], the diffusivity ψ can be directly approximated from the SPAR metadata as follows:
number
[0232] Here, pr slow,s pr s Same as or pr s It can be the long-term average of pd slow,s pd s Same as or pd s where s can be x, y, z.
[0233] 3.4.1.2 Alternative methods for diffuseness for 1-channel downmix with active W downmix The inverse matrix with active W calculation described in Section 3.3.1 is given.
number
[0234] Let w, x, y, z be the variances of W, X, Y, Z. In the case of 1-channel downmix, y can be approximated as follows:
number
[0235] Here, pr y is the predicted coefficient, and pd y is the decorrelation coefficient for the Y channel. Similarly, x and z can be calculated.
[0236] In this case, the reference power can be calculated as (w+x+y+z).
number
[0237] The strength can be calculated as follows:
number
[0238] Referring to equation [7], the diffusivity ψ can be approximated directly from the SPAR metadata as follows (when averaging w separately):
number
[0239] Here, pr slow,s pr s Same as or pr s It can be the long-term average of pd slow,s pd s Same as or pd s where s can be x, y, z.
[0240] 3.4.1.3 Alternative methods for diffuseness for any passive W downmix channel configuration The method is based on normalization of the input Ambisonics signal. For example, if the FoA input is normalized using Schmidt half-normalization (SN3D), assume that w=x+y+z, where w, x, y, and z are the variances of the W, X, Y, and Z channels, respectively. This results in w+x+y+z=2*w.
[0241] Substituting the dispersion assumptions and intensities from Eq.
[0108] in Section 3.4.1.1 into the diffusivity expression in Eq. [7] gives:
number
[0242] Here, pr slow,s pr s Same as or pr s where s can be x, y, z.
[0243] Encoding process example Figure 6 is a flow diagram of a process 600 for encoding an FOA input using an encoder as described with reference to Figures 2 and 4 in some embodiments. Process 600 can be implemented using the electronic device architecture described with reference to Figure 9.
[0244] The process 600 includes: receiving a multi-channel audio signal including a first set of channels (601); for the first set of frequency bands: computing directional audio coding (DirAC) metadata from the first set of channels (602); quantizing and encoding the DirAC metadata (603); converting the quantized and encoded DirAC metadata into two or more parameters of a first spatial reconstruction (SPAR) metadata (604); and for a second set of frequency bands lower than the first set of frequency bands: computing second SPAR metadata from the first set of channels (606). quantizing and encoding the second SPAR metadata (607); generating a downmix based on the first SPAR metadata and the second SPAR metadata (608); calculating frequency coefficients from the first set of channels (609); downmixing from the coefficients and the downmix to a second set of channels (610); encoding the second set of channels (611); and outputting a bitstream including the encoded second set of channels, the quantized and encoded second SPAR metadata, and the quantized and encoded DirAC metadata (612). Each of these steps was described above with reference to Figures 2 and 4.
[0245] Figure 7 is a flow diagram of a process 700 for encoding FOA+HOA input using the encoder described with reference to Figures 2 and 4 in some embodiments. Process 700 can be implemented using the electronic device architecture described with reference to Figure 9.
[0246] The process 700 includes: receiving a multi-channel audio signal including a first set of channels and a second set of channels different from the first set of channels (701); for the first set of frequency bands: calculating directional audio coding (DirAC) metadata from the first set of channels (702); quantizing and encoding the DirAC metadata (703); converting the quantized and encoded DirAC metadata into two or more parameters of first spatial reconstruction (SPAR) metadata (704); for a second set of frequency bands lower than the first set of frequency bands: calculating second SPAR metadata from the first set of channels and the second set of channels (705). The method includes: computing frequency coefficients from the first set of channels and the second set of channels (705); quantizing and encoding the second SPAR metadata (706); generating a downmix based on the first SPAR metadata and the second SPAR metadata (707); computing frequency coefficients from the first set of channels and the second set of channels (708); downmixing from the coefficients and the downmix to a third set of channels (709); encoding the third set of channels (710); and outputting a bitstream including the encoded third set of channels, the quantized and encoded second SPAR metadata, and the quantized and encoded DirAC metadata (711). Each of these steps was described above with reference to Figures 2 and 4.
[0247] Figure 8 is a flow diagram of a process 800 for decoding using a codec as described with reference to Figures 3 and 5 in some embodiments. The process 800 can be implemented using the electronic device architecture described with reference to Figure 9.
[0248] The process 800 includes: receiving an encoded bitstream including encoded audio channels and metadata, the metadata including first directional audio coding (DirAC) metadata associated with a first frequency band and first spatial reconstruction (SPAR) metadata associated with a second frequency band lower than the first frequency band (801); decoding and dequantizing the first DirAC metadata and the first SPAR metadata (802); for the first frequency band: converting the dequantized DirAC first metadata into two or more parameters of a second SPAR metadata (803); blending the first and second SPAR metadata into composite SPAR metadata (804); decoding the encoded audio channels (805); The method includes reconstructing downmix channels from the composite audio channels (806); transforming the downmix channels to a frequency banded domain (807); generating a SPAR upmix based on the composite SPAR metadata (808); upmixing the downmix channels in the frequency banded domain to a first set of channels based on the SPAR upmix (809); estimating second DirAC metadata in a second frequency band from the first set of channels and zero or more parameters in the first SPAR metadata (810); upmixing the first set of channels to a second set of channels in the frequency banded domain based on the first and second DirAC metadata (811); and transforming the second set of channels from the frequency banded domain to the time domain (812).
[0249] System architecture example 9 is a block diagram of an example electronic device architecture 900 suitable for implementing example embodiments of the present disclosure. Architecture 900 includes, but is not limited to, servers and client devices, as described above with reference to FIGS. 1-8.
[0250] As shown, the architecture 900 includes a central processing unit (CPU) 901 capable of executing various processes according to a program stored in, for example, a read-only memory (ROM) 902 or a program loaded from, for example, a storage device 908 into a random access memory (RAM) 903. Data required when the CPU 901 executes various processes is also stored in the RAM 903 as necessary. The CPU 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0251] Connected to the I / O interface 905 are components such as: an input unit 906, which may include a keyboard, a mouse, etc., an output unit 907, which may include a display, such as a liquid crystal display (LCD), and one or more speakers, a storage unit 908, which may include a hard disk or other suitable storage device, and a communication unit 909, which may include a network interface card, such as a network card (e.g., wired or wireless).
[0252] In some implementations, the input unit 906 includes one or more microphones in different positions (depending on the host device) that enable capturing audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0253] In some implementations, the output unit 907 includes a system having a varying number of speakers. The output unit 907 (depending on the capabilities of the host device) can render audio signals in a variety of formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0254] In some embodiments, the communication unit 909 is configured to communicate with other devices (e.g., via a network). Also, the drive 710 is connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive, or other suitable removable medium, is mounted on the drive 710, and a computer program read therefrom is installed in the storage unit 908 as needed. Although the system 900 is described as including the above components, in actual applications, it will be understood by those skilled in the art that some of these components can be added, removed, and / or replaced, and all such modifications or changes are within the scope of the present disclosure.
[0255] According to example embodiments of the present disclosure, the above-described processes may be implemented as a computer software program or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the method. In such an embodiment, the computer program may be downloaded and loaded from a network via a communication unit 709 and / or installed from a removable medium 911, as shown in FIG. 9.
[0256] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits (e.g., control circuitry), software, logic, or any combination thereof. For example, the above-mentioned units may be executed by control circuitry (e.g., CPU 901 in combination with other components of FIG. 9), and thus the control circuitry may perform the operations described in the present disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device (e.g., control circuitry). Although various aspects of example embodiments of the present disclosure are illustrated and described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, apparatus, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controller or other computing device, or any combination thereof, as non-limiting examples.
[0257] Additionally, the various blocks illustrated in the flowcharts may be considered as method steps and / or operations resulting from computer program code operations and / or as multiple combined logic circuit elements configured to perform the associated functions. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine-readable medium, the computer program including program code configured to perform the methods described above.
[0258] In the context of this disclosure, a machine-readable medium can be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, an instruction execution apparatus, or an instruction execution device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be non-transitory and can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of machine-readable storage media can include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0259] Computer program codes implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus having control circuitry such that, when the program code is executed by the processor of the computer or other programmable data processing apparatus, the program code implements the functions / operations specified in the flowcharts and / or block diagrams. The program code may be executed entirely on the computer as a stand-alone software package, or partly on the computer, partly on the computer and partly on a remote computer, entirely on a remote computer or remote server, or distributed over one or more remote computers and / or remote servers.
[0260] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or the scope that may be described in the claims, but rather as a description of features specific to a particular embodiment of a particular invention. Certain features described herein with respect to separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described with respect to a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, although features may be described above as operating in a particular combination and may even be initially claimed as such, one or more features from a combination of claims may in some cases be carved out of that combination, and the combination of claims may be directed to a subcombination or subcombination variation. The illustrated logic flow does not require the particular order depicted or sequential order to achieve the desired results. In addition, other steps may be provided to the described flow, steps may be deleted, and other components may be added or deleted to the described system. Accordingly, other implementations are within the scope of the appended claims.
Claims
1. A multi-channel audio signal containing a first set of channels is received by at least one processor, Regarding the first set of frequency bands, The at least one processor calculates directional audio coding (DirAC) metadata from the first set of channels, The at least one processor quantizes the DirAC metadata, wherein the DirAC metadata includes DirAC metadata parameters. The at least one processor encodes the quantized DirAC metadata, The at least one processor converts the quantized DirAC metadata into two or more parameters of a first spatial reconstruction (SPAR) metadata, wherein the first SPAR metadata includes two or more parameters of SPAR metadata parameters. Regarding a second set of frequency bands that are lower than the first set of frequency bands, The process involves calculating a second SPAR metadata from the first set of channels using the at least one processor, wherein the second SPAR metadata includes at least one parameter of the SPAR metadata parameters. The at least one processor quantizes the second SPAR metadata, The at least one processor encodes the quantized second SPAR metadata, The at least one processor generates a downmix from the first set of channels based on the first SPAR metadata and the second SPAR metadata, The at least one processor calculates the frequency coefficient from the first set of channels, The at least one processor downmixes the coefficients and the downmix into a second set of channels, The at least one processor encodes the second set of channels, Outputting a bitstream including the encoded second set of channels, the quantized and encoded second SPAR metadata, and the quantized and encoded DirAC metadata, Includes, The SPAR metadata parameters include a prediction parameter PR, a cross-prediction parameter C, and a decorrelation parameter P. The DirAC includes at least one diffusivity parameter, method.
2. The first set of channels is a primary ambisonic (FOA) channel. The method according to claim 1.
3. One or more parameters in the first SPAR metadata for the first set of frequency bands are encoded in the bitstream rather than being converted from the DirAC metadata. The method according to claim 1.
4. The first SPAR metadata parameter encoded within the bitstream is calculated from the combination of the DirAC metadata and the input covariance of the first set of channels. The method according to claim 3.
5. The second set of channels includes a primary downmix channel, which is obtained by applying gain to the first set of channels and summing the gain-adjusted first set of channels, the gain being calculated from the DirAC metadata, and the primary downmix channel is a representation of the dominant intrinsic signal for the first set of channels. The method according to claim 1.
6. A multichannel audio signal is received by at least one processor, which includes a first set of channels and a second set of channels different from the first set of channels. Regarding the first set of frequency bands, The at least one processor calculates directional audio coding (DirAC) metadata from the first set of channels, The at least one processor quantizes the DirAC metadata, The at least one processor encodes the quantized DirAC metadata, The at least one processor converts the quantized DirAC metadata into two or more parameters of the first spatial reconstruction (SPAR) metadata, Regarding a second set of frequency bands that are lower than the first set of frequency bands, The at least one processor calculates a second SPAR metadata from the first set of channels and the second set of channels, The at least one processor quantizes the second SPAR metadata, The at least one processor encodes the quantized second SPAR metadata, The at least one processor generates a downmix based on the first SPAR metadata and the second SPAR metadata, The at least one processor calculates frequency coefficients from the first set of channels and the second set of channels, The at least one processor downmixes the coefficients and downmixes into a third set of channels, The at least one processor encodes the third set of channels, Outputting a bitstream including the encoded third set of channels, the quantized and encoded second SPAR metadata, and the quantized and encoded DirAC metadata, Methods that include...
7. Two or more parameters in the first SPAR metadata are converted from the DirAC metadata, and the second SPAR data is calculated using the input covariance. The method according to claim 6.
8. One or more parameters in the first SPAR metadata for the first set of frequency bands are encoded in the bitstream rather than being converted from the DirAC metadata. The method according to claim 6.
9. The first SPAR metadata parameter encoded within the bitstream is calculated from the combination of the DirAC metadata and the covariance of the second set of channels. The method according to claim 8.
10. The first SPAR metadata parameter encoded within the bitstream is: The prediction coefficients, cross-prediction coefficients, and decorrelation coefficients for the second set of channels are included above. The method according to claim 8.
11. The first set of channels is a first-order ambisonic (FOA) channel, and the second set of channels includes at least one of planar or non-planar higher-order ambisonic (HOA) channels. The method according to claim 6.
12. The two or more parameters of the first SPAR metadata are converted from the DirAC metadata, and the second SPAR metadata is calculated and encoded for all frequency bands. The method according to claim 6.
13. The second SPAR metadata is calculated from the first and second sets of channels and the first SPAR metadata. The method according to claim 6.
14. Calculating a third SPAR metadata for the second set of channels and the first set of frequency bands, The prediction coefficients of the first set for the second set of channels in the third SPAR metadata are calculated from the first input covariance of the first set of channels and the second set of channels, Quantizing the first prediction coefficient in the third SPAR metadata, A first downmix is calculated from the quantized first prediction coefficients for the second set of channels and the first set of frequency bands, and the quantized DirAC metadata for the first set of channels and the first set of frequency bands. The first post-prediction is calculated using the first input covariance and the first downmix, From the first post-prediction, the cross-prediction coefficients of the first set in the third SPAR metadata are calculated, Quantizing the first cross-prediction coefficient in the third SPAR metadata, From the first input covariance, calculate the second set of prediction coefficients for the first set of channels and the first set of frequency bands, The second downmix is calculated from the first and second dequantized prediction coefficients for the channels of the first set and the channels of the second set and the frequency band of the first set, The second post-prediction is calculated using the first input covariance and the second downmix, The second set of cross-prediction coefficients is calculated from the second post-prediction, The first residual is calculated from the second cross-prediction coefficient and the second post-prediction, Calculating the decorrelation coefficient of the first set in the third SPAR metadata from the first residual and the first set of frequency bands, Quantizing the first uncorrelated coefficient in the third SPAR metadata, Encoding the first prediction coefficient, the first cross-prediction coefficient, and the first decorrelation coefficient in the third SPAR metadata, Outputting a bitstream including the encoded first prediction coefficient, the first cross-prediction coefficient, and the first decorrelation coefficient, The method according to claim 6, including the method described in claim 6.
15. The DirAC metadata is estimated based on the input covariance matrix. The method according to claim 4.
16. Generating the SPAR metadata from the DiRAC metadata is The second input covariance is approximated from the DirAC metadata and spherical harmonic response, Calculating the two or more parameters in the SPAR metadata from the second input covariance, including, The method according to claim 4.
17. One or more elements of the second input covariance are generated using the uncorrelated coefficients in the DirAC metadata and the second SPAR metadata. The method according to claim 16.
18. One or more elements of the second input covariance are generated from the DirAC metadata such that the uncorrelated coefficient in the SPAR metadata depends only on the at least one diffusivity parameter and the normalization of the ambisonics input and one or more constants in the DirAC metadata. The method according to claim 16.
19. The third set of channels includes a primary downmix channel, which is obtained by applying gain to the first set of channels and summing the gain-adjusted first set of channels, the gain being calculated from the DirAC metadata, and the primary downmix channel is a representation of the dominant intrinsic signal for the first set of channels. The method according to claim 6.
20. The at least one spreading parameter includes a spreading parameter calculated based on the reference power (E) and intensity (I) of the multi-channel audio signal, where E and I are calculated based on the input covariance. The method according to claim 4-5 or 7-19.
21. The first set of channels includes a primary ambisonic (FOA) channel, and the calculation of the reference power in the DirAC metadata ensures that the reference power is always greater than or equal to the variance of the W channel of the FOA channel. The method according to claim 20.
22. The downmix is energy compensated in the frequency band of the first set of channels based on the ratio of the total variance of the first set of channels to the total variance according to the second input covariance generated using the DirAC metadata. The method according to claim 16.
23. Receiving an encoded bitstream containing encoded audio channels and metadata by at least one processor, wherein the metadata includes first directional audio coding (DirAC) metadata associated with a first frequency band and first spatial reconstruction (SPAR) metadata associated with a second frequency band lower than the first frequency band, The at least one processor decrypts the first DirAC metadata and the first SPAR metadata, The at least one processor dequantizes the decoded first DirAC metadata and the first SPAR metadata, Regarding the first frequency band, The at least one processor converts the inversely quantized first DirAC metadata into two or more parameters of the second SPAR metadata, The at least one processor mixes the first and second SPAR metadata to form composite SPAR metadata, The at least one processor decodes the encoded audio channel, The at least one processor reconstructs the downmix channel from the decoded audio channel, The at least one processor converts the downmix channel into a frequency banded domain, The at least one processor generates a SPAR upmix based on the composite SPAR metadata, The at least one processor upmixes the downmix channels in the frequency banding region to a first set of channels based on the SPAR upmix, The at least one processor estimates the second DirAC metadata in the second frequency band from the first set of channels and zero or more parameters in the first SPAR metadata, The at least one processor upmixes the first set of channels to the second set of channels in the frequency banding region based on the first and second DirAC metadata, The at least one processor converts the second set of channels from the frequency-banded domain to the time domain, Methods that include...
24. The aforementioned downmix is converted to the frequency-banded domain using a filter bank (complex low-latency filter bank). The method according to claim 23.
25. The first set of channels includes a first-order ambisonics (FOA) channel and zero or more higher-order ambisonics (HOA) channels. The method according to claim 23.
26. The HOA channel among the first set of channels includes at least one planar HOA channel or a non-planar HOA channel. The method according to claim 23.
27. The bitstream includes a third SPAR metadata corresponding to the HOA channel and the first frequency band of the first set of channels, The method according to claim 23.
28. The DirAC metadata is estimated from the first-order ambisonics (FOA) channel in the frequency band domain for a third set of frequency bands, including the first set of frequency bands and the second set of frequency bands. The method according to claim 23.
29. The DirAC metadata is estimated for a fourth set of frequency bands, which is a subset of the second set of frequency bands, from the SPAR metadata and zero or more elements of the covariance generated using the downmix and upmix in the fourth set of frequency bands. The method according to claim 23.
30. Calculating the DirAC metadata from the SPAR metadata for the fourth set of frequency bands is: Calculating the direction of arrival angle in DirAC metadata solely from the prediction coefficients in SPAR metadata, The calculation of at least one diffusivity parameter in the DirAC metadata from the prediction coefficients and zero or more uncorrelated coefficients in the SPAR metadata, as well as the scale factor, including, The method according to claim 29.
31. The encoded channel includes a first-order ambisonic channel, and upmixing the downmix channel to a first set of channels in the first frequency band is: Calculating the upmix scaling gain from the first DirAC metadata, This includes applying the upmix scaling gain to the primary downmix channel to obtain the W channel of the first set of channels in the first frequency band, wherein the primary downmix channel is a representation of the dominant intrinsic signal for the first set of channels. The method according to any one of claims 23 to 26.
32. A non-temporary computer-readable storage medium that stores instructions causing a computing device to perform the method according to any one of claims 1 to 31 when the computing device is executed by the computing device.
33. At least one processor, A memory that stores instructions causing the computing device to perform the method according to any one of claims 1 to 31 when executed by at least one processor, A computing device equipped with the following features.