Methods, apparatus and systems for scene based audio mono decoding
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY LABORATORIES LICENSING CORP
- Filing Date
- 2024-07-03
- Publication Date
- 2026-05-13
AI Technical Summary
Current spatial audio coding technologies, such as Spatial Reconstruction (SPAR) and Directional Audio Coding (DirAC), have limitations in efficiently encoding and decoding audio scenes, particularly when transitioning from Ambisonics to Mono format, leading to undesirable audio characteristics due to inability to signal format changes effectively.
The development of a system that includes a mono detector in the encoder to signal mono mode through the bitstream, ensuring proper behavior in Scene Based Audio (SBA) mode, using implicit signaling by setting specific metadata parameters to zero, and explicit signaling with a mono flag, allowing for accurate detection and preservation of mono content within an Ambisonics format.
This solution ensures proper audio behavior when handling mono inputs in SBA mode, preventing misclassification and maintaining audio quality by accurately identifying and processing mono content within Ambisonics formats, thus enhancing the efficiency and effectiveness of audio encoding and decoding processes.
Smart Images

Figure US2024036799_09012025_PF_FP_ABST
Abstract
Description
METHODS, APPARATUS AND SYSTEMS FOR SCENE BASED AUDIO MONO DECODING CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from U.S. Provisional Patent Application Serial No. 63 / 570,117, filed on March 26, 2024, U.S. Provisional Patent Application Serial No. 63 / 593,278, filed on October 26, 2023 and U.S. Provisional Patent Application Serial No. 63 / 511,786, filed on July 3, 2023, each of which is incorporated by reference herein in its entirety. TECHNICAL FIELD
[0002] This disclosure relates generally to audio processing. BACKGROUND
[0003] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application, and are not admitted as prior art by inclusion in this section.
[0004] Spatial Reconstruction (SPAR) and Directional Audio Coding (DirAC) are separate spatial audio coding technologies that each seek to represent an input spatial audio scene in a compact way to enable transmission with a good trade-off between audio quality and bitrate. One such input format for a spatial audio scene is an Ambisonics representation (e.g., first-order Ambisonics (FOA) or higher-order Ambisonics (HOA)).
[0005] SPAR seeks to maximize perceived audio quality while minimizing bitrate by reducing the energy of the transmitted audio data while still allowing the second-order statistics of the Ambisonics audio scene (e.g., the covariance) to be reconstructed at the decoder side using transmitted metadata. SPAR seeks to faithfully reconstruct the input Ambisonics scene at the output of the decoder.
[0006] DirAC is a technology which represents spatial audio scenes as a collection of directions of arrival (DOA) in time-frequency tiles. From this representation, a similar-sounding scene can be reproduced in a different output format (e.g., binaural). Notably, in the context of Ambisonics, the DirAC representation allows a decoder to produce higher-order output from low-order input (blind upmix). DirAC seeks to preserve direction and diffuseness of the dominant sounds in the input scene.
[0007] Both DirAC and SPAR have different strengths and properties. It is therefore desirable to combine the complementary aspects of DirAC and SPAR (e.g., higher audio quality, reducedbitrate, input / output format flexibility and / or reduced computational complexity) into a coder / decoder (“codec”), such as an Ambisonics codec. SUMMARY
[0008] Techniques are described for processing audio signals. Examples found herein provide for systems, devices, and methods to encode a bitstream and / or decode a bitstream, where frames are marked with a mode indicator that is leveraged in rendering the audio.
[0009] Audio signal encoding and decoding methods are disclosed herein. Some disclosed methods for encoding an audio signal involve obtaining an audio signal that represents an input audio scene with a primary channel and side channels, analyzing the power of the primary channel and analyzing the powers of the side channels. Some such methods involve detecting a mono mode for encoding the audio signal based on analyzing the power of the primary channel and the powers of the side channels and computing one or more downmix channels and spatial metadata from the audio signal for a detected mono mode. Some such methods involve encoding the one or more downmix channels and spatial metadata in a bitstream for the detected mono mode and indicating the mono mode in the bitstream.
[0010] Some example embodiments describe methods for encoding an audio signal. In some instances, the audio signal may represent an input audio scene with a primary channel and side channels. In some example embodiments, the methods may involve obtaining the audio signal. According to some example embodiments, the methods may involve analyzing a power of the primary channel of the audio signal and analyzing powers of the side channels of the audio signal. In some example embodiments, the methods may involve detecting a mono mode for encoding the audio signal based on analyzing the power of the primary channel and the powers of the side channels. According to some example embodiments, the methods may involve computing one or more downmix channels and spatial metadata from the audio signal for a detected mono mode and encoding the one or more downmix channels and spatial metadata in a bitstream for the detected mono mode. In some example embodiments, the methods may involve indicating the mono mode in the bitstream.
[0011] In some example embodiments, the methods may involve outputting the bitstream, storing the bitstream, transmitting the bitstream, or combinations thereof.
[0012] According to some example embodiments, detecting the mono mode may be based on a determination that input audio signal has non-silent audio in the primary channel and that the side channels are silent.
[0013] In some example embodiments, detecting the mono mode may involve computing a primary channel power of the primary channel of the audio signal and computing a sum power asa summation of the powers of the side channels of the audio signal. In some such example embodiments, detecting the mono mode also may involve evaluating a ratio of the primary channel power to the sum power. In some such example embodiments, detecting the mono mode also may involve determining when the ratio exceeds a threshold.
[0014] According to some example embodiments, the methods may involve implicitly signaling the mono mode by setting one or more parameters of the spatial metadata to a value of zero. In some such example embodiments, the one or more parameters of the spatial metadata may include one or more Spatial Reconstruction (SPAR) metadata parameters, one or more Directional Audio Coding (DirAC) metadata parameters, or both.
[0015] In some example embodiments, the methods also may involve explicitly signaling the mono mode by setting a mono flag.
[0016] According to some example embodiments, the bitstream may be an Immersive Voice and Audio (IVAS) encoded bitstream.
[0017] Some additional example embodiments describe methods for audio signal decoding. In some example embodiments, the methods may involve obtaining an encoded bitstream and decoding the encoded bitstream to obtain downmix channels, spatial metadata, and a mono mode indicator. According to some example embodiments, the methods may involve setting one or more parameters of spatial metadata to zero upon detecting a mono mode. In some example embodiments, the methods may involve upmixing the downmix channels using the spatial metadata. According to some example embodiments, the methods may involve rendering upmixed channels to a desired audio format.
[0018] In some example embodiments, the rendering produces rendered audio data. According to some example embodiments, the methods also may involve transmitting the rendered audio data, storing the rendered audio data, transmitting the rendered audio data, or combinations thereof. Some example embodiments may involve providing the rendered audio data to one or more loudspeakers for playback.
[0019] According to some example embodiments, the mono mode indicator may be based on values of one or more spatial metadata parameters in the encoded bitstream that are set to values to indicate the mono mode. According to some such example embodiments, the one or more spatial metadata parameters may include one or more SPAR metadata parameters, one or more DirAC metadata parameters, or both.
[0020] In some example embodiments, the methods may involve setting the one or more parameters of spatial metadata to zero upon detecting the mono mode involves setting one or more energy ratio values to zero. In some such example embodiments, a received energy ratio metadatavalue may be non-zero due to an artifact of a quantization process. In some example embodiments, the methods also may involve setting one or more diffuseness values to one.
[0021] According to some example embodiments, the encoded bitstream may be an IVAS encoded bitstream.
[0022] Still other example embodiments describe a method of encoding an audio signal representing an input audio scene with a primary channel and side channels. In some example embodiments, the methods may involve obtaining a frame of the audio signal from a bitstream and identifying a plurality of frequency bands associated with the frame. According to some example embodiments, the methods may involve determining power associated with the primary channel and the side channels for each of the plurality of frequency bands associated with the frame. In some example embodiments, the methods may involve classifying each of the plurality of frequency bands as one of Silence, Mono, or Ambisonics based on a determined power for a corresponding frequency band. According to some example embodiments, the methods may involve marking a mode of the frame as one of Silence, Mono, or Ambisonics based on a plurality of classified frequency bands for the frame and encoding a marked mode of the frame in the bitstream.
[0023] In some example embodiments, the methods also may involve outputting the bitstream, storing the bitstream, transmitting the bitstream, or combinations thereof.
[0024] According to some example embodiments, marking the frame may involve marking the frame as Ambisonics when one or more classified frequency bands corresponds to Ambisonics.
[0025] According to some example embodiments, the methods may involve marking the frame as Mono when no classified frequency band corresponds to Ambisonics, and when one or more classified frequency bands corresponds to Mono. In some example embodiments, the methods may involve marking the frame as Silence when no classified frequency band corresponds to Ambisonics, and when no classified frequency band corresponds to Mono.
[0026] In some example embodiments, determining power of a frequency band also may involve calculating a primary channel power of the frequency band and calculating a sum power of the frequency band, wherein the sum power is a summation of power to the side channels. In some such example embodiments, classifying each of the plurality of frequency bands also may involve evaluating a ratio of the primary channel power to the sum power. In some such example embodiments, detecting a mono mode may involve determining when the ratio exceeds a threshold. In some such example embodiments, the threshold may be a linear function of the primary channel power, the primary channel power may be normalized to a value between 0 and 1, and the threshold may be clipped between maximum and minimum values.
[0027] According to some example embodiments, classifying each of the plurality of frequency bands also may involve determining that the sum power is below a minimum noise level and declaring the frequency band as Mono when the primary channel power is above a noise threshold.
[0028] In some example embodiments, classifying each of the plurality of frequency bands also may involve determining that the sum power is above a minimum noise level and declaring the frequency band as Mono when a normalized primary channel power is above a noise threshold.
[0029] According to some example embodiments, the bitstream may be an IVAS encoded bitstream.
[0030] In some example embodiments, marking the mode of the frame may involve setting one or more SPAR metadata parameters to zero, setting one or more DirAC metadata parameters to zero, or both.
[0031] Still other example embodiments describe an apparatus. In some example embodiments, the apparatus may include an input / output (I / O) system and one or more processors. According to some example embodiments, the one or more processors may be configured to obtain, via the I / O system, an audio signal representing an input audio scene with a primary channel and side channels. In some example embodiments, the one or more processors may be configured to analyze a power of the primary channel of the audio signal and to analyze powers of the side channels of the audio signal. According to some example embodiments, the one or more processors may be configured to detect a mono mode for encoding the audio signal based on analyzing the power of the primary channel and the powers of the side channels. In some example embodiments, the one or more processors may be configured to compute one or more downmix channels and spatial metadata from the audio signal for a detected mono mode. According to some example embodiments, the one or more processors may be configured to encode the one or more downmix channels and spatial metadata in a bitstream for the detected mono mode, and to indicate the mono mode in the bitstream.
[0032] According to some example embodiments, the one or more processors may be configured for outputting the bitstream, storing the bitstream, transmitting the bitstream, or combinations thereof.
[0033] In some example embodiments, marking the frame also may involve marking the frame as Ambisonics when one or more classified frequency bands corresponds to Ambisonics.
[0034] According to some example embodiments, marking the frame also may involve marking the frame as Mono when no classified frequency band corresponds to Ambisonics, and when one or more classified frequency bands corresponds to Mono.
[0035] Still other example embodiments describe an apparatus that includes an I / O system and one or more processors. In some example embodiments, the one or more processors may beconfigured to obtain, via the I / O system, a frame of an audio signal from a bitstream. According to some example embodiments, the one or more processors may be configured to identify a plurality of frequency bands associated with the frame. In some example embodiments, the one or more processors may be configured to determine power associated with the primary channel and the side channels for each of the plurality of frequency bands associated with the frame. According to some example embodiments, the one or more processors may be configured to classify each of the plurality of frequency bands as one of Silence, Mono, or Ambisonics based on a determined power for a corresponding frequency band. In some example embodiments, the one or more processors may be configured to mark a mode of the frame as one of Silence, Mono, or Ambisonics based on a plurality of classified frequency bands for the frame and to encode a marked mode of the frame in the bitstream.
[0036] According to some example embodiments, the one or more processors also may be configured for outputting the bitstream, storing the bitstream, transmitting the bitstream, or combinations thereof.
[0037] In some example embodiments, marking the frame also may involve marking the frame as Ambisonics when one or more classified frequency bands corresponds to Ambisonics. In some example embodiments, marking the frame also may involve marking the frame as Mono when no classified frequency band corresponds to Ambisonics, and when one or more classified frequency bands corresponds to Mono.
[0038] Still other example embodiments describe an apparatus that includes an I / O system and one or more processors. In some example embodiments, the one or more processors may be configured to obtain, via the I / O system, an encoded bitstream. According to some example embodiments, the one or more processors may be configured to decode the encoded bitstream to obtain downmix channels, spatial metadata, and a mono mode indicator. In some example embodiments, the one or more processors may be configured to set one or more parameters of spatial metadata to zero upon detecting a mono mode. According to some example embodiments, the one or more processors may be configured to upmix the downmix channels using the spatial metadata. In some example embodiments, the one or more processors may be configured to render upmixed channels to a desired audio format.
[0039] According to some example embodiments, the rendering may produce rendered audio data. According to some such example embodiments, the one or more processors may be further configured for transmitting the rendered audio data, storing the rendered audio data, transmitting the rendered audio data, or combinations thereof. In some example embodiments, the one or more processors also may be configured for providing the rendered audio data to one or more loudspeakers for playback.
[0040] The embodiments described herein may be generally described as techniques, where the term “technique” may refer to system(s), device(s), method(s), computer-readable instruction(s), module(s), component(s), hardware logic, and / or operation(s) as suggested by the context as applied herein.
[0041] Features and technical benefits other than those explicitly described above will be apparent from a reading of the following Detailed Description and a review of the associate drawings. This Summary is provided to introduce a selection of techniques in a simplified form, and not intended to identify key or essential features of the claimed subject matter, which are defined by the appended claims. DESCRIPTION OF DRAWINGS
[0042] FIG.1 is a block diagram of an example IVAS codec framework;
[0043] FIG.2A is a flow diagram that outlines various example methods that may be performed by what may be referred to herein as an “encoder mono detector”;
[0044] FIG.2B is a flow diagram that outlines additional example methods that may be performed by an encoder mono detector;
[0045] FIG. 3 is a flow diagram that outlinessome examples of band classification processes, which include an iteration loop over frequency bands for a given frame;
[0046] FIG.4 is a flow diagram that outlines some examples of frame marking processes, which include an iteration loop over frequency bands for a given frame;
[0047] FIG. 5 is a flow diagram that outlines some additional example methods that may be performed by a decoder such as those disclosed herein;
[0048] FIG.6 is a block diagram that illustrates some examples of an IVAS codec system;
[0049] FIG.7 is a block diagram of an example hardware architecture suitable for implementing the systems, devices and methods described herein;
[0050] FIG. 8 is a block diagram of an encoder with frequency-based and channel-based split between SPAR and DirAC, according to one or more embodiments;
[0051] FIG. 9 is a block diagram of decoder with frequency-based and channel-based split between SPAR and DirAC, according to one or more embodiments;
[0052] FIG. 10 is a block diagram of an alternate encoder with frequency-based and channel- based split between SPAR and DirAC, according to one or more embodiments; and
[0053] FIG. 11 is a block diagram of an alternate decoder with frequency-based and channel- based split between SPAR and DirAC, all arranged in accordance with embodiments described herein.
[0054] In the drawings, specific arrangements or orderings of schematic elements, such as those representing devices, units, instruction blocks and data elements, are shown for ease of description. However, it should be understood by those skilled in the art that the specific ordering or arrangement of the schematic elements in the drawings is not meant to imply that a particular order or sequence of processing, or separation of processes, is required. Further, the inclusion of a schematic element in a drawing is not meant to imply that such element is required in all embodiments or that the features represented by such element may not be included in or combined with other elements in some implementations.
[0055] Further, in the drawings, where connecting elements, such as solid or dashed lines or arrows, are used to illustrate a connection, relationship, or association between or among two or more other schematic elements, the absence of any such connecting elements is not meant to imply that no connection, relationship, or association can exist. In other words, some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the disclosure. In addition, for ease of illustration, a single connecting element is used to represent multiple connections, relationships or associations between elements. For example, where a connecting element represents a communication of signals, data, or instructions, it should be understood by those skilled in the art that such element represents one or multiple signal paths, as may be needed, to affect the communication.
[0056] The same reference symbol used in various drawings indicates like elements. DETAILED DESCRIPTION
[0057] In the following detailed description, numerous specific details are set forth to provide a thorough understanding of various described embodiments with reference to the accompanying drawings. The illustrative embodments in the detailed description, drawings, and claims are not neant to be limiting. Other embodiments may be utilized, and other changes made, without departing from the spirit or scope of the present disclosure. In light of the present disclosure, it will be apparent to one of ordinary skill in the art that the various described features and implementations may be practiced without many of these specific details. In some instances, well- known methods, procedures, components, and circuits, have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features are described hereafter that can each be used independently of one another or with any combination of other features. Thus, the features may be arranged, substituted, combined, separated or designed into other configurations, which is contemplated in light of the present disclosure.Nomenclature
[0058] As used herein, the term “includes” and its variants are to be read as open-ended terms that mean “includes, but is not limited to.” The term “or” is to be read as “and / or” unless the context clearly indicates otherwise. The term “based on” is to be read as “based at least in part on.” The term “one example implementation” and “an example implementation” are to be read as “at least one example implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “determined,” “determines,” or “determining” are to be read as obtaining, receiving, computing, calculating, estimating, predicting or deriving. In addition, in the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs. Acronyms IVAS – Immersive Voice and Audio Services CLDFB – Complex Low Delay Filter Bank SBA – Scene Based Audio ESD – Equivalent Spatial Domain HOA – Higher Order Ambisonics SN3D – Schmidt Semi-Normalization FOA – First Order Ambisonic COV – Covariance ACN- Ambisonic Channel Number DMX – Downmix DOA – Direction of Arrival AR – Augmented Reality FFT – Fast Fourier Transform VR –Virtual Reality DFT – Doscrete Fourier Transform CPU – Central Processing Unit MDFT – Modified Discrete Fourier Transform DSP – Digital Signal Processor MD – Metadata ASIC – Application-Specific Integrated Circuit BS - Bistream FPGA – Field-Programmable Gate Array EVS – Enhanced Voice Services RAM – Random Access Memory SPAR – Spatial Reconstruction ROM – Read Only Memory DIRAC – Directional Audio Coding EPROM – Erasable Programmable ROM D2S – DIRAC to SPAR CD-ROM – Compact Disc Read-Only Memory S2D – SPAR to DIRAC I / O – Input / Output LBRSBA – Low Bitrate SBA Example IVAS Codec Framework
[0059] FIG. 1 is a block diagram of an example immersive voice and audio services (IVAS) coder / decoder (“codec”) framework 100 for encoding and decoding IVAS bitstreams, according to one or more embodiments. IVAS is expected to support a range of audio service capabilities, including but not limited to mono to stereo upmixing and fully immersive audio encoding,decoding and rendering. IVAS is also intended to be supported by a wide range of devices, endpoints, and network nodes, including but not limited to: mobile and smart phones, electronic tablets, personal computers, conference phones, conference rooms, virtual reality (VR) and augmented reality (AR) devices, home theatre devices, and other suitable devices.
[0060] IVAS codec 100 includes IVAS encoder 101 and IVAS decoder 104. IVAS encoder 101 includes spatial encoder 102 that receives N channels of input spatial audio (e.g., FOA, HOA). In some implementations, spatial encoder 102 implements SPAR and DirAC for analyzing / downmixing N_dmx spatial audio channels, as described in further detail below. The output of spatial encoder 102 includes a spatial metadata (MD) bitstream (BS) and N_dmx channels of spatial downmix. The spatial MD is quantized and entropy coded. In some implementations, quantization can include fine, moderate, coarse and extra coarse quantization strategies and entropy coding can include Huffman or Arithmetic coding. In some implementations, the framework may permit not more than 3 levels of quantization at a given operating mode; however, with decreasing bitrates, in some such implementations the three levels become increasingly coarser overall, to meet bitrate requirements. Core audio encoder 103 (e.g., based on a mono Enhanced Voice Services (EVS) encoding unit) encodes N_dmx channels (N_dmx = 1-16 channels) of the spatial downmix into an audio bitstream, which is combined with the spatial MD bitstream into an IVAS encoded bitstream transmitted to IVAS decoder 104. As described below, given bitrate constraints for low bit rate Scene Based Audio (SBA), in some implementations the number of channels will be limited to a single channel.
[0061] IVAS decoder 104 includes core audio decoder 105 (e.g., EVS decoder) that decodes the audio bitstream extracted from the IVAS bitstream to recover the N_dmx audio channels. Spatial decoder / renderer 106 (e.g., SPAR / DirAC) decodes the spatial MD bitstream extracted from the IVAS bitstream to recover the spatial MD, and synthesizes / renders output audio channels using the spatial MD and a spatial upmix for playback on various audio systems with different speaker configurations and capabilities. Low Bitrate SBA (LBRSBA)
[0062] In some embodiments, it is desirable to implement LBRSBA (e.g., Ambisonics) using a SPAR-DIRAC codec. LBRSBA can be achieved using one or more of the following techniques: 1) reduced MD bitrate and band interleaving; 2) tuning of Active W; 3) extra covariance smoothing to facilitate the reduced MD bitrate; and 4) decoder side decorrelator coefficient smoothing.
[0063] Background information for the above techniques can be found in one or more of the following documents, all of which are hereby incorporated by reference: • PCT Application No.2023 / 063769, for “DirAC-SPAR Audio Processing”;• US 9,978,385 for “Parametric reconstruction of audio signals”; • International Application No. WO20212252748A1, for “Encoding of multi-channel audio signals comprising downmixing of a primary and two or more scaled non- primary input channels”; • International Application No. WO2022120093A1, for “Immersive Voice and Audio Services (IVAS) with Adaptive Downmix Strategies”; • International Application No. WO2021252811A2, for “Quantization and Entropy Coding of Parameters for a Low Latency Immersive Audio Codec”; • US Patent Publication No.20220406318A1, for “Bitrate Distribution in Immersive Voice and Audio Services”; and • US Patent Publication No.2022 / 0277757 for “Systems and Methods for Covariance Smoothing.”
[0064] LBRSBA can also be achieved using other techniques known to those of skill in the relevant arts.
[0065] When operating the IVAS Codec, there may be times when audio content has changed from Ambisonics format to Mono format, but the device that is streaming and encoding audio content cannot signal a format change to the codec. When this occurs, the decoder can treat the audio content as Ambisonics format, which may result in undesireable audio characteristics. Mono can be defined as a format with a single channel of audio content, whereas Ambisonics can be defined as a multi-channel audio format with 4 or more channels.
[0066] In this case, the single Mono channel is sent as the W channel of an Ambisonics input, with silence or near silence in the other channels. An example of near silence is when dither is applied to reduce quantisation error on silent channels. In some implementations, the IVAS codec SBA mode cannot handle this use case due to certain design choices that have already been made. SBA mode is the term for encoding and decoding Ambisonics content in the IVAS codec.
[0067] Below is a proposal of a system that will ensure proper behaviour of the IVAS Codec with mono input whilst operating in SBA mode. The system includes a mono detector in the encoder, is configured for signalling of mono through the bitstream (possibly without using explicit signalling bits) and includes a mono detector / preserver to ensure the W channel is treated as mono in the decoder. In some examples, the system works on a frame by frame basis, but has some inter- frame dependencies. Encoder / Encoding Mono Detection
[0068] Some example processes for detecting mono in the encoder are described below. For convenience, these example processes can be broadly considered as being two step processes: stepone being frequency band classification and step two being frame marking (or declaring). These two processes may be combined into a single process, or separated into additional processes or sub-processes, as may be desired in a particular implementation. Moreover, the functions described herein to implement the overall process may be substituted with other functions, where some of the described steps may be eliminated and / or replaced with other more suitable implemenations without departing from the spirit of the present disclosure.
[0069] FIG. 2A is a flow diagram that outlines various example methods 200 that may be performed by what may be referred to herein as an “encoder mono detector.” The example methods 200 may be partitioned into blocks, such as blocks 205, 210, 215, 220, 225, 230, and 235. The various blocks may be described as operations, processes, methods, steps, acts or functions. The blocks of methods 200, like other methods described herein, are not necessarily performed in the order indicated. In some implementations, one or more of the blocks of methods 200 may be performed concurrently. Moreover, some implementations of methods 200 may include more or fewer blocks than shown and / or described. The blocks of methods 200 involve encoding an audio signal, which may be performed by one or more devices, for example, the device that is shown in FIG.7. Processing may commence at block 205.
[0070] Block 205 involves “obtaining an audio signal.” In various examples, the audio signal represents an input audio scene with a primary channel and side channels. In some examples, the primary channel may be the W channel of an Ambisonics format, such as a first-order Ambisonics format, and the side channels may be, or may include, the X, Y and Z channels. In other examples, the primary channel may be the center channel of a multi-channel format and the side channels may include the left and right channels of the multi-channel format. In yet other examples, the primary channel may be the Mid channel of audio captured using a Mid-Sides microphone technique and the Sides channel could correspond with the side channels. Processing may continue to block 210.
[0071] Block 210 involves “analyzing power of the primary channel of the audio signal”. Block 215 involves “analyzing powers of the side channels of the audio signal.” In some examples, the analyses of blocks 210 and 215 may occur in the frequency domain. Various examples of frequency-domain analyses are disclosed herein. In some alternative examples, the analyses of blocks 210 and 215 may occur in the time domain. Processing may continue to block 220.
[0072] Block 220 involves “detecting a mono mode for encoding the audio signal based on analyzing the power of the primary channel and the powers of the side channels.” According to some examples, detecting the mono mode may be based on a determination that the input audio signal has non-silent audio in the primary channel and that the side channels are silent. Processing may continue to block 225.
[0073] In some examples, detecting the mono mode may involve computing a primary channel power of the primary channel of the audio signal and computing a sum power as a summation of the powers of the side channels of the audio signal. In some such examples, detecting the mono mode may involve evaluating a ratio of the primary channel power to the sum power. According to some such examples, detecting the mono mode may involve determining when the ratio exceeds a threshold.
[0074] Block 225 involves “computing one or more downmix channels and spatial metadata from the audio signal for a detected mono mode.” In some examples, the spatial metadata may include Spatial Reconstruction (SPAR) spatial metadata, Directional Audio Coding (DirAC) spatial metadata, or both. Processing may continue to block 230.
[0075] Block 230 involves “encoding the one or more downmix channels and spatial metadata in a bitstream for the detected mono mode.” In some examples, the bitstream may be an Immersive Voice and Audio (IVAS) encoded bitstream. Block 235 involves”indicating the mono mode in the bitstream.”
[0076] In some examples, methods 200 may involve implicitly signaling the mono mode according to one or more parameters of the spatial metadata. In some such examples, method 200 may involve implicitly signaling the mono mode by setting one or more parameters of the spatial metadata to a value of zero.
[0077] According to some examples, the encoder mono detector is configured to determine when an Ambisonics input has content in the W channel only. In some such example, for each audio block, the encoder mono detector analyses the content in all the channels and makes a decision about whether the block of audio can be declared as mono or Ambisonics.
[0078] In some examples, the encoder mono detector works in the Modified Discrete Fourier Transform (MDFT) domain by looping over each frequency band, calculating the power of the W channel (Wp) and the sum of the power of the other channels (Sump) in the input for each band. According to some examples, the encoder mono detector uses the Wp power, the sum of other channels’ power Sump and the ratio (Wp / Sump) of Wp power to the sum of other channels power Sump to determine if a band is mono, SBA (Ambisonics) or silence. In some examples, the classification decision (mono, multi-channel or silence) is based on two thresholds: a static NOISE_THRESHOLD and a dynamic RATIO_THRESHOLD.
[0079] The NOISE_THRESHOLD may be a static threshold that can be used to determine if a band is silent or has content. For example, a band that is below the NOISE_THRESHOLD may be considered as silent or near silent; while a band that is above the NOISE_THRESHOLD may be considered to include content. The RATIO_THRESHOLD may be a dynamic threshold that scales with content as the content gets quieter, and can be used to classify a band as eitherAmbisonics or mono. For example, a band that is below the RATIO_THRESHOLD may be considered as Ambisonics; while a band that is above the RATIO_THRESHOLD may be considered as mono.
[0080] The static threshold (e.g., NOISE_THRESHOLD) can be determined by running silent and near silent content through a codec and analysing the levels that are produced by the MDFT. According to some examples, the static threshold may be in the range from 40 dB to 45 dB, in the range from 42 dB to 48 dB, etc. For example, in some instances the static threshold may be 44 dB or 45 dB. The dynamic threshold (e.g., RATIO_THRESHOLD) can be calculated using a linear function with the W channel power (Wp) as input, where the power Wp can be clipped at a maximum and a minimum threshold (e.g., MAX_THRESHOLD, MIN_THRESHOLD) to improve stability. In some examples, the minimum threshold may be in the range from 15 dB to 22 dB, in the range from 18 dB to 25 dB, etc. For example, in some instances the minimum threshold may be 20 dB. According to some examples, the maximum threshold may be in the range from 55 dB to 70 dB, in the range from 50 dB to 65 dB, etc. For example, in some instances the maximum threshold may be 60 dB. In some examples, the minimum and maximum thresholds are applied to normalized power values. All of the threshold values may be dependent on scaling within the IVAS codec.
[0081] In various experiments, the max and min thresholds were empirically determined by analysing mono and normal ambisonic content and ensuring that the discrimination level met the requirements of accurate detection. In some examples, the requirements of accurate detection may be that the leakage to non-W channels after decoding should be inaudible or nearly inaudible. Alternatively, or additionally, the requirements of accurate detection may be that there is no misclassification of multi-channel audio as mono. However, the disclosure is not limited to this specific method of determining max and min thresholds, and the thresholds may be adjusted accordingly.
[0082] According to some examples, after determining the state of all the frequency bands, the encoder mono detector is configured to check if any of the frequency bands was classified (or declared) as Ambisonics. If any frequency band is classified (or declared) as Ambisonics, then the encoder mono detector may determine that the entire audio frame can be declared (or marked) as non-mono. If none of the frequency bands are classified (or declared) as Ambisonics, the encoder mono detector may determine that all of the frequency bands must be considered (or declared) as either silent or mono. If there is at least one frequency band that is considered (or declared) as mono, then the encoder mono detector may determine that the frame will be classified (or declared) as mono.
[0083] The frame declaration as Ambisonics or Mono can be implemented with a mono flag, in some examples. In some implementations, there is a further process that can be used to ensure that the mono flag does not change states too quickly (e.g., toggle on and off rapidly), which would cause an instability. In some examples, the further process may involve analyzing a history of mono flags. In some such examples, only after there have been consecutive mono flags for some time period (e.g., 20 milliseconds(ms), 25 ms, 30 ms, 35 ms, 40 ms, etc.) does the encoder mono detector declare a frame as mono and signal this mono declaration to the decoder.
[0084] FIG.2B is a flow diagram that outlines additional example methods that may be performed by an encoder mono detector. The example methods 240 may be partitioned into blocks, such as blocks 245, 250, 255, 260, 265 and 270. The various blocks may be described as operations, processes, methods, steps, acts or functions. The blocks of methods 240, like other methods described herein, are not necessarily performed in the order indicated. In some implementation, one or more of the blocks of method 240 may be performed concurrently. Moreover, some implementations of method 240 may include more or fewer blocks than shown and / or described. The blocks of method 240 involve encoding an audio signal representing an input audio scene with a primary channel and side channels. The primary channel and side channels may, for example, be any of the types that are described with reference to FIG.2A. The blocks of method 240 may be performed by one or more devices, such as the device that is shown in FIG.7. Processing may commence at block 245.
[0085] Block 245 involves “obtaining a frame of an audio signal from a bitstream.” Processing may continue to block 250.
[0086] Block 250 involves “identifying a plurality of frequency bands associated with the frame.” Processing may continue to block 255. Block 255 involves “determining power associated with the primary channel and the side channels for each of the plurality of frequency bands associated with the frame.” In some examples, determining the power of a frequency band may involve calculating a primary channel power of the frequency band and calculating a sum power of the frequency band. The sum power may be a summation of the power to the side channels. Processing may continue to block 260.
[0087] Block 260 involves “classifying each of the plurality of frequency bands as one of Silence, Mono, or Ambisonics based on a determined power for a corresponding frequency band.” According to some examples, classifying each of the plurality of frequency bands may involve evaluating a ratio of the primary channel power to the sum power. In some such examples, detecting a mono mode may involve determining when the ratio exceeds a threshold. In some such examples, the threshold may be a linear function of the primary channel power. According to someexamples, the primary channel power may be normalized to a value between 0 and 1. In some examples, the threshold may be clipped between maximum and minimum values.
[0088] In some examples, classifying each of the plurality of frequency bands may involve determining that the sum power is below a minimum noise level and declaring the frequency band as Mono when the primary channel power is above a noise threshold. According to some examples, classifying each of the plurality of frequency bands may involve determining that the sum power is above a minimum noise level and declaring the frequency band as Mono when a normalized primary channel power is above a noise threshold. Processing may continue to block 265.
[0089] Block 265 involves “marking a mode of the frame as one of Silence, Mono, or Ambisonics based on a plurality of classified frequency bands for the frame.” In some examples, marking the frame may involve marking the frame as Ambisonics when one or more classified frequency bands corresponds to Ambisonics. According to some examples, marking the frame may involve marking the frame as Mono when no classified frequency band corresponds to Ambisonics, and when one or more classified frequency bands corresponds to Mono. In some examples, marking the frame may involve marking the frame as Silence when no classified frequency band corresponds to Ambisonics, and when no classified frequency band corresponds to Mono. According to some examples, marking the frame may involve setting one or more SPAR metadata parameters to zero, setting one or more DirAC metadata parameters to zero, or both. Processing may continue to block 270.
[0090] Block 270 involves “encoding a marked mode of the frame in the bitstream.” According to some examples, the bitstream may be an Immersive Voice and Audio (IVAS) encoded bitstream. In some examples, marking the mode of the frame may involve setting one or more Spatial Reconstruction (SPAR) metadata parameters to zero, setting one or more Directional Audio Coding (DirAC) metadata parameters to zero, or both.
[0091] Another example mono detection process for the encoder may involve the following: • For each frame of the encoder loop over each MDFT frequency band. • For each frequency band calculate: o W channel power (W_p); and o Sum of the power of each individual non-W channel. • Check if the sum power is close to zero to prevent divide by zero errors. o If the sum power is close to zero, check if the W power is greater than the NOISE_THRESHOLD and declare the band as mono if it is. o Otherwise start the loop iteration again.• Calculate the ratio of W power to sum power. • Calculate the RATIO_THRESHOLD: o Normalise W power to be between 0 and 1; o Scale MAX_TRESHOLD by norm W power; and o Clip the calculated threshold to ensure the range is within MIN_THRESHOLD and MAX_TRESHOLD. • Check if the calculated ratio is above RATIO_THRESHOLD: o If ratio is above RATIO_THRESHOLD declare the band as mono; o Otherwise declare the band as Ambisonics • If none of these checks pass, then the band is declared as silence. • After the frequency band loop check if any band was declared as Ambisonics: o If true encoder frame is non-mono; o Otherwise check if any band was declared as mono: ^ If true encoder frame is mono; ^ Otherwise encoder frame is non-mono. • If the frame is mono increment mono frame counter. o If the frame counter is > 30ms then set the global mono flag to 1. • If the frame is non-mono reset the mono frame counter to zero.
[0092] FIG. 3 is a flow diagram that outlines some examples of band classification processes, which include iteration loops over frequency bands for a given frame. The example methods 300 may be partitioned into blocks, such as blocks 305, 310, 315, 320, 325, 327, 330, 335, 340, 345, 350 and 355. The various blocks may be described as operations, processes, methods, steps, acts or functions. Methods 300 may be performed by an apparatus or system that is configured to implement an encoder mono detector. The blocks of methods 300, like other methods described herein, are not necessarily performed in the order indicated. In some implementation, one or more of the blocks of methods 300 may be performed concurrently. Moreover, some implementations of methods 300 may include more or fewer blocks than shown and / or described. The blocks of methods 300 involve band classification processes, which may be performed by one or more devices, such as the device that is shown in FIG.7. Processing may commence at block 305.
[0093] Block 305 involves obtaining a frequency band of an audio signal. According to this example, the frequency band is an MDFT frequency band, but in other examples the frequency band may be another type of frequency band. Processing may continue to block 310. In this example, methods 300 involve the following:• For each frame of an audio signal, where the audio signal represents an audio scene with a primary channel and side channels (e.g., a frame loop); o Classify (Classifying) each of the MDFT frequency bands associated with the frame (e.g., a frequency band loop) based on power of the corresponding MDFT frequency band, by: ^ Calculate (calculating) power of the corresponding MDFT frequency band as: • W channel power (Wp) in block 310; and • Sum power (Sump) as a summation of the power of non-W channels in block 315. ^ Analyze (analyzing) the power calculations (Wp , Sump) to determine if the band is Silence, Mono, or Ambisonics, by: • Determine (determining), in block 320, if the sum power (Sump) is nominally below a minimum noise level ^, where ^ is a minimum value that may be close to zero. In some examples, the minimum noise level ^ may be equal to the NOISE_THRESHOLD; • When the sum power (Sump) is determined to be nominally below a miniumum noise level: o Determine (determining), in block 325, if the W channel power (Wp) is greater than or less than the NOISE_THRESHOLD; o Declare (declaring), in block 327, the band as Silence when the channel power is determined to be below the noise threshold (e.g., Wp≤ NOISE_THRESHOLD); o Declare (declaring), in block 330, the band as Mono when the chanel power is determined to be above the noise threshold (e.g., Wp > NOISE_THRESHOLD); and o Proceed (proceeding) to the next frequency band of the frame and revert to block 310. • When the sum power is determined, in block 320, to be nominally above the minimum noise level (e.g., Sump ≥ ^): o Calculate (calculating): ^ The ratio (R) of W channel power (Wp) to sum power (Sump) in block 335;^ The RATIO_THRESHOLD, in block 340. In some examples, block 340 may involve using a linear function with the W channel power (Wp) as input, by: • Normalize (normalizing) the W channel power (Norm_Wp) to be between 0 and 1 • Scale (scaling) the MAX_THRESHOLD by normalized W power • Clip (clipping) the calculated threshold to ensure the range is between MIN_THRESHOLD and MAX_TRESHOLD. o Determine (determining), in block 345, if the calculated ratio (R) is above or below the RATIO_THRESHOLD o Declare (declaring), in block 355, the band as Ambisonics when the calculated ratio is determined to be below the ratio threshold (e.g., R ≤ RATIO_THRESHOLD); o Declare (declaring), in block 350, the band as Mono when the calculated ratio is determined to be above the ratio threshold (e.g., R > RATIO_THRESHOLD); and o Proceed (proceeding) to the next frequency band of the frame, for example by reverting to block 310. o Some disclosed examples of the methods 300 involve marking the frame as one of Ambisonics, Mono or Silence based on the classified MDFT frequency bands for the frame, by: ^ Mark (marking) the frame as Ambisonics when one or more of the classified MDFT frequency bands for the frame is declared as Ambisonics; ^ Mark (marking) the frame as Mono when none of the classified MDFT frequency bands for the frame are declared as Ambisonics, and one or more of the classified MDFT frequency bands for the frame is declared as Mono; and ^ Mark (marking) the frame as Silence when none of the classified MDFT frequency bands for the frame are declared as Ambisonics, and none of the classified MDFT frequency bands for the frame are declared as Mono.
[0094] FIG.4 is a flow diagram that outlines some examples of frame marking processes, which include an iteration loop over frequency bands for a given frame. The example methods 400 may be partitioned into blocks, such as blocks 405, 410, 415, 420, 425, 430, 435, 440, 445 and 450.The various blocks may be described as operations, processes, methods, steps, acts or functions. Methods 400 may, for example, be performed after a frame classification process such as that of FIG.2B or FIG.3. Methods 400 may, in some examples, be performed by an apparatus or system that is configured to implement an encoder mono detector. The blocks of method 400, like other methods described herein, are not necessarily performed in the order indicated. In some implementation, one or more of the blocks of method 400 may be performed concurrently. Moreover, some implementations of method 400 may include more or fewer blocks than shown and / or described. The blocks of method 400 involve frame marking processes, which may be performed by one or more devices, such as the device that is shown in FIG. 7 Processing may commence at block 405.
[0095] Block 405 involves obtaining a classified frequency band of an audio frame. The classified frequency band may have previously been classified according to any of the disclosed classification processes. In this example, the classified frequency band has previously been classified as Ambisonics, Mono or Silence. Processing may continue to block 410.
[0096] Block 410 involves determining the classification of the classified frequency band obtained in block 405. In this example, the results of the classification are provided to block 415. Processing may continue to block 415.
[0097] Block 415 involves determining whether a frequency band has been classified as Ambisonics. In this example, if it is determined in block 415 that the frequency band has been classified as Ambisonics, the process flow continues from block 415 to block 420 wherein it is determined that the entire frame that includes the classified frequency band is not Mono. No further analysis is required, so in this example the process flow continues to “exit loop” block 425. However, if it is determined in block 415 that the frequency band has not been classified as Ambisonics, the process flow may continue from block 415 to block 430.
[0098] Block 430 involves determining whether a frequency band has been classified as Mono. According to this example, if it is determined in block 430 that the frequency band has been classified as Mono, the process flow may continue from block 430 to block 435, wherein the entire frame is provisionally marked as Mono. However, if it is determined in block 430 that the frequency band has not been classified as Mono, the process flow may continue from block 430 to block 440.
[0099] Block 440 involves determining whether the classified frequency band that has been evaluated in block 430 is the last classified frequency band of the audio frame. If not, the next classified frequency band is obtained and the process may revert from block 440 to block 410. However, if it is determined in block 440 that the classified frequency band that has been evaluatedin block 430 is the last classified frequency band of the audio frame, the process may continue from block 440 to block 445.
[0100] Block 445 involves determining whether the entire frame has been provisionally marked as Mono. According to this example, if it is determined in block 445 that the entire frame has been provisionally marked as Mono, the process may continue from block 445 to block 450. In this example, the entire frame is classified as Mono in block 450. The process flow may continue from block 450 to “exit loop” block 425. However, if it is determined in block 445 that the entire frame has not been provisionally marked as Mono, the process flow may continue to “exit loop” block 425. Transmit / Transmitting
[0101] In some examples, the mono detected information identified herein may be transmitted from the encoder to the decoder using an implicit signaling technique. In some such examples, existing metadata may be modified to specific values to support this technique, without the need for any additional data bit(s) in the bitstream. Thus, implicit signaling eliminates any potential increase in data transmission overhead such as potential increased bandwidth, speed, or data storage / buffering requirements, etc.
[0102] According to some examples, implicit signalling may be achieved by using the existing SBA metadata, and setting the SBA metadata to specific values. By using the existing SBA metadata, bit-exactness on existing tests can be preserved. In one example, 6 different SPAR and DIRAC metadata fields are set to zero at the encoder and then checked at the decoder. If the 6 metadata fields are zero in the decoder, then the decoder will process the frame as mono. The signalling also considers quantisation of the metadata values through the codec. Due to quantisation, certain values cannot be encoded in the bitstream. The detector in the decoder takes into account this quantisation to detect mono correctly.
[0103] Some example metadata values for implicit signaling include: • SPAR metadata o Prediction coefficients o Cross-prediction coefficients o Decorrelator coefficients • DiRAC metadata o Energy ratio o Azimuth o Elevation
[0104] In some examples, these metadata values are all set to zero when mono is detected. The SPAR values are related to extracting higher order channels out of the downmixed channels sent in the bitstream. When these values are set to zero, the extraction is not completed. The DiRACvalues are related to diffuseness and position. When these values are set to zero, the mono channel is not spread across other channels in the Ambisonics output. Decoder / Decoding Mono Detector / Preserver
[0105] According to some examples, the decoder is configured to analyze the signalling through the bitstream (or implicit signalling through the metadata). If mono is signalled, then in some implementations the decoder is configured to sets the DIRAC energy ratio to be zero and diffuseness to be one for every frequency band and subframe in the decoder. This ensures that the W channel energy is not spread across the other channels and the audio signal remains as mono content.
[0106] According to some examples, the decoder is configured to look for the 6 values mentioned above in each frequency band. In some such examples, if any of the values in any band do not match the mono expected values, then the whole block is declared non-mono. Otherwise, in some examples, the block is declared mono. According to some such examples, all of the expected values for mono are zero, except for energy_ratio which is expected to be < 0.15 due to quantisation of that value.
[0107] When mono is detected in the decoder, in some examples the energy_ratio values are reset to zero for all bands and time slices in the block. In some such examples, the decoded metadata does not apply to all frequency bands in the decoder and therefore the decoded metadata is expanded to cover all frequency bands. This may be the reason that the energy_ratio values are reset to zero. In some such examples, a diffuseness vector that applies to all frequency bands in the decoder is set to one. This results in the correct behaviour in the decoder for mono content contained within an Ambisonics format.
[0108] FIG.5 is a flow diagram that outlines additional example methods that may be performed by a decoder such as those disclosed herein. The example methods 500 may be partitioned into blocks, such as blocks 505, 510, 515, 520 and 525. The various blocks may be described as operations, processes, methods, steps, acts or functions. The blocks of methods 500, like other methods described herein, are not necessarily performed in the order indicated. In some implementation, one or more of the blocks of methods 500 may be performed concurrently. Moreover, some implementations of methods 500 may include more or fewer blocks than shown and / or described. The blocks of methods 500 involve audio signal decoding, which may be performed by one or more devices, such as the device that is shown in FIG. 7. Processing may commence at block 505.
[0109] Block 505 involves “obtaining an encoded bitstream.” In this example, the bitstream was previously encoded according to one of the disclosed encoding methods. In some examples, theencoded bitstream may be an Immersive Voice and Audio (IVAS) encoded bitstream. Processing may continue to block 510. Block 510 involves “decoding the encoded bitstream and obtaining downmix channels, spatial metadata, and a mono mode indicator.” In some examples, the mono mode indicator may be based on values of one or more spatial metadata parameters in the encoded bitstream that are set to values to indicate the mono mode. According to some examples, the one or more spatial metadata parameters may include one or more SPAR metadata parameters, one or more DirAC metadata parameters, or both. Processing may continue to block 515.
[0110] Block 515 involves “setting one or more parameters of spatial metadata to zero upon detecting a mono mode.” According to some examples, setting the one or more parameters of spatial metadata to zero upon detecting the mono mode may involve setting one or more energy ratio values to zero. In some examples, methods 500 may involve setting a diffuseness vector for all frequency bands to one upon detecting a mono mode. Processing may continue to block 520.
[0111] Block 520 involves “upmixing the downmix channels using the spatial metadata.” Processing may continue to block 525. Block 525 involves “rendering upmixed channels to a desired audio format.” In some examples, methods 500 may involve providing audio in the desired audio format to one or more loudspeakers of an audio environment. According to some examples, methods 500 may involve playing back the audio in the desired audio format by the one or more loudspeakers. Example System
[0112] FIG. 6 is a block diagram that illustrates some examples of an IVAS codec system. As with other disclosed implementations, the numbers, types and arrangements of elements shown in FIG.6 are merely examples. Other implementations may include more elements, fewer elements and / or different arrangements of elements.
[0113] In this example, the IVAS codec system 600 includes an example IVAS encoder 605 that is configured to encode dual ended detection, and an example IVAS decoder 655 that is configured to decode based on dual ended detection. The IVAS encoder 605 and the example IVAS decoder 655 may, in some examples, be implemented by instances of the electronic device architecture 700 that is shown in FIG. 7. In some examples, the IVAS codec system 600 may be configured for encoding and decoding according to one or more of the implicit signaling methods described above.
[0114] In the example shown in FIG.6, the IVAS encoder 605 includes a mono detector module 610 and an encoder module 615. In these examples, mono detector module 610 is configured to receive audio input signal 601 and is configured to output modified metadata 612 to the encoder module 615. The modified metadata 612 may include DirAC metadata, SPAR metadata, or a combination thereof. According to these examples, the audio data 601 is in Ambisonics format.Accordingly, the audio data 601 FOA or some type of HOA. The mono detector module 610 may be configured to perform any of the disclosed mono detection methods, dependiong on the particular implementation. According to these examples, the mono detector module 610 is configured for analyzing power of the primary channel of the audio signal, which is the Ambisonics W channel in this example. In these examples, the mono detector module 610 is also configured for analyzing powers of the side channels of the audio signal, which are at least the Ambisonics X, Y and Z channels in this instance. According to these examples, the mono detector module 610 is configured for detecting a mono mode for encoding the audio signal based on analyzing the power of the primary channel and the powers of the side channels.
[0115] In these examples, the encoder module 615 is configured to receive the audio input signal 601 and the modified metadata 612, and is is configured to output an encoded bitstream 620. According to these examples, the encoder module 615 is configured for computing one or more downmix channels and spatial metadata from the audio signal for a mono mode detected by the mono detector module 610. In these examples, the encoder module 615 is configured for encoding the one or more downmix channels and spatial metadata in a bitstream for the detected mono mode, and for indicating the mono mode in the encoded bitstream 620. According to these examples, the encoder module 615 is configured for both SPAR and DirAC encoding. In this example, the encoded bitstream 620 is an IVAS bitstream.
[0116] In the examples shown in FIG. 6, the IVAS decoder 655 includes a mono detector and preserver module 660 and a decoder module 655. Here, the IVAS decoder 655 is configured to receive the encoded bitstream 620 and to output rendered audio signals 670. According to this example, the IVAS decoder 655 is configured for decoding the encoded bitstream 620 and obtaining downmix channels, spatial metadata, and one or more mono mode indicators.
[0117] In these examples, the mono detector and preserver module 660 is configured to receive the encoded bitstream 620 and to output modified metadata 662 to the decoder module 655. The modified metadata 662 may include DirAC metadata, SPAR metadata, or a combination thereof. According to these examples, the mono detector and preserver module 660 is configured for detecting a mono mode based on the one or more mono mode indicators. In this example, the one or more mono mode indicators are based on values of one or more spatial metadata parameters in the encoded bitstream that are set to values to indicate the mono mode. According to these examples, the one or more mono mode indicators include one or more SPAR metadata parameters, one or more DirAC metadata parameters, or both. According to some examples, the mono detector and preserver module 660 may be configured for setting one or more parameters of spatial metadata to zero upon detecting a mono mode. For example, the mono detector and preserver module 660 may be configured for setting one or more energy ratio values to zero. In someexamples, the mono detector and preserver module 660 may be configured for setting a diffuseness value to one.
[0118] According to these examplse, the decoder module 655 is configured to receive the encoded bitstream 620 and the modified metadata 662, and is configured to output rendered audio signals 670. In this example, the decoder module 655 is configured for upmixing the downmix channels of the encoded bitstream 620 using spatial metadata of the the modified metadata 662, and is configured for rendering upmixed channels to a desired audio format, to produce rendered audio signals 670. According to these examples, the decoder module 655 is configured for upmixing the downmix channels using the spatial metadata and for rendering upmixed channels to a desired audio format based, at least in part, on modified spatial metadata received from the mono detector and preserver module 660. Example System Architecture
[0119] FIG. 7 shows a block diagram of an example electronic device architecture 700 suitable for implementing the systems, devices and methods described herein. Architecture 700 includes but is not limited to servers and client devices, as described herein. As shown, the architecture 700 includes central processing unit (CPU) 701 which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 702 or a program loaded from, for example, storage unit 708 to random access memory (RAM) 703. The CPU 701 may include one or more general purpose single- or multi-chip processors, one or more digital signal processor (DSPs), one or more application specific integrated circuit (ASICs), one or more field programmable gate array (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof. In RAM 703, the data required when CPU 701 performs the various processes is also stored, as required. CPU 701, ROM 702 and RAM 703 are connected to one another via bus 804. Input / output (I / O) interface 705 is also connected to bus 704.
[0120] The following components are connected to I / O interface 705: input unit 706, that may include a keyboard, a mouse, or the like; output unit 707 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 708 including a hard disk, or another suitable storage device; and communication unit 709 including a network interface card such as a network card (e.g., wired or wireless).
[0121] In some implementations, input unit 706 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0122] In some implementations, output unit 707 include systems with various number of speakers. Output unit 707 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0123] In some embodiments, communication unit 709 is configured to communicate with other devices (e.g., via a network). Drive 710 is also connected to I / O interface 705, as required. Removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 710, so that a computer program read therefrom is installed into storage unit 708, as required. A person skilled in the art would understand that although system 700 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure. 1.0 Algorithm Analysis
[0124] As previously described, SPAR seeks to maximize perceived audio quality while minimizing bitrate by reducing the energy of the transmitted audio data while still allowing the second-order statistics of the Ambisonics audio scene (i.e. the covariance) to be reconstructed at the decoder side using transmitted metadata. DirAC seeks to preserve direction and diffuseness of the dominant sounds in the input scene. Summaries of DirAC and SPAR technologies are described below in sections 1.1 and 1.3, respectively. 1.1 DirAC Technology 1.1.1 Reference Paper
[0125] DirAC technology is described in V. Pulkki, “Directional Audio Coding in Spatial Sound Reproduction and Stereo Upmixing,” in Laboratory of Acoustics and Audio Signal Processing, Helsinki University of Technology, Finland, 2006, which is hereby incorporated by reference. 1.2 Example Implementation of DirAC Analysis in MDFT Domain 1.2.1 DirAC Analysis in MDFT Domain
[0126] In an example implementation, the DirAC analysis block takes the time domain FOA channels of the ambisonics as an input and converts the FOA channels into frequency domain using Modified Discrete Fourier Transform (MDFT). Then, intensity and reference power is computed in the MDFT domain. Let ^^, ^^, ^^, ^^, ^^, ^^, ^^, ^^be the real and imaginary bin samples of the W, X, Y and Z channels of the FOA component of the Ambisonics input in the MDFT domain, then the intensity corresponding to the frequency bin f of the channel X iscomputed as^(^)^ = ^^ ∗ ^^ + ^^ ∗ ^^, [1]Similarly, intensity corresponding to Y and Z channels are computed.
[0127] The reference power computation ^ in frequency bin f is computed as^(^) = ^^∗ ^^+ ^^∗ ^^+ ^^∗ ^^+ ^^∗ ^^. [2]
[0128] The direction vector, dv, corresponding to X channel (or front-back direction) andfrequency bin f is computed as^^(^) ^(^)^^ =^(^)^^^^, where [3][4]Similarly, direction vector, to and Z channel (or top-bottom direction) are computed. 1.2.2 DirAC Parameter Estimation in Banded Domain
[0129] The intensity, reference power and direction vector per bin are then converted to the banded domain by applying the absolute response of a filterbank to the above computed values in [1], [2] and [3]. Let the banded intensity, reference power and direction vector in a particular frequency band be ^#, ^, ^^#, respectively, where s can be x, y or z. 1.2.2.1 DoA Angles (Azimuth and Elevation Angles) Computation
[0130] The azimuth and elevation of the dominant sound source within the scene for a particulartime-frequency tile are computed in degrees as$^ = arctan *+,-+,^. ∗ / 012 [5]^3 = ∗ 1.2.2.2
[0131] For diffuseness, long-term averaging of E and I is computed over N frames or M subframes. In an example implementation, a frame represents 20 ms of audio data and a subframe represents 5 ms of audio data and the long-term averaging of E and I is done over 160 ms of audio data, i.e., 8 frames or 32 subframes. Let the long-term average be ^#:^;,#, ^#:^;, then the diffuseness is given as 6*^7 7 7EF^G,^ 8 ^EF^G,- 8 ^EF^G,5.[7][8]^A@MN^ MOP=Q = 1 − B . [9]
[0132] The DirAC metadata parameters, i.e., DoA angles and diffuseness parameter, are quantized and coded by a metadata quantization and coding block. Based on the available bitrate, DirAC chooses N_dmx audio channels (also refered to as N_dmx downmix channels) out of N channel input, here N_dmx<= N and one of the channels in the N_dmx downmix channels is the W channel of the Ambisonics input to be coded by a core coder. Core coder bits and DirACmetadata bits are multiplexed into a bitstream and transmitted to a decoder. The decoder decodes the bitstream and reconstructs N_dmx downmix channels using a core decoder and DirAC metadata parameters using a metadata unquantization and decoding block. The N_dmx downmix channels and DirAC metadata parameters are fed into a DirAC synthesis and rendering block. The DirAC synthesis and rendering block computes the directional component of the output spatial audio scene using the W channel and spherical harmonics as per DoA angles. The DirAC synthesis and rendering block also computes the diffused component of the output spatial audio scene using a decorrelated version of the W channel, which is generated using a decorrelator block, and the diffuseness parameter in the DirAC metadata. The N_dmx downmix channels and directional and diffused components are then used to output the desired audio output format. 1.3 Example Implementation of SPAR (Spatial Reconstruction) With FOA Input
[0133] SPAR is a technology for efficient coding of spatial audio input. SPAR takes a multi- channel input and generates spatial metadata and a downmix signal such that the combination of spatial metadata and downmix signal can be coded with higher coding efficiency compared to coding each channel of the multi-channel input separately. SPAR aims to reproduce the covariance of an N channel multi-channel input and computes the spatial metadata and N_dmx channel downmix signal (where N_dmx <= N) based on the parameterized input covariance. The spatial metadata and downmix is quantized, coded and sent to the decoder. The decoder decodes the bitstream and unquantizes the spatial metadata and reconstructs downmix signal. The decoder then utilizes the spatial metadata and downmix and zero or more decorrelator(s) to reconstruct the multi-channel input audio scene. Example implementations of SPAR are further described in in PCT Patent Application No. PCT / US2023 / 010415, filed on January 9, 2023, for “Spatial Coding Of Higher Order Ambisonics For A Low Latency Immersive Audio CODEC.” 1.3.1 First Order Ambisonics (FOA) Input
[0134] With FOA input, consisting of channels W, Y, Z, X (in the ACN channel ordering convention), SPAR downmix signals can vary from 1 to 4 channels and the spatial metadata parameters include prediction parameters PR, cross-prediction parameters C, and decorrelation parameters P. These parameters are calculated from a covariance matrix of a windowed input audio signal and are calculated in a specified number of frequency bands (e.g., 12 frequency bands). An example representation of SPAR parameters extraction is described below. 1.3.1.1 Side Signal Prediction
[0135] Predict all side signals (Y, Z, X) from the primary audio signal W and compute the prediction coefficients for the residual channels using Equation
[0011] :S 1 0 0 0TU S= Y−−ZZMM[ 1 0 0 RTX ,
[0010] where, as an calculated as shownin Equation
[0011] :ZM = _`aand RYW = to channels Y and W.prZ and prX. PR is the vector of the predictions coefficients PR= [ZM[, ZM\, ZM]]k.
[0136] The above mentioned downmixing is also referred to as passive W downmixing in which W does not get changed during the downmix process. Another way of downmixing is active Wdownmixing which allows some mixing of Y, X and Z channels into the W channel as follows:SU = S + l[ ∗ T + l\ ∗ V + l] ∗ W,
[0012] where l[ is computed as a function of normalized input covariance RYW as l[ =^∗_`abcd (^∗_GG, (_``8 _gg8 _ff)) , here f and m are constants (e.g., f = 0.50, m = 3) and similarly l]and l\are computed as well. In some embodiments, l[, l], l\are computed as a function of active prediction coefficients as l[= ^ ∗ ZM[, l]= ^ ∗ ZM], l\= ^ ∗ ZM\, here, ZM[, ZM], ZM\are the active downmixing prediction coefficients and f is a constant (e.g., 0.50). In passive W, f=0 so there is no mixing of X, Y, Z channels into the W channel and W’ = W . 1.3.1.2 W Channel and Predicted Channels (mU, nU, oU) Remixed
[0137] The W channel and predicted channels ( TU, VU, WU) are remixed from most to least acoustically relevant, where remixing includes reordering or recombining channels based on some methodology, as shown in Equation
[0013] : SS
[0138] Note thatinput channels to W,TU, WU, VU, given the assumption that audio cues from left and right are more important than frontto back, and lastly up and down cues. 1.3.1.3 Post Prediction Covariance Computation
[0139] The covariance of the 4-channel post-prediction and remixing downmix are computed asshown in Equations
[0014] and
[0015] :u= [M@s=^][ZM@^=wP]. u. y yv^ [ZM@^=wP] [M@s=^] ,
[0014] u{{ u{+ u{|uv^= zu+{u++u+|},
[0015] where dd represents N-dmxth channels), and u represents theth to 4 channels).
[0140] For the example of a WABC downmix ith 1-4 downmix channels, d and u represent the following channels, where the placeholder variables A, B, C can be any combination of X, Y, Z channels in FOA): TABLE I N Residual Channels Predicted Channels 1 -- $U, pU, qU2 $UpU, qU3$U, pUqU4 $U, pU, qU-- 1.3.1.4 Extra C Coefficients
[0141] From these calculations, it is determined if it is possible to cross-predict any remaining portion of the fully parametric channels from the residual channels being sent. The required extraC coefficients are:q= u (u + ^sO^(~, PM(u ) ∗ 0.00 ) ^ / |+ ++ ++ 5 ) .
[0016]
[0142] Therefore, C has the shape (1x 2) for a 3-channel downmix, and (2x1) for a 2-channel downmix. One embodiment of spatial noise filling does not require these C parameters and these parameters can be set to 0. An alternate embodiment of spatial noise filling may also include C parameters. 1.3.1.5 Remaining Energy In Parameterized Channels
[0143] The remaining energy in parameterized channels that must be filled by decorrelators is calculated. The residual energy in the upmix channels Resuu is the difference between the actualenergy Ruu (post-prediction) and the regenerated cross-prediction energy Reguu:u@N|| = qu++qy,
[0017] u@?|| = u|| − u@N||,
[0018] where ?wO3@ is a normalization scaling factor. Scale can be a broadband value (e.g., ?wO3@ = 0.01) or frequency dependent, and may take a different value in different frequency bands (e.g., ?wO3@ = linspace (0.5, 0.01, 12) when the spectrum is divided into 12 bands). 1.4 Merging DirAC and SPAR
[0144] As stated previously, both DirAC and SPAR have different strengths and properties, and it is desired to combine the complementary aspects of each technology to produce a merged system that is advantageous in one or more of the following dimensions: higher audio quality, reduced bitrate, input / output format flexibility and / or reduced computational complexity. Some of the embodiments to efficiently merge these two technologies are listed below. 1.4.1 Frequency Based SPAR-DirAC Split
[0145] It has been observed that coding lower frequency bands with SPAR and higher frequency bands with DirAC improves the coding efficiency and quality at the decoder while reconstructing the spatial audio scene. It may also be desirable to code low frequency bands with SPAR to reconstruct an input covariance at the output, or to code higher frequency bands with DirAC, with the same, or finer, time resolution, to perform efficient upmix of SPAR reconstructed FoA signals to HOA.
[0146] At the encoder, a first embodiment: 1) uses a filterbank to convert time domain broadband Ambisonics input into a frequency banded domain; 2) performs DirAC analysis in high frequency bands and obtains DirAC MD parameters in high frequency bands; 3) performs SPAR analysis in low frequency bands and obtains SPAR MD parameters in low frequency bands; 4) obtains SPAR MD parameters in high frequency bands by converting DirAC MD parameters into SPAR MD using a MD conversion routine (D2S) (mentioned in sections 1.5 to 2.4); 5) generates a downmix matrix from SPAR MD and applying the downmix matrix to input channels obtains downmix channels as mentioned in section 2.3; 6) quantizes and encodes the SPAR MD parameters in low frequency bands and DirAC MD parameters in high frequency bands; 7) encodes downmix channels using a core audio coder; and 8) multiplexes MD bits and core coder bits into a bitstream and transmits the bitstream to a decoder.
[0147] At the decoder, a second embodiment: 1) obtains MD bits and core coder bits from the bitstream; 2) decodes the downmix channels using a core audio decoder; 3) decodes and unquantizes the low frequency SPAR MD parameters and the high frequency DirAC MD parameters from the MD bits; 4) obtains high frequency band SPAR MD from DirAC MD using a D2S conversion routine; 5) performs filter bank analysis on the decoded downmix channels; 6) generates a SPAR upmix in the filterbank domain using the SPAR MD in all frequency bands; and 7) generates spatial audio output at the decoder. In some embodiments, as part of step 7), filterbank synthesis is done on SPAR upmixed channels to reconstruct Ambisonics channels at the decoder.In some other embodiments, as part of step 7), DirAC analysis is done on the upmix channels generated by SPAR, obtaining DirAC MD parameters in all frequency bands and performing a DirAC upmix to a desired output format including but not limited to HOA2 / HOA3.
[0148] At the decoder, a third embodiment: 1) obtains MD bits and core coder bits from the bitstream; 2) decodes downmix channels using a core audio decoder; 3) decodes and unquantizes the low frequency SPAR MD parameters and the high frequency DirAC MD parameters from MD bits; 4) obtains high frequency band SPAR metadata (MD) from DirAC MD using a D2S conversion routine and low frequency band DirAC MD from the SPAR MD and / or the downmix covariance using a SPAR to DiRAC (S2D) MD conversion routine (mentioned in section 2.4); 5) performs filterbank analysis on the decoded downmix channels; 6) generates a SPAR upmix in filterbank domain using SPAR MD in all frequency bands; and 7) generates spatial audio output at the decoder. In some embodiments, as part of step 7), filterbank synthesis is done on SPAR upmixed channels to reconstruct Ambisonics channels at the decoder. In some other embodiments, as part of step 7), DirAC MD parameters in all frequency bands, including the low frequency DirAC MD obtained in step 4) are applied to the SPAR upmix to perform a DirAC upmix to desired output format including but not limited to HOA2 / HOA3. 1.4.2 Channel Based SPAR-DirAC Split (same processing in all bands)
[0149] In some embodiments, a subset of Ambisonics input channels may be reconstructed via SPAR (either residually or parametrically), and some channels are reconstructed by DirAC. Any further upmix to a higher order is also handled by DirAC. SPAR reconstructs at least enough channels for DirAC analysis to be performed in the decoder, where generally DirAC analysis requires FOA channels (or planar FOA channels for the planar case). As used herein, residual coding is direct audio coding of the residual from which the output channel is reconstructed along with the predicted component from W, and parametric coding is coding of cross-prediciton and decorrelation parameters from which the output is reconstructed, along with the predicted component from W and the cross-predicted component of residuals and decorrelated version of W.
[0150] SPAR generally operates with a B-format representation of input and output Ambisonics audio. DirAC, in some cases, reconstructs the audio signal in A-format or Equivalent Spatial Domain (ESD), and in other cases, in B-format. The description that follows addresses mainly the latter B-format case. However, analogous embodiments are possible for DirAC synthesis in A- format or ESD. To that end, the SPAR reconstructed B-format channels may be used to generate a relatively sparse set of DirAC prototype signals in B-, A-format or ESD from which DirAC synthesis generates a denser set of upmix signals, where each of the upmix signals may drive a speaker of a multi-loudspeaker system. Such a multi-loudspeaker system may correspond to a realloudspeaker setup like, e.g., 7.1.4 or 5.1 or a virtual loudspeaker system which is an intermediate step to immersive binaural rendering of the synthesized audio signal.
[0151] Various embodiments of the above are described and tabulated in Table II below. TABLE II Ambisonics Input SPAR DirAC DirAC Reconstruction Reconstruction Blind Upmix Channels Channels Channels HOA3 FOA All 2nd+ 3rdorder N / A FOA + 2ndorder 2ndorder non-planar N / A planar + all 3rdorder FOA + 2ndand 3rd2ndand 3rdorder N / A order planar non-planar HOA2 FOA All 2ndorder All 3rdorder FOA + 2ndorder 2ndorder non-planar All 3rdorder planar Planar HOA N > 1 Planar channels, all Planar channels, Planar channels, orders up to n<N orders n<m <= N orders N<p<= 3 FOA, n_dmx = 4 FOA N / A All 2ndand / or 3rdorder FOA, n_dmx < 4 FOA N / A All 2ndand / or 3rdorder Planar FOA Z All 2ndand / or 3rdorder Planar FOA, n_dmx = 3 Planar FOA N / A Planar 2ndand / or 3rdorder Planar FOA, n_dmx < 3 W, Y X Planar 2ndand / or 3rdorder
[0152] For HOA3 input, channels are reconstructed according to the following options: FOA, or HOA2, or FOA + 2nd order planar channel, or FOA + 2nd + 3rd order planar channels are reconstructed with SPAR, while the HOA2 and HOA3, or HOA3, or 2nd order non-planar and HOA3, or 2nd and 3rd order non-planar channels are reconstructed using DirAC to reduce computational complexity without compromising the quality.
[0153] For FOA input, bitrates where SPAR has less than 4 downmix channels: 1) FOA with SPAR is implemented with a DirAC blind upmix to HOA2 / HOA3, or 2) planar FOA with SPAR, a DirAC upmix to full FOA, with possible blind upmix HOA2 / HOA3 with DirAC.
[0154] For FOA input, bitrates where SPAR has 4 downmix channels: 1) FOA with SPAR is implemented with a blind upmix to HOA2 / HOA3 with DirAC.
[0155] For planar FOA input, bitrates where SPAR has less than 3 downmix channels, WY reconstruction with SPAR is implemented, along with an upmix to planar FOA, with possible blind upmix to planar HOA2 / HOA3 with DirAC.
[0156] For planar FOA input, bitrates where SPAR has 3 downmix channels, WYX reconstruction with SPAR is implemented with a blind upmix to planar HOA2 / HOA3 with DirAC. 1.4.3 Reconstruction of Individual Channels in Part by SPAR and by DirAC
[0157] In conjunction with the channel based SPAR-DirAC split technique disclosed in section 1.4.2, a further category of channel could be introduced that is parametrically reconstructed in part from both SPAR and DirAC methods. The motivation for this is to reduce reliance on large numbers of decorrelator outputs at the decoder, which could reduce mixing complexity figures. This approach uses SPAR prediction and cross-prediction to reconstruct the majority of a particular parametrically reconstructed signal, and then relies on DirAC diffuseness to restore any missing covariance. 1.4.4 Alternative Method To Reduce Usage of Decorrelators in Decoder
[0158] Instead of adding decorrelation in proportion with the decorrelation coefficients, in this embodiment energy matching of the cross- / prediction parametrically constructed channel is achieved by applying a gain derived from the SPAR coefficients. A particular Ambisonics signalS can be parametrically reconstructed as follows:^^$u u@wQA?PM>wP@^ ^ = ZM#S + ∑^^^^^ / U^^ / q^#u^ + ^# ∗ ^(S)
[0021] wherewith S, and residual signal R (e.g. Y’, Z’, X’, …).^O=A O^^>?P@^ ^ = N#^^, ^ℎ@M@(|v^E|78|^E|7){78∑^^^^^^^^ |^^E|7^_^^ ^71.4.5
[0159] In some embodiments, sections 1.4.1 and 1.4.2 are combined to get the benefit of merging SPAR and DirAC by doing a combination of frequency based split and channel based split. In an example implementation, input to merged SPAR-DirAC system is an N channel Ambisonics signal. Out of these N channels, M channels are fed into a SPAR subsystem, where M <= N. In an embodiment, these M channels contain FoA channels. In some other embodiments, these M channels include FOA and planar HOA channels. SPAR can then operate in any downmix configuration based on the operating bitrate, where the number of downmix channels ^+^^is such that 1<=^+^^<=M. For low frequencies, SPAR computes SPAR parameters including prediction, cross-prediction and decorrelation parameters based on methods described in section 1.3, whereasfor higher frequencies DirAC parameters are computed as described in section 1.2, and SPAR parameters are estimated from DirAC parameters as described in sections 1.5 to 2.4 below. In some embodiments, SPAR computes SPAR parameters for high frequencies as well for a subset of input channels based on methods described in section 1.3.
[0160] The M channels reconstructed by SPAR on decoder side are then used by DirAC to reconstruct a representation of original N channel input scene.
[0161] Example implementations of a combined frequency based and channel-based split with HOA3 input to a SPAR-DirAC merged system are given below. 1.4.5.1 Example Encoder Embodiment 1
[0162] FIG.8 is a block diagram of an encoder 800 with frequency-based and channel-based split between SPAR and DirAC, according to one or more embodiments. In this embodiment, SPAR is operating in 4 channel downmix mode. Input into encoder 800 is a HOA3 (3rd Order Ambisonics) signal. DirAC parameter estimator 801 estimates the DirAC parameters which are limited to high frequencies and computed as per section 1.2 based on FOA channels in the Ambisonics input. The estimated DirAC parameters are quantized and coded 802 and the quantized DirAC MD are converted 803 to SPAR MD.
[0163] SPAR analysis and metadata computation 804 is based on FOA, planar HOA2 and planar HOA3 channels in the low frequencies as per section 1.3. The SPAR metadata is quantized and coded 805 and the quantized SPAR metadata in low frequencies and SPAR MD obtained from DirAC MD in high frequencies is converted into a downmix matrix 806. An MDFT transform 807 is applied to the FOA, planar HOA2 and planar HOA3 signals. The MDFT coefficients and downmix matrix are frequency band mixed with cross-fades using a filterbank mixer 808 to generate a 4-channel downmix. The 4-channel downmix is coded by one or more core codecs 809 (e.g., Enhanced Voice Services (EVS) encoder). The SPAR metadata coded in low frequencies and the DirAC metadata coded in high frequencies are packed together with the core codec coded bits to form final bitstream 810 output by encoder 800.
[0164] Encoder 800 is one example embodiment of an encoder that combines DirAC and SPAR. In other embodiments, SPAR and DirAC are combined by only frequency splitting or by only channel splitting. 1.4.5.1.2 Example Decoder Embodiment 1
[0165] FIG. 9 is a block diagram of decoder 900 with frequency-based and channel-based split between SPAR and DirAC, according to one or more embodiments. In this embodiment, decoder 900 receives bitstream 901(810) and provides the core codec encoded bits to one or more core codec decoder(s) 907 (e.g., EVS decoder(s)). DirAC MD 902 in the high frequencies is decoded and then converted to SPAR MD 903 in the high frequencies using DirAC MD to SPAR MDconversion 913, in an embodiment 913 at the decoder is same as 803 at the encoder. SPAR MD in the bitstream is decoded to reconstruct SPAR metadata 904 in low frequencies. A SPAR upmix matrix 905 is generated using the low frequency SPAR metadata 904 extracted from bitstream 910 and the high frequency SPAR metadata 903 converted from the high frequency DirAC metadata. The downmix channels are reconstructed by one or more instances of core decoders 907 and converted into a frequency banded domain by filterbank 908 (e.g., CLDFB filterbank, Quadrature Mirror filterbank (QMF), etc.).
[0166] In some embodiments, the primary downmix channels are input into decorrelator(s) 909 and the outputs of decorrelator(s) 909 are input together with the upmix matrix into SPAR upmixing unit 906 to reconstruct the FOA, planar HOA2 and planar HOA3 channels. The decorrelation can be implemented in the time domain or frequency banded domain (e.g., CLDFB domain). The decorrelator(s) may either generate time domain decorrelated output and then convert it into frequency banded domain, or convert input into frequency banded domain and generate decorrelated outputs in frequency banded domain. The output channels of 906 are fed into DirAC parameter estimator 910, which estimates the DirAC metadata in low frequencies based on the reconstructed FOA signal in the frequency banded domain. DirAC upmixer 911 uses the low frequency DirAC metadata and the high frequency DirAC metadata to upmix the FOA, planar HOA2 and planar HOA3 channels into the 16 HOA3 channels, which is a frequency banded domain representation of the original 16 channel HOA3 input to encoder 200. Synthesizer 312 (e.g., CLDFB synthesizer) synthesizes / renders the 16 channel HOA3 frequency banded domain representation into the time domain representation for playback on various audio systems with different speaker configurations and capabilities. 1.4.5.2.1 Example Encoder Embodiment 2
[0167] FIG. 10 is a block diagram of an alternate encoder 1000 with frequency-based and channel-based split between SPAR and DirAC, according to one or more embodiments. In this embodiment, input into encoder 1000 is an HOA3 signal. The DirAC parameters are estimated 1001 and quantized and coded 1002. The DirAC parameter estimation is limited to high frequencies and is done as per section 1.2 based on FOA channels. The SPAR analyses and metadata computation 1004 and quantization and coding 1005 is done in the low frequencies based on FOA, planar HOA2 and HOA3 channels plus zero or more non-planar channels (e.g., non- planar channels), as per section 1.3.
[0168] For high frequencies, SPAR analysis and parameter estimation is done for non-FoA channels (this is not done in system 200 ) as per section 2.2.7.2. In this embodiment, SPAR is operating in 4 channel downmix mode and to obtain a SPAR downmixing matrix for allfrequencies, SPAR FoA metadata at high frequencies is estimated based on DirAC metadata using the methods described in section 2.2.
[0169] The quantized and coded SPAR metadata is used to generate a downmix matrix 1007. An MDFT transform 1006 is applied to the FOA, planar HOA2 and planar HOA3 signals. The MDFT coefficients and downmix matrix are frequency band mixed with cross-fades 1008 to generate a 4- channel downmix. The 4-channel downmix is coded by one or more core codecs 1009. The SPAR metadata coded in low frequencies for FOA channels and all frequencies for HOA channels and the DirAC metadata coded in high frequencies are packed together with the core codec coded bits to form final bitstream 1010 output by encoder 1000.
[0170] Downmixed channels are coded 1009 by one or more core codecs (e.g., EVS). For FOA channels, SPAR metadata is coded for low frequencies whereas DirAC metadata is coded for high frequencies, while for non-FOA channels SPAR metadata is coded for the entire frequency range, and packed together with core codec coded bits to form the final bitstream 1010 output by encoder 1000. In this embodiment, SPAR metadata computation for HOA2 and HOA3 channels in high frequencies is done as per methods described in section 2.2.7.2. Further in this embodiment, as per methods described in section 2.2.7.2, SPAR metadata computation for HOA2 and HOA3 channels in high frequencies 1004 depends on SPAR MD for FOA channels in high frequencies that is estimated form DirAC MD in high frequencies 303.
[0171] Note that in Embodiment 2 DirAC MD to SPAR MD conversion only happens for FOA channels, such that fullband SPAR MD is used for any HOA channels handled by SPAR. In general, any number of non-planar HOA channels could be handled by SPAR. In embodiment 2 only 1 non-planar HOA channel was added. Also, while these embodiments focus on 4 downmix channels (Ndmx = 4), any number of transport channels (e.g., from 1-16) is possible. 1.4.5.2.2 Example Decoder Embodiment 2
[0172] FIG. 11 is a block diagram of an alternate decoder 1100 with frequency-based and channel-based split between SPAR and DirAC, all arranged in accordance with embodiments described herein.
[0173] In this embodiment, decoder 1100 receives the coded bitstream 1104 and provides core codec coded bits to one or more core decoders 1105. DirAC MD 1102 in the high frequencies is decoded and then converted to SPAR MD 1103 in the high frequencies using DirAC MD to SPAR MD conversion 1113. In an embodiment, DirAC MD to SPAR MD conversion 1113 at the decoder is same as DirAC MD to SPAR MD conversion 903 at the encoder. SPAR MD 1104 corresponding to FOA and planar HOA and zero or more non-planar HOA channels is decoded and fed into SPAR mixing matrix 1106. Missing SPAR MD 1103 for the FOA channels in high frequencies is estimated from DirAC MD in the same way as the encoder 900. A SPAR upmix matrix 1106 isgenerated using the SPAR MD 1104 extracted from bitstream 1110 and the high frequency SPAR MD 1103 converted from the high frequency DirAC MD. Downmix channels that are reconstructed by one or more instances of core decoders 1105 are converted into frequency banded domain with the help of a filterbank analyses 1107, and the upmix matrix 1106 is applied to reconstruct FOA, planar HOA2, planar HOA3 channels and zero or more non-planar channels.
[0174] The decoded downmix channels output from the one or more core decoders 1105 are fed into decorrelator(s) 1109 and the outputs of decorrelator(s) 1109 are input together with the upmix matrix into SPAR upmixing unit 1108 to reconstruct the FOA, planar HOA2 and planar HOA3 channels. The decorrelation can be implemented in the time domain or frequency banded domain (e.g., CLDFB domain). The decorrelator(s) may either generate time domain decorrelated output and then convert it into the frequency banded domain, or convert the input into the frequency banded domain and generate decorrelated outputs in the frequency banded domain. The output channels of 1108 are fed into DirAC parameter estimator 1110, which estimates the DirAC metadata in low frequencies based on the reconstructed FOA signal in the frequency banded domain and uses the DirAC parameters in high frequencies extracted from bitstream 1101. Alternatively, DirAC upmixer 1108 may estimate DirAC parameters in the entire frequency range based on the FOA signal in the frequency band domain (e.g., CLDFB domain) and ignore the DirAC parameters in high frequencies from the bitstream 1101.
[0175] DirAC upmixer 1111, uses the DirAC metadata from 1110 and 1102 and converts the FOA, planar HOA2, planar HOA3 and zero or more non-planar channels into an HOA3 output which is a frequency band domain (e.g., CLDFB domain) representation of the original 16 channel HOA3 input to encoder 900. Synthesizer 1112 (e.g., CLDFB synthesizer) synthesizes / renders the 16 channel HOA3 frequency band domain representation into the time domain representation for playback on various audio systems with different speaker configurations and capabilities. It should be noted that output of decorrelator(s) 1109 is in CLDFB domain such that it covers embodiments where a time domain decorrelator is followed by CLDFB analyses and CLDFB analyses with CLDFB domain decorrelation. 1.4.6 Diffuseness in DirAC Upmix Channels
[0176] When estimating higher order channels from first order channels using the DirAC approach, the directional panning in the higher order channels that are upmixed using DirAC approach can be done by using DOA angles and spherical harmonics responses. However, the addition of diffuseness and decorrelation to these higher order channels should be handled carefully as it is known that too much decorrelation may hurt the audio quality and too little decorrelation may cause spatial collapse.
[0177] Below mentioned are embodiments for adding diffuseness to the higher order channels that are upmixed by DirAC approach: 1.4.6.1 Add Uniform Decorrelation to All HOA Upmixed Channels
[0178] In some embodiments, Nd decorrelated channels that are uncorrelated with respect to W channel are computed, where Nd is the number of HOA channels that are to be upmixed by DirAC from FOA channels. The B (diffuseness) is computed using one of ways described in thisdocument and then compute:<=^^>?@A@??^^^^^^(=) = B ∗ ^QMs(=),
[0024] where i is the channel index and Norm is the corresponding normalization factor that iscomputed as per given Ambisonics normalization, e.g., SN3D normalization.<=^^>?@A@??^^^^^^(=) is applied to the =th decorrelated channel to get the diffused componentfor the corresponding HOA channel. In an example embodiment, if input to the DirAC upmixer is FOA channels (4 channels) and HOA3 channels are to be upmixed from input FOA channels,then the number of decorrelator outputs needed are 12 (Nd = 12). The upmixed HOA channel^(=) can represented as:^(=) = @A@MN^_uOP=Q_^OwPQM ∗ u@?Z^ ∗ S + <=^^>?@A@??^^^^^^(=) ∗ <^(S)
[0025] computed using DOA angle ^ , where in ^ can be represented in terms of azimuth and elevation angles. The @A@MN^_uOP=Q_^OwPQM can be computed as (1 − B). <^(S) is the ith decorrelated channel.
[0179] The above approach may result in too much decorrelation and may make the reconstructed scene more diffused than desired. Also, it may be computationally expensive to generate too many decorrelator outputs and scaling them to get desired diffuseness levels. 1.4.6.2 Add Directional Decorrelation to All Upmixed Channels
[0180] In this embodiment, directional diffuseness information is sent from the encoder to the decoder. The decoder uses this directional diffuseness information, and adds only a desired amount of decorrelation to the upmixed HOA channel. This method is applicable to cases where input to the encoder is HOA and due to bitrate and complexity limitation, only a few selected channels are reconstructed using SPAR, whereas the remaining channels are upmixed using DirAC. In an example implementation, the encoder can compute directional diffuseness using P (decorrelation) coefficients computed by SPAR in section 1.3. This method uses additional information to be sent to the decoder from the encoder.1.4.6.3 Add Decorrelation to Selected Upmixed Channels
[0181] In this embodiment, the addition of diffuseness is limited to a few selected channels to keep the overall diffuseness within desired limits. This method also reduces computational complexity. The selection of channels for diffuseness addition can be static or dynamic based on signal characteristics. 1.4.6.3.1 Static Selection of Channels
[0182] In this embodiment, decorrelation is added to a selected few HOA channels. These channels are chosen based on perceptual importance. In an example implementation, if FOA and planar HOA channels are reconstructed by SPAR, and only non-planar HOA channels are to be upmixed using DirAC to get HOA3 output in ACN-SN3D format, then channel index 6, 10, 12, 14 (channel index ranging from 0 to 15) can be chosen to add decorrelation. This method does not require any additional information to be sent to decoder. 1.4.6.3.2 Dynamic Selection of Channels
[0183] In this embodiment, the directional diffuseness information is computed at the encoder and sent to the decoder to select the channels to which diffuseness is to be added while upmixing. This embodiment is only applicable to cases where input to the encoder is HOA. Only the channels in which the amount of decorrelation needed is higher than a first threshold value are chosen at the DirAC decoder to add decorrelation. In an example implementation, the encoder computes directional diffuseness using P (decorrelation) coefficients computed by SPAR in section 1.3, compares the P coefficients values against a first threshold and codes the channel indices which have P coefficients higher than a first threshold value. These indices are read by the decoder. If the number of channel indices exceeds a second threshold value, then limited indices can be chosen based on P coefficients values and perceptual importance of a given channel. This embodiment requires additional information to be sent to decoder from encoder.
[0184] To perform frequency based split as mentioned in section 1.4.1., an efficient mechanism is desired to convert DirAC metadata to SPAR in DirAC frequency bands and SPAR metadata to DirAC in SPAR frequency bands, so that DirAC and SPAR metadata can be reconstructed in all bands when required to perform upmix or downmix. Below are example embodiments to convert DirAC metadata to SPAR and SPAR metadata to DirAC. 1.5 DirAC to SPAR Conversion
[0185] In some embodiments, an approximation of input covariance matrix is computed based on quantized DirAC MD parameters (Azimuth angle (Az), Elevation angle (El), diffuseness). Az and El are also referred to as DOA angle θ in this document.1.5.1 Equations
[0186] In some embodiments, the model-covariance blocks calculate the covariance matrix and prediction coefficients from the DirAC DOAs and diffuseness as follows u;,; u;,^ u;,! u;,"æ u^,! u^," öHere, R is a covariance is estimated using DirAC metadata. Example1.5.2 Example Covariance Computations
[0187] In some embodiments, the covariance is computed as follows:u@?Z; = T1,1(^ ), u@?Z^ = T / , / (^ ), u@?Z! = T / ,^ / (^ ), u@?Z" = T / ,1(^ )
[0027] where u@?Z; = 1 and ^(u@?Z^ + u@?Z! + u@?Z" ) = 1,
[0028] for S variance u;; = u;; + B^
[0029] for side channel variance u / / ^^ = u ^^ + B^ ∗ √³ ∗ √³
[0030] wherein, i and j can be ^, ^, ^, ^.
[0188] In the above equation, E is an approximation of overall signal energy (as given in
[0033] below). This is obtained by adding a rough estimation of directional energy and diffused energy. Let ^^be the real bin sample of W channel in MDFT domain, the energy corresponding to eachbin is computed as follows^(^) = ^^ ∗ ^^.
[0031]
[0189] The energy is then converted into frequency banded power by applying filterbank responses of each band. The frequency banded energy in each band is extrapolated to computeoverall signal energy as follows^= ^ ∗ (u@?Z; + u@?Z^ + u@?Z! + u@?Z" )
[0032]
[0190] The^= ^ ∗ (1 + B ).
[0033]
[0191] In some embodiments, when the covariance smoothing is turned off, the above computed covariance is used to calculate SPAR coefficients as usual.2.0 Other Embodiments 2.1 DirAC MD Computation 2.1.1 Improved Computation of DirAC Diffuseness
[0192] DirAC needs time smoothing to compute diffuseness parameter. In some embodiments, a simple parameter averaging is performed over 160ms (Eqn.12 from Section 1.2.2.2 )
[0193] In some embodiments, SPAR’s Covariance smoothing and / or the transient detector- ducker algorithms can be used to improve computation of the DirAC diffuseness parameter. For example, SPAR’s covariance smoothing algorithm, described in PCT Application No. PCT / 2020 / 044670, filed July 31, 2020, for “Systems and Methods for Covariance Smoothing,” can be adapted to weigh recent audio events more heavily that events further into the past, and can do this differently at each frequency band. This may be advantageous over a simple averaging operation. Using transient detection and ducking, the diffuseness value could be instantaneously reduced during short transients without disturbing the long-term smoothing process.
[0194] Because the time smoothing causes long-term time dependence as well as smoothness over time, in another embodiment differential coding can be used to reduce MD bitrate and improve frame loss resilience. 2.1.2 Computation of DirAC Metadata in Frequency Banded Covariance Domain
[0195] Based on the DirAC analysis captured in section 1.0, in some embodiments DirAC MD can be computed based on input frequency banded covariance matrix instead of computing DirAC MD in the FFT (Fast Fourier Transform) or MDFT domain and then converting it into a frequency banded domain.
[0196] In some embodiments, computation of SPAR metadata can be done based on an input frequency banded covariance as shown in section 1.0.
[0197] In some embodiments, computing both SPAR and DirAC metadata from the input covariance allows for better conversion of SPAR to DirAC and DirAC to SPAR MD in the desired bands. It is also computationally efficient. Below is an example of how DirAC MD can be computed from input covariance. 1. Compute an N*N frequency banded covariance matrix, where N is the number of input channels. 2. Smooth the covariance matrix as mentioned in section 2.1.1. 3. Compute reference power as trace of covariance matrix. 4. Compute intensity as u;^, u;!, u;", here u;^, u;!, u;"are covariance of W channel and X, Y, Z channels. 5. Compute intensity norm as ^^^^^=^(u;^+ u;!+ u;")6. Compute direction vector as ^^;#=_GE7 , here, s can be ^, ^, ^, and then 6(_G^ 8 _G7- 8 _G75 )compute azimuth and elevation and [6] in Section 1.2.2.1.
[0198] Similarly, diffuseness computationon frequency banded covariance matrix as follows.
[0199] For diffuseness, first, reference power E and intensity I of input signal are computed ina given frequency band.^ = u;;+ u!!+ u^^+ u"",
[0034] ^^ = u;^, ^! = u;! , ^" = u;".
[0035]
[0200] Given that covariance is already smoothed as mentioned in section 2.1.1, then the diffuseness can be computed as 6^^^78 ^-78 ^57^<=^^>?@A@?? = B = 1 − ,^A@MN^ MOP=Q = 1 − B.
[0038]
[0201] In some embodiments, before computing diffuseness, E and I are further averaged usinga long term averaging filter as given below:^^= (1 − ^^) ∗ ^^^ / + ^^∗ ^
[0039]
[0202] Here, ^^and these values are then used, instead of E and I, in the computation of diffuseness computation equation
[0036] . The factors ^^and ^^in
[0039] and
[0040] are examples of smoothing factors. 2.1.3 Improvement to reference power (E) computation
[0203] In some embodiments, an alternate method can be used to compute reference power that results in better estimates of diffuseness and leads to better estimates of SPAR coefficients when they are derived from DirAC coefficients.
[0204] For diffuseness, first, reference power E and intensity I of input signal are computed ina given frequency band:^ = u;;+ u!!+ u^^+ u"",
[0041] ^^ = u;^, ^! = u;! , ^" = u;".
[0042] Here, u^©is the covariance between ith and jth channel.
[0205] The reference power is computed as^= max (u;;, 0.5 ∗ ^).
[0043]
[0206] E computed in
[0043] provides better estimates for diffuseness and SPAR coefficients in cases where W channel energy is higher than 0.5*E. Diffuseness is computed as 6^^·7^ 8 ^·7- 8 ^·75 ^<=^^>?@A@?? = B = 1 −H·.
[0044] Here, ^^, ^^^, ^^!, ^" . Alternatively,^ , ^ , ^^^^ ^!, ^^" may
[0207] Diffuseness can then be limited asB= max (0, min(1, B)),
[0045] ^A@MN^ = 1 − B.
[0208] SPAR can any of the methods described in this document. 2.2 Improvements to DirAC to SPAR MD Conversion 2.2.1 Alternative ways to compute covariance / spar MD from DirAC
[0209] In some embodiments, passive prediction coefficients can also be computed as u@?Z^ ∗u@?Z© , wherein i and j can be ^, ^, ^, ^, which should be similar to the direction vector, ^^, fora given side channel. This way prediction coefficients will be close to actual SPAR prediction coefficients when variance of W channel is less than ^^^^^in frequency banded domain. In some embodiments, the additional parameter_GG^^^^^, can be sent to the decoder for a better estimate of prediction coefficients when theof W channel is greater than ^^^^^. In some embodiments, prediction coefficients may also be computed directly from DirAC metadata. 2.2.2 Quantization of DirAC metadata
[0210] In some embodiments, SPAR MD is computed based on quantized DirAC MD. 2.2.3 Generic reconstruction of SPAR coefficients from DirAC metadata for any downmix configuration
[0211] In some embodiments, the input covariance, R, is a 4x4 matrix computed based on DirACparameters as follows:u^© = (1 − wB) ∗ ^ ∗ u@?Z^ ∗ u@?Z© when = ! = ^,u;; = ^ ∗ u@?Z; ∗ u@?Z;,^ when =! = w ,
[0047] (^ ), u@?Z^ = T / , / (^ ), u@?Z! =T / ,^ / (^ ), u@?Z" = T / ,1(^ ) are the spherical harmonics, º^ and c are constants in the range 0and 1. Setting c = 1 and º^= 1 would make it similar to equations mentioned in section 1.5, in which case both encoder and decoder would have prior knowledge about these constants. Insome embodiments º^OA^ w can be dynamically computed based on the actual input covariance matrix and above mentioned approximation of input matrix from DirAC parameters.
[0212] Given that SPAR coefficients are normalized with respect to covariance, SPAR coefficients derived from input covariance u are equal to SPAR coefficients derived from ^ ∗ u, where E can either be the variance of the W channel or overall signal energy or any constant.
[0213] In some embodiments, a normalized covariance matrix R_norm is derived based on DirAC parameters only. R_norm is a 4x4 covariance matrix for FOA channels and is an approximation of actual normalized input covariance matrix, where the actual input covariancematrix is given as:u^^ = »»k, 4x4 covariance matrix for FOA input channels, whereU = [S W T V]k, FOA input, and R_normcan be computed based on DirAC parameters only as given below:u^^^^^©= (1 − wB) ∗ u@?Z^∗ u@?Z©when = ! = ^, and u_AQMs;;= u@?Z;∗u@?Z^ + º^ when =! = ^ wℎOAA@3 =A^@^.
[0048]
[0214] SPAR coefficients, including prediction, cross prediction and decorrelation coefficients, are computed from normalized covariance u_AQMs^©as disclosed in section 1.3. 2.2.3.1 Example Reconstruction of Prediction and Decorrelation Coefficients Directly From DirAC Metadata for 1 Channel Downmix
[0215] From the above-mentioned normalized covariance matrix, SPAR coefficients can be computed based on computations in section 1.3 as follows.
[0216] The prediction coefficient is computed as^u^ / ! / " = (1 − wB) ∗ u@?Z^ / ! / "
[0049]
[0217] as^^ = ?¾MP((1 − wB) ∗ u@?Z^ + º^B − (1 − wB) ∗ u@?Z^ ),
[0050] ^! = ?¾MP((1 − wB) ∗ u@?Z! + º!B − (1 − wB) ∗ u@?Z! ),
[0051] ^" = ?¾MP((1 − wB) ∗ u@?Z" + º"B − (1 − wB) ∗ u@?Z" ).
[0052] Here, decorrelation coefficients depend on spherical harmonics response. To avoid this dependency, 2.2.4 can be used. 2.2.4 Another Variant of Generic Reconstruction of SPAR Coefficients From DirAC Metadata for Any Downmix Configuration
[0218] In this embodiment, a 4x4 covariance matrix, R, that is an approximation of actual input covariance u^^, is computed based on DirAC parameters as follows, where the elements of thematrix are approximated asu^© = (1 − wB) ∗ ^ ∗ u@?Z^ ∗ u@?Z© when = ! = ^, and u;; = ^ ∗ u@?Z; ∗ u@?Z; andu^^=(1 − wB)∗ ^ ∗ u@?Z^+ º^∗ E ∗(1 −(1 − wB) )when =! = ^ channel index,
[0053] where u@?Z;= T1,1(^ ), u@?Z^= T / , / (^ ), u@?Z!= T / ,^ / (^ ), u@?Z"= T / ,1(^ ) are the spherical harmonics, º^and w are constants in the range 0 and 1 (e.g., c = 1 and º^= 1 / 3), in which case both the encoder and the decoder would have prior knowledge about these constants. In some implementation, º^and w can be dynamically computed based on actual input covariance matrix and above mentioned approximation of input matrix from DirAC parameters.
[0219] Given that SPAR coefficients are normalized, the SPAR coefficients are derived from u similar to SPAR coefficients derived from ^ ∗ u, where E can be variance of just W channel or overall signal energy or any constant.
[0220] The elements of a normalized 4x4 covariance matrix for FoA channels are derived basedon DirAC parameters only:u^^^^^© = (1 − wB) / ∗ u@?Z^ ∗ u@?Z© ^ℎ@A = ! = ^, OA^ u^^^^;; = u@?Z; ∗ u@?Z; OA^are computed from u_AQMs as disclosed in section 1.3. 2.2.4.1 Example Reconstruction of Prediction and Decorrelation Coefficients Directly From DirAC Metadata for 1 Channel Downmix
[0222] From the above-mentioned normalized covariance, SPAR coefficients can be computed based on computations in section 1.3 as follows below.
[0223] The prediction coefficient can be computed as^u^ / ! / " = (1 − wB) ∗ u@?Z^ / ! / " .
[0055]
[0224] For aas^^= ?¾MP(º^(1 − (1 − wB) ),
[0056] Here,and only depend on diffuseness and some constants. Computation of Constant “c” – Solution 1
[0225] In an example implementation, to further improve the prediction coefficients, (1 − wB)can be set such that the passive W prediction coefficients are^u^ = ?¾MP(1 − B) ∗ u@?Z^, here i can be x, y , z
[0059]
[0226] Based on equation
[0053] , this will result in value of c asw = / ^ #À^^( / ^ Á)Á .
[0060]
[0227] In an embodiment, to improve SPAR coefficients that are computed from DirAC MD, a4x4 covariance matrix, R_norm, that is an approximation of actual normalized input covarianceu_AQMs^^, is computed based on DirAC parameters as follows, where the elements of the matrixare approximated as per
[0054] and
[0061] as given belowu^^^^^©= ?¾MP(1 − B) ∗ u@?Z^ ∗ u@?Z© when = ! = ^, OA^ u^^^^;;= u@?Z; ∗ u@?Z;OA^ u^^^^^^ = (1 − B) ∗ u@?Z^ + º^ ∗ B ^ℎ@A =! = ^ channel index.
[0061]
[0228] For a one channel downmix, based on Equations
[0062] -
[0064] , the decorrelation coefficientsare computed as^^ = ?¾MP(º^ ∗ B),
[0062] ^! = ?¾MP(º! ∗ B),
[0063] ^" = ?¾MP(º" ∗ B).
[0064]
[0229] In some embodiments, the values of º^, º!, º"can be set to 1 / 3. Computation of Constant “c” - Solution 2
[0230] In another example embodiment, w can be computed such that1− wB = ^^^^^bcd (__^^GG,^^^^^).
[0065]
[0231] + u_=A;" ), here,u_=A^© are the actual input covariance values and
[0232] Substituting this value of c in the prediction coefficient computation Equation^u^ =^^^^^ bcd (__^^GG,^^^^^) ∗ u@?Z^ , which is a close approximation ofbe x, y, z.
[0066]
[0233] This prediction coefficient
[0066] is similar to the passive prediction coefficient computation disclosed in section 1.3.1.1. For this solution the value of c can be transmitted to the decoder. 2.2.5 Energy compensation of DirAC based downmix
[0234] The covariance computation from DirAC metadata (MD) as disclosed in section 1.5.2 and in the solutions described in sections 2.2.3 and 2.2.4 assumes the signal to be perfectly SN3D normalized such that w = x+y+z,
[0067] where, w, x, y, z is the variance of W, X, Y, and Z channel respectively.
[0235] This assumption is not true in real life FoA captures, e.g., in overtalk situations, diffused background noise captures, etc. The above method results in spatial collapse especially when the number of downmix channels are limited to 1.
[0236] Energy compensation can be applied to prevent spatial collapse by scaling the downmix signal such that the upmixed signal is energy matched with respect to the input. Below is an example implementation of energy compensation with 1 channel downmix.
[0237] The actual input covariance matrix, u^^^^^, is computed, such that N is the number of input channels and, u^^^,©, is the frequency banded or broadband covariance of ith and jth input channel. For FOA input N = 4, and i and j can be W, X, Y, Z.
[0238] The normalized actual input covariance matrix u_AQMs_=A^^^, is computed asu_AQMs_=A^,© = u^^^,© / u^^;,;
[0068]
[0239] The DirAC metadata based normalized covariance estimate, u_AQMs^^^, is computed as per either of the techniques mentioned in sections 1.5.2, 2.2.3 and 2.2.4.
[0240] The scaling factor is obtained as?wO3@ = ?¾MP(PMOw@(u_AQMs_=A^^^) / max (@Z?, PMOw@(u_AQMs^^^))),
[0069] ?wO3@ = max (PℎM@?ℎ:^; , min (PℎM@?ℎÃ^ÄÃ, ?wO3@)).
[0070] Here, PℎM@?ℎ:^;OA^ PℎM@?ℎÃ^ÄÃare lower and upper bounds to the scale factor. In an example embodiment, PℎM@?ℎ:^;= 1 OA^ PℎM@?ℎÃ^ÄÃ= 2.
[0241] The SPAR downmix matrix and SPAR coefficients, including prediction, cross prediction and decorrelation coefficients are computed as disclosed in section 1.3, using the DirAC estimated normalized input covariance matrix.
[0242] Let the downmix matrix be <Q^As=^ / ^^. The downmix matrix is scaled by scale computed in equation
[0070] in section 2.2.5. Let the actual downmix matrix be <Q^As=^_OwP / ^^and<Q^As=^_OwP / ^^ = ?wO3@ ∗ <Q^As=^ / ^^.
[0071]
[0243] In an example embodiment, for 1 channel downmix, <Q^As=^ / ^^is given as follows asper Equation
[0072] ,<Q^As=^ / ^^ = [l{, l[, l", l^]
[0072] Here, l{, l[, l", l^are the gains that are used to mix Y, Z, and X channel, respectively, into W channel to form a downmix channel.
[0244] Post scaling with “scale” value, the downmix channel is computed as SU= ?wO3@ ∗ (l{∗ S + l[∗ T + l\∗ V + l]∗ W).
[0073]
[0245] In another example implementation, l{= 1, l[= l"= l^= 0, and SU= ?wO3@ ∗ S . Another example implementation with computation of l{, l[, l", l^is described in section 2.3.
[0246] The metadata parameters are unmodified with this scaling. The encoder encodes metadata parameters and the scaled downmix and the bitstream are transmitted to decoder.
[0247] The decoder decodes the scaled downmix channel W’ and spatial parameters including the prediction and decorrelation parameters, and applies the prediction and decorrelationparameters to reconstruct the original input scene such thatSÅÆÇ = S′(1 − ^#(ZMd + ZMÉ + ZMÊ )),
[0074] =+,VÅÆÇ = ZMÊS′ + ZÊ<³(S′).
[0077] Here, prx, pry and prz are prediction parameters, px, py, and pz are decorrelation parameters, D1(W’), D2(W’), D3(W’) are 3 decorrelated channels decorrelated with respect to W’, fs is active scaling as described in section 2.3. In an example implementation 0 ≤ ^#≤ 1.
[0248] This approach will scale the reconstructed signal by scale factor computed in equation
[0070] in this section, thereby energy matching the reconstructed scene with respect to the input without sending any additional parameter in the bitstream. 2.2.6 Extrapolating Directional Diffuseness in DirAC Bands
[0249] DirAC based covariance estimates assumes uniform diffuseness in all directions which may not be true with real life signals for e.g., overtalk scenarios. Adding directional information on top of the diffuseness parameter computed in Equation [7] in section 1.2 would result in additional metadata to be coded in the bitstream. SPAR does provide directional diffuseness information in its metadata and the directional information in high bands may be extrapolated using the directional information in lower bands.
[0250] In an example embodiment for FOA input with one channel downmix, if SPAR is coding up to the 6 kHz frequency range and the DirAC parameters are sent for 6-24 kHz frequency range,then the directional information in the SPAR frequency bands can be extracted as follows:^=M^^7+^^^|##^^^##,^ = bcd ^- ,
[0078] Here, px, py,
[0251] This directional information can be used in high frequency bands while computing downmix using DirAC parameters. An example estimation of normalized covariance matrix from DirAC metadata with directional diffuseness is as follows. R_norm is a 4x4 matrix for FOAchannels that is computed asu^^^^^© = (1 − wB) ∗ u@?Z^ ∗ u@?Z© when = ! = ^, and u^^^^;; = u@?Z; ∗ u@?Z;OA^ u^^^^^^ = (1 − wB) ∗ u@?Z^ + ^=M+^^^|##^^^##,^(1 − (1 − wB) ) when =! =^channel index
[0081] where, , u@?Z!= T / ,^ / (^ ), u@?Z"= T / ,1(^ ) are the spherical harmonics, w is a constant in the range 0 and 1 (e.g., c = 1), in which case both the encoder and the decoder would have prior knowledge about this constant. In some implementation, w can be dynamically computed based on actual input covariance matrix and above mentioned approximation of input matrix from DirAC parameters.
[0252] The downmix matrix and SPAR coefficients, including prediction, cross prediction and decorrelation coefficients, are computed from u_AQMs as disclosed in section 1.3. Example computation of Prediction coefficients and decorrelation coefficients for 1 channel downmix is given in
[0055] to
[0058] . Downmix matrix can be further scaled as per
[0070] to better energy match the reconstructed Ambisonics signal at the decoder with the Ambisonics signal at the encoder input. 2.2.7 DirAC to SPAR Metadata Conversion for HoA Channels 2.2.7.1 Estimating HoA Input Covariance Matrix from DirAC Parameters
[0253] In this method DirAC parameters are used to estimate the input covariance matrix.
[0254] The NxN covariance u is computed based on DirAC parameters, where N is number of input channels in HOA signal, here u is an approximation of actual input covariance matrix. Insome embodiments, the covariance, u, can be computed asu^©= (1 − wB) ∗ ^ ∗ u@?Z^∗ u@?Z©when = ! = ^, and u;;= ^ ∗ u@?Z;∗ u@?Z;andu^^ = (1 − wB) ∗ ^ ∗ u@?Z^ + º^ ∗ E ∗ (1 − (1 − wB) ) when =! = ^ channel index.Here, u@?Z^are the spherical harmonics, º^and w are constants in the range 0 and 1, e.g., c = 1 and º^= 1 / 3 for first order channels, i.e., 0 <= i <= 3, º^= 1 / 5 for second order channels , i.e., 4<= i <= 8, º^= 1 / 7 for third order channels, i.e., 9<= i <= 15, in which case both the encoder and the decoder have prior knowledge about these constants. In some embodiments, º^and w are dynamically computed based on actual input covariance matrix and above mentioned approximation of input matrix from DirAC parameters.
[0255] Given that SPAR coefficients are normalized, the SPAR coefficients derived from, u are equal to SPAR coefficients derived from ^ ∗ u, where E can be variance of just W channel or overall signal energy or any constant.
[0256] The covariance R in equation
[0081] is normalized and the elements of this NxN normalized covariance matrix, u_AQMs, are derived based on DirAC parameters only as follows:u^^^^^© = (1 − wB) / ∗ u@?Z^ ∗ u@?Z© when = ! = ^, and u^^^^;; = u@?Z; ∗ u@?Z;OA^ u_AQMs^^ = (1 − wB) ∗ u@?Z^ + º^ (1 − (1 − wB) ) when =! =^ channel index.
[0083]
[0257] SPAR coefficients, including prediction, cross prediction and decorrelation coefficients, are computed from u_AQMs as disclosed in section 1.3. 2.2.7.2 Improving Spatial Resolution of DirAC to SPAR Conversion by Limiting DirAC Covariance Estimation to FOA Channels Only
[0258] It has been observed that covariance estimation for HOA channels from DirAC parameters is not optimal when there is critical information in HOA channels. Loss of ambiance has been observed when estimating the entire NxN covariance matrix (or all HOA SPAR parameters) from the DirAC parameters. A separate approach is desired for such HOA signals. Below mentioned are few embodiments for DirAC to SPAR conversion with improved spatial resolution. 2.2.7.2.1 By computing and coding SPAR HoA parameters independently
[0259] In this method DirAC parameters are used to estimate input covariance matrix for only FOA channels and then from that estimate SPAR parameters corresponding to FOA channels are computed. This is done by methods described in section 2.2. and 2.2.4.
[0260] SPAR parameters including prediction coefficients, cross-prediction coefficients and decorrelation coefficients for HoA channels are computed independently based on actual covariance matrix of the input signal based on methods described in section 1.3.
[0261] This method will require coding of SPAR HoA parameters into bitstream for all frequencies. 2.2.7.2.2 Alternate Computation of SPAR HOA Parameters Based on DirAC Estimated FOA
[0262] This method is applicable to SPAR modes where number of downmix channels are less than number of input channels to SPAR, that is cases where SPAR has cross-prediction and / or decorrelation coefficients to code for HOA channels. In this method DirAC parameters are used to estimate the input covariance matrix for only FOA channels and then from that SPAR parameters corresponding to FOA channels are computed. This is done by methods described in sections 2.2. and 2.2.4. Computation of HOA Prediction Coefficients
[0263] SPAR prediction coefficients for HOA channels are computed independently based on the actual covariance matrix of the input signal based on methods described in section 1.3.Computation of HOA Cross-Prediction Coefficients
[0264] Section 1.3 shows that cross-prediction coefficients in SPAR MD depend on predicted side channels or residual channels in the downmix. Furthermore, the residual channels in FOA component of the Ambisonics input depends on SPAR MD that is derived from DirAC MD in a set of frequency bands. Hence, cross-prediction coefficients in HOA channels can be dependent on DirAC MD in FOA channels and it has been observed that computing cross-prediction coefficients in HOA channels based on DirAC MD in FOA channels and SPAR MD in FOA and HOA channels can lead to a better estimate of these coefficients. In an example implementation, HOA channels (4 to N) prediction coefficients are computed from an actual input covariance matrix as described in section 1.3. These prediction coefficients are quantized based on a quantization strategy. DirAC estimated FoA prediction coefficients along with SPAR estimated HOA quantized prediction coefficients are used to generate the downmix matrix as described in section 1.3. A post prediction covariance matrix is computed from the actual input covariance and downmix matrix computed above. Cross-prediction coefficients are then computed from post prediction matrix as described in section 1.3. Computation of HOA Decorrelation Coefficients
[0265] It has been observed that computing HOA decorrelation coefficients directly from Ambisonics input covariance as described in section 1.3 without having any dependency on DirAC MD in FOA channels leads to better estimation of decorrelation coefficients and results in desired amount of decorrelation in the recontructed HOA channels at the decoder. This is helpful in reducing the audio artifacts that can arise due to too much decorrelation and also avoids spatial collapse due to too less decorrelation. In an example implementation, first, the prediction coefficients corresponding to all side channels are computed from actual input covariance matrix as described in section 1.3, where side channels in Ambisonics are all channels input except the W channel. Then, the computation of decorrelation coefficients from the prediction coefficients and the covariance matrix is the same as described in section 1.3. This method will code the SPAR HoA parameters into the bitstream for all frequencies. 2.3 Active W Downmix Based on DirAC Metadata 2.3.1 Based on DirAC Based Covariance Estimation
[0266] From DirAC metadata, the input covariance may be estimated as a DirAC metadata-based input signal (4 x 4) covariance matrix estimation as given in section 2.2.3 or 2.2.4: ^(1 − wB)^>Ï∗where >Ï is2.2.3, and Sis a 3x3 matrix where the elements of the matrix are given by^^© = (1 − wB) ∗ ^ ∗ u@?Z^ ∗ u@?Z© when = ! = ^, and^^^ = (1 − wB) ∗ ^ ∗ u@?Z^ + º^ ∗ B ∗ E
[0085] Alternatively, ^ can be computed as given in section 2.2.4 as^^© = (1 − wB) ∗ ^ ∗ u@?Z^ ∗ u@?Z© when = ! = ^, and^^^ = (1 − wB) ∗ ^ ∗ u@?Z^ + º^ ∗ E ∗ (1 − (1 − wB) ).
[0086]
[0267] One possible approach to performing active downmix based on above covariance matrix is by having following prediction matrix [Ï∗^M@^ Í^Í] = Ò 1 l>−N>Ï ^³ − Nl>Ï>Ï∗Ó , ^ℎ@M@ l = (1 − wψ),>Ï is [3x1] unit vector u@?Z^, u@?Z!, u@?Z".
[0087] Then post prediction matrix can be given as^Q?P_ZM@^=wP=QA[Í^Í]= ^M@^ ∗ u[Í^Í]∗ ^M@^U,
[0088] ⋯ Other to the activedownmixing gains computation.
[0268] Minimizing M̂ by setting >Ï∗ ∗ M̂, results in a linear equation given by3=A@OM(N) = N×l + 2ØNl + ^N − ×l − Ø, N = Ù8 ÚÛÚÛ78 ÙÛ 8 H
[0090] Here, × = >Ï∗in
[0090] , ^ cancels out in the denominator and numerator and g can be computed directly from DirAC metadata on both the encoder and decoder side.
[0269] Actual downmix matrix for 1 channel downmix, post scaling, is given as^M@^[ / ^Í] = (M Ml>Ï∗), where M is a scaling factor.
[0091] The computation of the post prediction scaling factor “r” is done by matching the reconstructed W variance at decoder with the variance of W encoder input, ÄG8^Ä 78Í^ Ä7M =G E,
[0092] and, ^#, is a
[0270] The scaled prediction coefficients are computed as followsNU = Ä^
[0093] Here, N′>Ï = [ZMd; ZMÉ; ZMÊ] are the active prediction coefficients.
[0271] Computation of decorrelation coefficients is as follows^Q?P_ZM@^=wP=QA = ^M@^ U[Í^Í] ∗ u[Í^Í] ∗ ^M@^Here, ^M@^ is the prediction matrix given in
[0091] , decorrelation coefficients are computed from^Q?P_ZM@^=wP=QA[Í^Í] as follows^u@?>> = _^#||^^^*^,^^#^_v^^+^^^^^^G,G,^^(|_^#|||).,Here, Resuu pz] are the decorrelation
[0272] Computation of active W downmix channel from FOA input [W, Y, Z, X] is given asSU= ?wO3@ ∗ M ∗ (S + l[ ∗ T + l\ ∗ V + l] ∗ W)l[ = l*u@?Z!l] = l*u@?Z^l\ = l*u@?Z"
[0273] Computation of ?wO3@ is given in
[0070] , computation of another scale factor M is given in
[0092] , SUis encoded with a core coder, DirAC MD is coded and together these coded bits are sent to decoder
[0274] The inverse prediction matrix at the decoder is given as follows: ^A^^M@^[Í^Í]= Ò(1 − ^#NU ) 0 Ä^ Ó , NU=^
[0095]
[0275] SÅÆÇ = S′(1 − ^#(ZMd + ZMÉ + ZMÊ )),
[0096] Here, prx, pry and prz are prediction parameters that are computed from DirAC MD as given in
[0090] , px, py, and pz are decorrelation parameters that are computed from DirAC MD as given in
[0094] , D1(W’), D2(W’), D3(W’) are 3 decorrelated channels decorrelated with respect to W’, fs is the scaling constant used in
[0092] . 2.4 SPAR to DirAC Metadata Conversion
[0276] It may be desired to convert SPAR MD to DirAC MD in a set of frequency bands such that DirAC MD is available at all required frequency bands in order to perform upmix to desired output format at the decoder. Direct conversion from SPAR MD to DirAC MD also savescomplexity. In an example implementation, It is possible to derive directional vector ^^ fromprediction coefficients^^# = ^_E .
[0100] 6(^_^78 ^_-78 ^_57)Here, ? wOA Þ@ ^, ^, ^ azimuth and elevation can then be computed based on equations [5] and [6]. 2.4.1 Diffuseness Computation From SPAR Metadata
[0277] Assuming that SPAR perfectly reconstructs the covariance (COV) matrix, the output covariance matrix can be computed at the decoder from the input (DMX + decorrelators) covariance and upmix matrix. From the output COV, the reference power and intensity are computed and averaged over N frames (e.g., 8 frames). From that, diffuseness is computed as per Equation [7].
[0278] There are other embodiments to directly compute DirAC diffuseness from SPAR metadata without computing output covariance matrix as disclosed below. 2.4.1.1 Alternate Method for Diffuseness for 1 Channel Downmix With Passive W Downmix (where W channel in downmix is same or just delayed version of W channel in input)
[0279] Let the variance of W, X, Y, Z be w, x, y, z. In a 1 channel downmix case, ^ can beapproximated as^= ^(ZM! + Z^! )
[0101] where, pry is the prediction coefficient and pdy is the decorrelation coefficient for the Y channel. Similarly, x and z can be calculated for the X and Z channels. The reference power E can then becomputed as (w + x + y + z),^= ^(1 + ZM! + Z^! + ZM^ + Z^^ + ZM" + Z^" ).
[0102]
[0280] Intensity can be computed as^! = ^ ∗ ZM!, ^^ = ^ ∗ ZM^ , ^" = ^ ∗ ZM".
[0103]
[0281] Referring to Equation [7], diffuseness B may be approximated directly from SPAR metadata as follows: 6*v^7EF^G,^ 8 v^7 7EF^G,- 8 v^EF^G,5.Here, ZM#:^;,#same as Z^#or it could be a long time average of Z^#, Here, ? wOA Þ@ ^, ^, ^. 2.4.1.2 Alternate Method for Diffuseness for 1 Channel Downmix With Active W Downmix
[0282] Given the inverse matrix with active W computation as mentioned in section 2.3.1.^A^^M@^[Í^Í] = Ò(1 − ^ U#N ) 0N′>Ï ^ Ó , NU=Ä[10³ ^ 5]
[0283] Let case, y can beapproximated^= ^(ZM! + Z^! ).
[0106] Here, pry is the prediction coefficient and pdy is the decorrelation coefficient for the Y channel. Similarly, x and z can be calculated as well.
[0284] The reference power can then be computed as (w + x + y + z),^= ^((1 − ^#NU ) + ZM! + Z^! + ZM^ + Z^^ + ZM" + Z^" ).
[0107]
[0285] Intensity can be computed as^! = ^ ∗ ZM!, ^^ = ^ ∗ ZM^ , ^" = ^ ∗ ZM".
[0108]
[0286] Referring to Equation [7], diffuseness B may be approximated directly from SPAR metadata as (if we average w separately), 6*v^7 8 v^7 8 v^7EF^G,^ EF^G,- .B = 1 −EF^G,5. Here, issame as Z^#or it could be a long time average of Z^#, Here, ? wOA Þ@ ^, ^, ^. 2.4.1.3 Alternate Method for Diffuseness for Any Passive W Downmix Channel Configuration
[0287] This method is based on the normalization of the input Ambisonics signal. For example, if the FoA input is normalized using Schmidt semi-normalization (SN3D), then it assumes that w = x+y+z, where w, x, y, z is the variance of W, X, Y, and Z channels, respectively. This makes w+x+y+z = 2*w.
[0288] Substituting the variance assumption and intensity from equation
[0108] in section 2.4.1.1. into the diffuseness formula in Equation [7] gives, ;6*v^7 8 v^7 8 7EF^G,^ EF^G v^ .<=^^>?@A@?? = ψ = 1 −,- EF^G,51.¶∗( ∗;),
[0110] <=^^>?@A@?? = ψ = 1 − 6^ZM#:^;,^+ ZM#:^;,!+ ZM#:^;,"^.
[0111] Here, ZM#:^;,#is either same as ZM#or it could be a long time average of ZM#, here, ? wOA Þ@ ^, ^, ^.
[0289] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, thecomputer program may be downloaded and mounted from the network via the communication unit 709, and / or installed from the removable medium 711, as shown in FIG.7.
[0290] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 701 in combination with other components of FIG.7), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non- limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0291] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.
[0292] In the context of the disclosure, a machine readable medium may be any tangible medium that may contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine readable medium may be a machine readable signal medium or a machine readable storage medium. A machine readable medium may be non- transitory and may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0293] Computer program code for carrying out methods of the present disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general purpose computer, special purpose computer, or otherprogrammable data processing apparatus that has control circuitry, such that the program codes, when executed by the processor of the computer or other programmable data processing apparatus, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a computer, partly on the computer, as a stand-alone software package, partly on the computer and partly on a remote computer or entirely on the remote computer or server or distributed over one or more remote computers and / or servers.
[0294] While this document contains many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can, in some cases, be excised from the combination, and the claimed combination may be directed to a sub combination or variation of a sub combination. Logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.
Claims
CLAIMS What is claimed is:
1. A method for encoding an audio signal, comprising: obtaining the audio signal, wherein the audio signal represents an input audio scene with a primary channel and side channels; analyzing a power of the primary channel of the audio signal; analyzing powers of the side channels of the audio signal; detecting a mono mode for encoding the audio signal based on analyzing the power of the primary channel and the powers of the side channels; computing one or more downmix channels and spatial metadata from the audio signal for a detected mono mode; encoding the one or more downmix channels and spatial metadata in a bitstream for the detected mono mode; and indicating the mono mode in the bitstream.
2. The method of claim 1, further comprising outputting the bitstream, storing the bitstream, transmitting the bitstream, or combinations thereof.
3. The method of claim 1 or claim 2, wherein detecting the mono mode is based on a determination that input audio signal has non-silent audio in the primary channel and that the side channels are silent.
4. The method of claim 1 or claim 2, wherein detecting the mono mode comprises: computing a primary channel power of the primary channel of the audio signal; and computing a sum power as a summation of the powers of the side channels of the audio signal.
5. The method of claim 4, wherein detecting the mono mode further comprises evaluating a ratio of the primary channel power to the sum power.
6. The method of claim 5, wherein detecting the mono mode further comprises determining when the ratio exceeds a threshold.
7. The method of any one of claims 1–6, further comprising implicitly signaling the mono mode by setting one or more parameters of the spatial metadata to a value of zero.
8. The method of claim 6, wherein the one or more parameters of the spatial metadata include one or more Spatial Reconstruction (SPAR) metadata parameters, one or more Directional Audio Coding (DirAC) metadata parameters, or both.
9. The method of any one of claims 1–8, further comprising explicitly signaling the mono mode by setting a mono flag.
10. The method of any one of claim 1–9, wherein the bitstream is an Immersive Voice and Audio (IVAS) encoded bitstream.
11. A method for audio signal decoding, comprising: obtaining an encoded bitstream; decoding the encoded bitstream and obtaining downmix channels, spatial metadata, and a mono mode indicator; setting one or more parameters of spatial metadata to zero upon detecting a mono mode; upmixing the downmix channels using the spatial metadata; and rendering upmixed channels to a desired audio format.
12. The method of claim 11, wherein the rendering produces rendered audio data, further comprising transmitting the rendered audio data, storing the rendered audio data, transmitting the rendered audio data, or combinations thereof.
13. The method of claim 11, wherein the rendering produces rendered audio data, further comprising providing the rendered audio data to one or more loudspeakers for playback.
14. The method of any one of claims 11–13, wherein the mono mode indicator is based on values of one or more spatial metadata parameters in the encoded bitstream that are set to values to indicate the mono mode.
15. The method of claim 14, wherein the one or more spatial metadata parameters include one or more Spatial Reconstruction (SPAR) metadata parameters, one or more Directional Audio Coding (DirAC) metadata parameters, or both.
16. The method of any one of claims 11–15, wherein setting the one or more parameters of spatial metadata to zero upon detecting the mono mode involves setting one or more energy ratio values to zero.
17. The method of claim 16, wherein a received energy ratio metadata value is non-zero due to an artifact of a quantization process.
18. The method of claim 17, further comprising setting one or more diffuseness values to one.
19. The method of any one of claims 11–18, wherein the encoded bitstream is an Immersive Voice and Audio (IVAS) encoded bitstream.
20. A method of encoding an audio signal representing an input audio scene with a primary channel and side channels, the method comprising: obtaining a frame of the audio signal from a bitstream; identifying a plurality of frequency bands associated with the frame; determining power associated with the primary channel and the side channels for each of the plurality of frequency bands associated with the frame; classifying each of the plurality of frequency bands as one of Silence, Mono, or Ambisonics based on a determined power for a corresponding frequency band; marking a mode of the frame as one of Silence, Mono, or Ambisonics based on a plurality of classified frequency bands for the frame; and encoding a marked mode of the frame in the bitstream.
21. The method of claim 20, further comprising outputting the bitstream, storing the bitstream, transmitting the bitstream, or combinations thereof.
22. The method of claim 20 or claim 21, wherein marking the frame further comprises: marking the frame as Ambisonics when one or more classified frequency bands corresponds to Ambisonics.
23. The method of claim 20 or claim 21, wherein marking the frame further comprises: marking the frame as Mono when no classified frequency band corresponds to Ambisonics, and when one or more classified frequency bands corresponds to Mono.
24. The method of any one of claims 20–23, wherein marking the frame further comprises: marking the frame as Silence when no classified frequency band corresponds to Ambisonics, and when no classified frequency band corresponds to Mono.
25. The method of any one of claims 20–24, wherein determining power of a frequency band further comprises: calculating a primary channel power of the frequency band; and calculating a sum power of the frequency band, wherein the sum power is a summation of power to the side channels.
26. The method of claim 25, wherein classifying each of the plurality of frequency bands further comprises: evaluating a ratio of the primary channel power to the sum power.
27. The method of claim 26, wherein detecting a mono mode comprises determining when the ratio exceeds a threshold.
28. The method of claim 27, wherein the threshold is a linear function of the primary channel power, the primary channel power is normalized to a value between 0 and 1, and the threshold is clipped between maximum and minimum values.
29. The method of claim 25, wherein classifying each of the plurality of frequency bands further comprises: determining that the sum power is below a minimum noise level; and declaring the frequency band as Mono when the primary channel power is above a noise threshold.
30. The method of claim 25, wherein classifying each of the plurality of frequency bands further comprises: determining that the sum power is above a minimum noise level; and declaring the frequency band as Mono when a normalized primary channel power is above a noise threshold.
31. The method of any one of claims 20–30, wherein the bitstream is an Immersive Voice and Audio (IVAS) encoded bitstream.
32. The method of any one of claims 20–31, wherein marking the mode of the frame involves setting one or more Spatial Reconstruction (SPAR) metadata parameters to zero, setting one or more Directional Audio Coding (DirAC) metadata parameters to zero, or both.
33. An apparatus configured to perform the method of any one of claims 1–32.
34. A system configured to perform the method of any one of claims 1–32.
35. One or more non-transitory, computer-readable media having instructions encoded thereon for controlling one or more devices to perform the method of any one of claims 1–32.
36. An apparatus, comprising: an input / output (I / O) system; and one or more processors configured to:obtain, via the I / O system, an audio signal, wherein the audio signal represents an input audio scene with a primary channel and side channels; analyze a power of the primary channel of the audio signal; analyze powers of the side channels of the audio signal; detect a mono mode for encoding the audio signal based on analyzing the power of the primary channel and the powers of the side channels; compute one or more downmix channels and spatial metadata from the audio signal for a detected mono mode; encode the one or more downmix channels and spatial metadata in a bitstream for the detected mono mode; and indicate the mono mode in the bitstream.
37. The apparatus of claim 36, wherein the one or more processors are further configured for outputting the bitstream, storing the bitstream, transmitting the bitstream, or combinations thereof.
38. The apparatus of claim 36 or claim 37, wherein marking the frame further comprises marking the frame as Ambisonics when one or more classified frequency bands corresponds to Ambisonics.
39. The apparatus of any one of claims 36–38, wherein marking the frame further comprises marking the frame as Mono when no classified frequency band corresponds to Ambisonics, and when one or more classified frequency bands corresponds to Mono.
40. An apparatus, comprising: an input / output (I / O) system; and one or more processors configured to: obtain, via the I / O system, a frame of an audio signal from a bitstream; identify a plurality of frequency bands associated with the frame; determine power associated with the primary channel and the side channels for each of the plurality of frequency bands associated with the frame; classify each of the plurality of frequency bands as one of Silence, Mono, or Ambisonics based on a determined power for a corresponding frequency band; mark a mode of the frame as one of Silence, Mono, or Ambisonics based on a plurality of classified frequency bands for the frame; and encode a marked mode of the frame in the bitstream.
41. The apparatus of claim 40, wherein the one or more processors are further configured for outputting the bitstream, storing the bitstream, transmitting the bitstream, or combinations thereof.
42. The apparatus of claim 40 or claim 41, wherein marking the frame further comprises marking the frame as Ambisonics when one or more classified frequency bands corresponds to Ambisonics.
43. The apparatus of any one of claims 40–42, wherein marking the frame further comprises marking the frame as Mono when no classified frequency band corresponds to Ambisonics, and when one or more classified frequency bands corresponds to Mono.
44. An apparatus, comprising: an input / output (I / O) system; and one or more processors configured to: obtain, via the I / O system, an encoded bitstream; decode the encoded bitstream and obtaining downmix channels, spatial metadata, and a mono mode indicator; set one or more parameters of spatial metadata to zero upon detecting a mono mode; upmix the downmix channels using the spatial metadata; and render upmixed channels to a desired audio format.
45. The apparatus of claim 44, wherein the rendering produces rendered audio data and wherein the one or more processors are further configured for transmitting the rendered audio data, storing the rendered audio data, transmitting the rendered audio data, or combinations thereof.
46. The apparatus of claim 44 or claim 45, wherein the rendering produces rendered audio data and wherein the one or more processors are further configured for providing the rendered audio data to one or more loudspeakers for playback.