Method, apparatus, and system for scene-based audio mono decoding

A mono detector in the encoder corrects format misinterpretation in spatial audio coding by identifying and signaling mono modes, ensuring accurate decoding and rendering of audio signals, addressing the challenges faced by SPAR and DirAC in handling format transitions.

JP2026524629APending Publication Date: 2026-07-23DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2024-07-03
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing spatial audio coding techniques such as SPAR and DirAC face challenges in efficiently handling audio signals that transition from ambisonic to mono format, leading to undesirable audio characteristics due to format misinterpretation by decoders.

Method used

Implementing a mono detector in the encoder to identify and signal mono modes through spatial metadata parameters, ensuring the W channel is treated as mono in the decoder, and using frequency band analysis to classify frames as silent, mono, or ambisonic, thereby correcting format interpretation.

Benefits of technology

Ensures proper decoding and rendering of audio signals, maintaining audio quality by accurately distinguishing between mono and ambisonic formats, preventing misinterpretation and enhancing decoder performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026524629000001_ABST
    Figure 2026524629000001_ABST
Patent Text Reader

Abstract

This specification discloses methods for encoding and decoding audio signals. Some disclosed methods for encoding audio signals include obtaining an audio signal representing an input audio scene with a primary channel and side channels, analyzing the power of the primary channel, and analyzing the power of the side channels. Some such methods include detecting a mono mode for encoding the audio signal based on the analysis of the primary channel power and side channel power, and calculating one or more downmix channels and spatial metadata from the audio signal for the detected mono mode. Some such methods include encoding the one or more downmix channels and spatial metadata in a bitstream for the detected mono mode, and representing the mono mode in a bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-references to related applications This application claims priority from U.S. Provisional Patent Application No. 63 / 570,117 filed on 26 March 2024, U.S. Provisional Patent Application No. 63 / 593,278 filed on 26 October 2023, and U.S. Provisional Patent Application No. 63 / 511,786 filed on 3 July 2023, each of which is incorporated herein by reference in whole.

[0002] Technical field This disclosure relates, in general terms, to audio processing. [Background technology]

[0003] Unless otherwise stated herein, the materials described in this section are not prior art to the claims of this application, nor are they recognized as prior art by their inclusion in this section.

[0004] Spatial Reconstruction (SPAR) and Directional Audio Coding (DirAC) are distinct spatial audio coding techniques, each attempting to represent an input spatial audio scene in a compact way that allows for a good trade-off between audio quality and bitrate. One such input format for a spatial audio scene is ambisonic representation (e.g., first-order ambisonics (FOA) or higher-order ambisonics (HOA)).

[0005] SPAR attempts to maximize perceived audio quality while minimizing bitrate by reducing the energy of transmitted audio data, while allowing secondary statistics (e.g., covariance) of the ambisonic audio scene to be reconstructed on the decoder side using transmitted metadata. SPAR attempts to faithfully reconstruct the input ambisonic scene at the decoder output.

[0006] DirAC is a technique that represents spatial audio scenes as a collection of direction of arrival (DOA) values ​​in time-frequency tiles. This representation allows similar-sounding scenes to be reproduced in different output formats (such as binaural). In particular, in the context of ambisonics, the DirAC representation allows the decoder to generate higher-order outputs from lower-order inputs (blind upmixing). DirAC attempts to maintain the dominant sonic direction and diffuseness in the input scene.

[0007] Both DirAC and SPAR have different strengths and characteristics. Therefore, it is desirable to combine the complementary aspects of DirAC and SPAR (e.g., higher audio quality, reduced bitrate, input / output format flexibility, and / or reduced computational complexity) with a coder / decoder ("codec") such as an Ambisonics codec. [Overview of the project] [Problems that the invention aims to solve]

[0008] Techniques for processing audio signals are described. The examples described herein provide systems, devices, and methods for encoding and / or decoding bitstreams, where frames are marked with mode indicators used when rendering audio. [Means for solving the problem]

[0009] Methods for encoding and decoding audio signals are disclosed herein. Some disclosed methods for encoding audio signals involve obtaining an audio signal representing an input audio scene with a primary channel and side channels, analyzing the power of the primary channel, and analyzing the power of the side channels. Some such methods involve detecting mono modes for encoding the audio signal based on the analysis of the power of the primary channel and side channels, and for the detected mono modes, calculating one or more downmix channels and spatial metadata from the audio signal. Some such methods involve encoding the one or more downmix channels and spatial metadata in a bitstream for the detected mono modes, and indicating the mono modes in the bitstream.

[0010] Several exemplary embodiments describe a method for encoding an audio signal. In some cases, the audio signal may represent an input audio scene with a primary channel and side channels. In some exemplary embodiments, the method may involve acquiring the audio signal. According to some exemplary embodiments, the method may involve analyzing the power of the primary channel of the audio signal and analyzing the power of the side channels of the audio signal. In some exemplary embodiments, the method may involve detecting mono modes for encoding the audio signal based on the analysis of the primary channel power and side channel power. According to some exemplary embodiments, the method may involve calculating one or more downmix channels and spatial metadata from the audio signal for the detected mono modes, and encoding the one or more downmix channels and spatial metadata in a bitstream for the detected mono modes. In some exemplary embodiments, the method may involve representing the mono modes in the bitstream.

[0011] In some exemplary embodiments, the method may involve outputting a bitstream, storing a bitstream, transmitting a bitstream, or a combination thereof.

[0012] According to some exemplary embodiments, mono mode detection may be based on the determination that the input audio signal has non-silent audio in the primary channel and that the side channel is silent.

[0013] In some exemplary embodiments, mono-mode detection may involve calculating the primary channel power of the primary channel of the audio signal and calculating the total power as the sum of the powers of the side channels of the audio signal. In some such exemplary embodiments, mono-mode detection may also involve evaluating the ratio of the primary channel power to the total power. In some such exemplary embodiments, mono-mode detection may also involve determining when the ratio exceeds a threshold.

[0014] According to some exemplary embodiments, the method may involve implicitly signaling mono-mode by setting one or more parameters of spatial metadata to a value of 0. In some such exemplary embodiments, the one or more parameters of spatial metadata may include one or more spatial reconstruction (SPAR) metadata parameters, one or more directional audio coding (DirAC) metadata parameters, or both.

[0015] In some exemplary embodiments, the method may involve explicitly signaling to mono mode by setting a mono flag.

[0016] According to some exemplary embodiments, the bitstream may be an Immersive Voice and Audios (IVAS) encoded bitstream.

[0017] Several additional exemplary embodiments describe methods for decoding audio signals. In some exemplary embodiments, the method may involve obtaining an encoded bitstream and decoding the encoded bitstream to obtain downmix channels, spatial metadata, and a mono mode indicator. According to some exemplary embodiments, when mono mode is detected, the method may involve setting one or more parameters of the spatial metadata to 0. In some exemplary embodiments, the method may involve upmixing the downmix channels using the spatial metadata. According to some exemplary embodiments, the method may involve rendering the upmixed channels to a desired audio format.

[0018] In some exemplary embodiments, rendering generates rendered audio data. According to some exemplary embodiments, the method may also involve transmitting the rendered audio data, storing the rendered audio data, transmitting the rendered audio data, or a combination thereof. Some exemplary embodiments may also involve providing the rendered audio data to one or more loudspeakers for playback.

[0019] According to some exemplary embodiments, the mono-mode indicator may be based on the value of one or more spatial metadata parameters in the encoded bitstream, which are set to a value indicating mono-mode. According to some such exemplary embodiments, the one or more spatial metadata parameters may relate to one or more SPAR metadata parameters, one or more DirAC metadata parameters, or both.

[0020] In some exemplary embodiments, the method may involve setting one or more parameters of the spatial metadata to 0 when it detects that a mono-mode is involved in setting one or more energy ratio values ​​to 0. In some such exemplary embodiments, the received energy ratio metadata values ​​may be non-zero due to artifacts of the quantization process. In some exemplary embodiments, the method may also involve setting one or more diffusivity values ​​to 1.

[0021] According to some exemplary embodiments, the encoded bitstream may be an IVAS-encoded bitstream.

[0022] Further exemplary embodiments describe a method for encoding an audio signal that represents an input audio scene with a primary channel and side channels. In some exemplary embodiments, the method may involve obtaining a frame of the audio signal from a bitstream and identifying a plurality of frequency bands associated with the frame. According to some exemplary embodiments, the method may involve determining the power associated with the primary channel and side channels for each of the plurality of frequency bands associated with the frame. In some exemplary embodiments, the method may involve classifying each of the plurality of frequency bands as one of silence, mono, or ambisonics based on the determined power for the corresponding frequency band. According to some exemplary embodiments, the method may involve marking the mode of the frame as one of silence, mono, or ambisonics based on the plurality of classified frequency bands for the frame, and encoding the marked mode of the frame in the bitstream.

[0023] In some exemplary embodiments, the method may also involve outputting a bitstream, storing a bitstream, transmitting a bitstream, or a combination thereof.

[0024] According to some exemplary embodiments, marking a frame may involve marking a frame as ambisonics if one or more classified frequency bands correspond to ambisonics.

[0025] According to some exemplary embodiments, the method may involve marking a frame as mono if none of the classified frequency bands correspond to ambisonics and one or more classified frequency bands correspond to mono. In some exemplary embodiments, the method may involve marking a frame as silent if none of the classified frequency bands correspond to ambisonics and none of the classified frequency bands correspond to mono.

[0026] In some exemplary embodiments, determining the power of a frequency band may also involve calculating the primary channel power of the frequency band and the total power of the frequency band, where the total power is the sum of the powers for each side channel. In some such exemplary embodiments, classifying each of a plurality of frequency bands may also involve evaluating the ratio of the primary channel power to the total power. In some such exemplary embodiments, detecting a mono-mode may involve determining when the ratio exceeds a threshold. In some such exemplary embodiments, the threshold may be a linear function of the primary channel power, the primary channel power may be normalized to a value between 0 and 1, and the threshold may be clipped between a maximum and a minimum value.

[0027] According to some exemplary embodiments, classifying each of multiple frequency bands may involve determining that the sum power is below a minimum noise level and declaring the frequency band as mono when the primary channel power exceeds a noise threshold.

[0028] In some exemplary embodiments, classifying each of the multiple frequency bands may involve determining whether the sum power exceeds the minimum noise level and declaring the frequency band as mono when the normalized primary channel power exceeds a noise threshold.

[0029] According to some exemplary embodiments, the bitstream may be an IVAS-encoded bitstream.

[0030] In some exemplary embodiments, marking the mode of a frame may involve setting one or more SPAR metadata parameters to 0, setting one or more DirAC metadata parameters to 0, or both.

[0031] Further exemplary embodiments describe the apparatus. In some exemplary embodiments, the apparatus may include an input / output (I / O) system and one or more processors. According to some exemplary embodiments, one or more processors may be configured to acquire an audio signal representing an input audio scene with primary and side channels via the I / O system. In some exemplary embodiments, one or more processors may be configured to analyze the power of the primary channel of the audio signal and to analyze the power of the side channels of the audio signal. According to some exemplary embodiments, one or more processors may be configured to detect a mono mode for encoding the audio signal based on the analysis of the primary channel power and the side channel power. In some exemplary embodiments, one or more processors may be configured to calculate one or more downmix channels and spatial metadata from the audio signal for the detected mono mode. According to some exemplary embodiments, one or more processors may be configured to encode one or more downmix channels and spatial metadata in the bitstream for the detected mono mode to indicate the mono mode in the bitstream.

[0032] According to some exemplary embodiments, one or more processors may be configured to output a bitstream, store a bitstream, transmit a bitstream, or a combination thereof.

[0033] In some exemplary embodiments, marking a frame may include marking the frame as ambisonics if one or more classified frequency bands correspond to ambisonics.

[0034] According to some exemplary embodiments, marking a frame may involve marking a frame as mono if the classified frequency bands do not correspond to ambisonics, and if one or more classified frequency bands correspond to mono.

[0035] Further exemplary embodiments describe a device comprising an I / O system and one or more processors. In some exemplary embodiments, one or more processors may be configured to acquire frames of audio signals from a bitstream via the I / O system. According to some exemplary embodiments, one or more processors may be configured to identify a plurality of frequency bands associated with a frame. In some exemplary embodiments, one or more processors may be configured to determine the power associated with the primary channel and side channels for each of the plurality of frequency bands associated with a frame. According to some exemplary embodiments, one or more processors may be configured to classify each of the plurality of frequency bands as one of silence, mono, or ambisonics based on the determined power for the corresponding frequency band. In some exemplary embodiments, one or more processors may be configured to mark the mode of a frame as one of silence, mono, or ambisonics based on the plurality of classified frequency bands for the frame, and to encode the marked mode of the frame in the bitstream.

[0036] According to some exemplary embodiments, one or more processors may be configured to output a bitstream, store a bitstream, transmit a bitstream, or a combination thereof.

[0037] In some exemplary embodiments, marking a frame may involve marking the frame as ambisonic if one or more classified frequency bands correspond to ambisonics. In some exemplary embodiments, marking a frame may involve marking the frame as mono if none of the classified frequency bands correspond to ambisonics, and one or more classified frequency bands correspond to mono.

[0038] Further exemplary embodiments describe a device including an I / O system and one or more processors. In some exemplary embodiments, one or more processors may be configured to receive an encoded bitstream via the I / O system. According to some exemplary embodiments, one or more processors may be configured to decode the encoded bitstream to receive downmix channels, spatial metadata, and mono mode indicators. In some exemplary embodiments, one or more processors may be configured to set one or more parameters of the spatial metadata to 0 when mono mode is detected. According to some exemplary embodiments, one or more processors may be configured to upmix the downmix channels using the spatial metadata. In some exemplary embodiments, one or more processors may be configured to render the upmixed channels to a desired audio format.

[0039] According to some exemplary embodiments, rendering may generate rendered audio data. According to some such exemplary embodiments, one or more processors may be further configured to transmit the rendered audio data, store the rendered audio data, transmit the rendered audio data, or a combination thereof. In some exemplary embodiments, one or more processors may be configured to provide the rendered audio data to one or more loudspeakers for playback.

[0040] Embodiments described herein may generally be described as techniques, where “technique” may mean a system, device, method, computer-readable instruction, module, component, hardware logic, and / or operation, as suggested by the context in which it is applied herein.

[0041] Features and technical benefits not explicitly stated above will become apparent upon reading the detailed description below and examining the relevant drawings. This abstract is provided to present a simplified selection of techniques and is not intended to identify key or essential features of the claimed subject matter as defined by the attached claims. [Brief explanation of the drawing]

[0042] [Figure 1] This is a block diagram of an example IVAS codec framework.

[0043] [Figure 2A] This flowchart outlines various exemplary methods that can be performed by what may be referred to herein as an "encoder-mono detector."

[0044] [Figure 2B]This flowchart outlines additional exemplary methods that can be performed by an encoder-mono detector.

[0045] [Figure 3] This flowchart outlines several examples of a band classification process, including sequential iterative loops across various frequency bands for a given frame.

[0046] [Figure 4] This flowchart outlines several examples of frame marking processes, including sequential iterative loops across various frequency bands for a given frame.

[0047] [Figure 5] This flowchart outlines some additional exemplary methods that may be performed by decoders such as those disclosed herein.

[0048] [Figure 6] This block shows some examples of the IVAS codec system.

[0049] [Figure 7] This is a block diagram of an exemplary hardware architecture suitable for implementing the systems, devices, and methods described herein.

[0050] [Figure 8] This is a block diagram of an encoder having frequency-based and channel-based splitting between SPAR and DirAC according to one or more embodiments.

[0051] [Figure 9] This is a block diagram of a decoder having frequency-based and channel-based splitting between SPAR and DirAC according to one or more embodiments.

[0052] [Figure 10]This is a block diagram of an alternative encoder having frequency-based and channel-based splitting between SPAR and DirAC according to one or more embodiments.

[0053] [Figure 11] This is a block diagram of an alternative decoder having frequency-based and channel-based splitting between SPAR and DirAC. Both are configured according to the embodiments described herein.

[0054] In drawings, the specific arrangement or order of schematic elements, such as those representing devices, units, instruction blocks, and data elements, is shown for the purpose of facilitating explanation. However, it should be understood by those skilled in the art that the specific arrangement or order of schematic elements in the drawings is not intended to imply that a specific order or sequence of processing or separation of processes is required. Furthermore, the inclusion of schematic elements in the drawings is not intended to imply that such elements are required in all embodiments, or that features represented by such elements are not included in or combined with other elements in some implementations.

[0055] Furthermore, when connecting elements such as solid or dashed lines or arrows are used in drawings to indicate connections, relationships, or associations between two or more other schematic elements, the absence of such connecting elements is not intended to imply that such connections, relationships, or associations cannot exist. In other words, some connections, relationships, or associations between elements are not shown in the drawings so as not to obscure the disclosure. Moreover, for the sake of clarity, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, if a connecting element represents the communication of signals, data, or instructions, it should be understood by those skilled in the art that such an element may, as necessary, represent one or more signal paths affecting the communication.

[0056] The same reference symbols used in various drawings represent similar elements. [Modes for carrying out the invention]

[0057] The following detailed description includes many specific details to provide a full understanding of the various embodiments described, with reference to the accompanying drawings. The detailed description, drawings, and exemplary embodiments in the claims are not intended to limit. Other embodiments can be utilized and other modifications made without departing from the spirit or scope of this disclosure. It will be apparent to those skilled in the art that many of the various features and implementations described can be carried out without many of these specific details. In some cases, well-known methods, procedures, components, and circuits are not described in detail so as not to unnecessarily obscure aspects of the embodiments. Several features are described below, each of which may be used independently of others or in any combination of other features. Thus, features may be arranged, replaced, combined, separated, or designed in other configurations considered in light of this disclosure.

[0058] term Where used herein, the term “including” and its variations should be read as an open-ended expression meaning “including, but not limited to.” The term “or” should be read as “and / or” unless the context explicitly indicates otherwise. The term “based on…” should be read as “based at least in part on….” The terms “one exemplary implementation” and “a certain exemplary implementation” should be read as “at least one exemplary implementation.” The term “another implementation” should be read as “at least one other implementation.” The terms “determined,” “determine,” or “determining” should be read as “obtain,” “receive,” “compute,” “calculate,” “estimate,” “predict,” or “derive.” Furthermore, in the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meanings as generally understood by those skilled in the art to which this disclosure belongs. Abbreviation IVAS Immersive Voice and Audio Services SBA Scene Based Audio HOA Higher Order Ambisonics FOA First Order Ambisonic ACN - Ambisonic Channel Number DOA Direction of Arrival FFT (Fast Fourier Transform) DFT (Discrete Fourier Transform) MDFT (Modified Discrete Fourier Transform) MD Metadata BS Bistream EVS Enhanced Voice Services SPAR Spatial Reconstruction DIRAC Directional Audio Coding D2S DIRAC to SPAR S2D SPAR to DIRAC (SPAR to DIRAC conversion) LBRSBA (Low Bitrate SBA) CLDFB Complex Low Delay Filter Bank ESD Equivalent Spatial Domain SN3D Schmidt Semi-Normalization COV Covariance DMX Downmix AR (Augmented Reality) VR (Virtual Reality) CPU (Central Processing Unit) DSP (Digital Signal Processor) ASIC (Application-Specific Integrated Circuit) FPGA Field-Programmable Gate Array RAM (Random Access Memory) ROM (Read Only Memory) EPROM (Erasable Programmable ROM) CD-ROM Compact Disc Read-Only Memory I / O Input / Output

[0059] Exemplary IVAS codec framework Figure 1 is a block diagram of an Immersive Voice and Audio Services (IVAS) coder / decoder ("codec") framework 100 for encoding and decoding IVAS bitstreams, according to one or more embodiments. IVAS is expected to support a range of audio service functions, including but not limited to mono-to-stereo upmixing and fully immersive audio encoding, decoding, and rendering. IVAS is also intended to be supported by a wide range of devices, endpoints, and network nodes, including but not limited to mobile phones and smartphones, electronic tablets, personal computers, conference phones, conference rooms, virtual reality (VR) and augmented reality (AR) devices, home theater devices, and other suitable devices.

[0060] The IVAS codec 100 includes an IVAS encoder 101 and an IVAS decoder 104. The IVAS encoder 101 includes a spatial encoder 102 that receives N channels of input spatial audio (e.g., FOA, HOA). In some implementations, the spatial encoder 102 implements SPAR and DirAC for parsing / downmixing N_dmx spatial audio channels, as will be described in more detail below. The output of the spatial encoder 102 includes a spatial metadata (MD) bitstream (BS) and N_dmx channels of spatial downmix. The spatial MD is quantized and entropy coded. In some implementations, quantization can include fine, moderate, coarse, and extra-coarse quantization strategies, and entropy coding can include Huffman coding or arithmetic coding. In some implementations, the framework can only allow up to 3 levels of quantization in a given operating mode. However, as the bitrate decreases, in some such implementations, those 3 levels become progressively coarser overall to meet the bitrate requirements. The core audio encoder 103 (e.g., based on an Enhanced Voice Services (EVS) encoding unit) encodes N_dmx channels (N_dmx = 1 to 16 channels) of the spatial downmix into an audio bitstream, which is combined with the spatial MD bitstream to become the IVAS-encoded bitstream sent to the IVAS decoder 104. Given bitrate constraints for low-bitrate scene-based audio (SBA), as described below, some implementations limit the number of channels to a single channel.

[0061] The IVAS decoder 104 includes a core audio decoder 105 (e.g., an EVS decoder) that decodes the audio bitstream extracted from the IVAS bitstream to recover N_dmx audio channels. The spatial decoder / renderer 106 (e.g., SPAR / DirAC) decodes the spatial MD bitstream extracted from the IVAS bitstream to recover the spatial MD and synthesizes / renders the output audio channels using the spatial MD and spatial upmix for playback on various audio systems with different speaker configurations and capabilities.

[0062] Low bitrate SBA (LBRSBA) In some embodiments, it is desirable to implement LBRSBA (e.g., Ambisonics) using the SPAR-DIRAC codec. LBRSBA can be achieved using one or more of the following techniques: 1) reduced MD bitrate and bandwidth interleaving; 2) active W tuning; 3) additional covariance smoothing to facilitate the reduced MD bitrate; and 4) decoder-side decorrelation coefficient smoothing.

[0063] Background information regarding the above techniques may be found in one or more of the following documents, all of which are incorporated herein by reference. [Patent Document 1] PCT Application No. 2023 / 063769, "DirAC-SPAR Audio Processing" [Patent Document 2] U.S. Patent No. 9,978,385, "Parametric reconstruction of audio signals" [Patent Document 3] International Publication No. 20212252748A1, "Encoding of multi-channel audio signals comprising downmixing of a primary and two or more scaled non-primary input channels" [Patent Document 4] International Publication No. 2022120093A1, "Immersive Voice and Audio Services (IVAS) with Adaptive Downmix Strategies" [Patent Document 5] International Publication No. 2021252811A2, "Quantization and Entropy Coding of Parameters for a Low Latency Immersive Audio Codec" [Patent Document 6] U.S. Patent Application Publication No. 20220406318A1, “Bitrate Distribution in Immersive Voice and Audio Services” [Patent Document 7] U.S. Patent Application Publication No. 2022 / 0277757, "Systems and Methods for Covariance Smoothing"

[0064] LBRSBA can also be achieved using other techniques known to those skilled in the art.

[0065] When operating the IVAS codec, audio content may be changed from ambisonic format to mono format, but the device streaming and encoding the audio content cannot signal the format change to the codec. When this occurs, the decoder may treat the audio content as ambisonic format, which can lead to undesirable audio characteristics. Mono can be defined as a format with a single channel of audio content, while ambisonic can be defined as a multi-channel audio format with four or more channels.

[0066] In this case, a single mono channel is transmitted as the W channel of the ambisonic input, while the other channels are silent or nearly silent. An example of nearly silent channels is when dithering is applied to reduce quantization errors on the silent channels. In some implementations, the IVAS codec SBA mode cannot handle this use case due to certain design choices already made. SBA mode is a term used in the IVAS codec for encoding and decoding ambisonic content.

[0067] The following is a proposed system to ensure proper behavior of the IVAS codec with mono input when operating in SBA mode. This system includes a mono detector in the encoder, configured for mono signal transmission through the bitstream (possibly without using an explicit signal transmission bit), and includes a mono detector / storage to ensure that the W channel is treated as mono in the decoder. In some examples, the system works frame by frame, but there are some inter-frame dependencies.

[0068] Encoder / Encoder Mono Detection Several exemplary processes for detecting objects in an encoder are described below. For convenience, these exemplary processes can be considered as roughly two-step processes. Step 1 is frequency band classification, and Step 2 is frame marking (or declaration). These two processes may be combined into a single process or separated into additional processes or subprocesses, as may be desired in a particular implementation. Furthermore, the functions described herein for implementing the overall process may be replaced by other functions, and some of the steps described may be omitted and / or replaced by other more suitable implementations without departing from the spirit of this disclosure.

[0069] Figure 2A is a flowchart outlining various exemplary methods 200 that may be performed by what may be referred to herein as “encoder-mono detector”. Exemplary methods 200 may be divided into blocks such as blocks 205, 210, 215, 220, 225, 230, and 235. Various blocks may be described as operations, processes, methods, steps, steps, or functions. The blocks of method 200, as with other methods described herein, may not necessarily be performed in the order shown. In some implementations, one or more blocks of method 200 may be performed concurrently. Furthermore, some implementations of method 200 may include more or fewer blocks than those illustrated and / or described. The blocks of method 200 are involved in encoding an audio signal, which may be performed by one or more devices such as the device shown in Figure 7, for example. Processing may begin in block 205.

[0070] Block 205 deals with “acquisition of the audio signal”. In various examples, the audio signal represents the input audio scene with a primary channel and side channels. In some examples, the primary channel may be the W channel of an ambisonics format such as a first-order ambisonics format, and the side channels may be the X, Y, and Z channels, or include them. In other examples, the primary channel may be the center channel of a multi-channel format, and the side channels may include the left and right channels of a multi-channel format. In yet another example, the primary channel may be the mid channel of audio captured using the mid-side microphone technique, and the side channels may correspond to the side channels. Processing may continue to Block 210.

[0071] Block 210 relates to “Analysis of the power of the main channel of the audio signal.” Block 215 relates to “Analysis of the power of the side channel of the audio signal.” In some examples, the analyses in blocks 210 and 215 may be performed in the frequency domain. Various examples of frequency domain analysis are disclosed herein. In some alternative examples, the analyses in blocks 210 and 215 may be performed in the time domain. The processing may then proceed to block 220.

[0072] Block 220 includes "detecting mono mode for encoding an audio signal based on an analysis of the power of the main channel and the power of the side channels." In some examples, mono mode detection may be based on the determination that the input audio signal has non-silent audio in the main channel and silent in the side channels. Processing may then proceed to Block 225.

[0073] In some examples, mono-mode detection may involve calculating the primary channel power of the primary channel of the audio signal and calculating the total power as the sum of the powers of the side channels of the audio signal. In some such examples, mono-mode detection may involve evaluating the ratio of the primary channel power to the total power. According to some such examples, mono-mode detection may involve determining when the ratio exceeds a threshold.

[0074] Block 225 involves "calculating one or more downmix channels and spatial metadata from the audio signal for the detected mono mode." In some examples, the spatial metadata may involve spatial reconstruction (SPAR) spatial metadata, directional audio coding (DirAC) spatial metadata, or both. Processing may then proceed to block 230.

[0075] Block 230 involves "encoding one or more downmix channels and spatial metadata in the bitstream for the detected mono mode." In some examples, the bitstream may be an immersive audio and audio (IVAS) encoded bitstream. Block 235 involves "indicating mono mode in the bitstream."

[0076] In some examples, method 200 may involve implicitly signaling mono-mode according to one or more parameters of spatial metadata. In some such examples, method 200 may involve implicitly signaling mono-mode by setting one or more parameters of spatial metadata to a value of 0.

[0077] In some examples, an encoder mono detector is configured to determine when an ambisonic input has content only in the W channel. In some such examples, for each audio block, the encoder mono detector analyzes the content in all channels and determines whether the audio block can be declared as mono or ambisonic.

[0078] In some examples, the encoder-mono detector operates in the Modified Discrete Fourier Transform (MDFT) domain by looping across each frequency band and calculating the power of the W channel (Wp) and the sum of the powers of the other channels (Sump) at the input for each band. According to some examples, the encoder-mono detector uses the Wp power, the sum of the powers of the other channels (Sump), and the ratio of the Wp power to the sum of the powers of the other channels (Wp / Sump) to determine whether the band is mono, SBA (ambisonics), or silent. In some examples, the classification decision (mono, multichannel, or silent) is based on two thresholds: a static NOISE_THRESHOLD and a dynamic RATIO_THRESHOLD.

[0079] NOISE_THRESHOLD can be a static threshold that can be used to determine whether a frequency band is silent or contains content. For example, frequency bands below NOISE_THRESHOLD may be considered silent or nearly silent, while frequency bands above NOISE_THRESHOLD may be considered to contain content. RATIO_THRESHOLD may be a dynamic threshold that scales with the content as the content becomes quieter, and can be used to classify frequency bands as ambisonic or mono. For example, frequency bands below RATIO_THRESHOLD may be considered ambisonic, while frequency bands above RATIO_THRESHOLD may be considered mono.

[0080] Static thresholds (e.g., NOISE_THRESHOLD) can be determined by passing silent and near-silent content through the codec and analyzing the levels generated by the MDFT. According to some examples, static thresholds may range from 40dB to 45dB, 42dB to 48dB, etc. For example, in some cases, the static threshold may be 44dB or 45dB. Dynamic thresholds (e.g., RATIO_THRESHOLD) can be calculated using a linear function with W channel power (Wp) as input, where power Wp can be clipped at maximum and minimum thresholds (e.g., MAX_THRESHOLD, MIN_THRESHOLD) to improve stability. In some examples, minimum thresholds may range from 15dB to 22dB, 18dB to 25dB, etc. For example, in some cases, the minimum threshold may be 20dB. According to some examples, maximum thresholds may range from 55dB to 70dB, 50dB to 65dB, etc. For example, in some cases, the maximum threshold may be 60 dB. In some cases, the minimum and maximum thresholds apply to the normalized power value. All of these thresholds may depend on the scaling within the IVAS codec.

[0081] In various experiments, maximum and minimum thresholds were determined empirically by analyzing mono and normal ambisonic content and ensuring that the discrimination level met the requirements for accurate detection. In some examples, the requirement for accurate detection may be that there is no or very little leakage into non-W channels after decoding. Alternatively or additionally, the requirement for accurate detection may be that multi-channel audio is not misclassified as mono. However, this disclosure is not limited to this particular method of determining maximum and minimum thresholds, and the thresholds may be adjusted accordingly.

[0082] According to some examples, after determining the state of all frequency bands, the encoder mono detector is configured to check whether any frequency band has been classified (or declared) as ambisonic. If any frequency band has been classified (or declared) as ambisonic, the encoder mono detector may determine that the entire audio frame can be declared (or marked) as non-mono. If no frequency band has been classified (or declared) as ambisonic, the encoder mono detector may determine that all frequency bands must be considered (or declared) as either silent or mono. If at least one frequency band has been considered (or declared) as mono, the encoder mono detector may determine that the frame is classified (or declared) as mono.

[0083] Declaring a frame as ambisonic or mono can be implemented in some examples using mono flags. Some implementations have additional processes that can be used to ensure that the mono flag does not change state too quickly (e.g., rapidly switching on and off), which would cause instability. In some examples, the additional process may involve parsing the history of the mono flag. In some such examples, the encoder mono detector declares the frame as mono and signals this mono declaration to the decoder only after there have been consecutive mono flags over a certain time period (e.g., 20 milliseconds (ms), 25 milliseconds, 30 milliseconds, 35 milliseconds, 40 milliseconds, etc.).

[0084] Figure 2B is a flowchart outlining additional exemplary methods that may be performed by the encoder-mono detector. Exemplary Method 240 may be divided into blocks such as blocks 245, 250, 255, 260, 265, and 270. The various blocks may be described as operations, processes, methods, steps, steps, or functions. The blocks of Method 240, as with other methods described herein, may not necessarily be performed in the order shown. In some implementations, one or more blocks of Method 240 may be performed concurrently. Furthermore, some implementations of Method 240 may involve more or fewer blocks than those illustrated and / or described. The blocks of Method 240 involve encoding an audio signal representing an input audio scene using a primary channel and side channels. The primary channel and side channels may be any of the types described, for example, with reference to Figure 2A. The blocks of Method 240 may be performed by one or more devices, such as the device shown in Figure 7. Processing may begin in block 245.

[0085] Block 245 deals with "getting frames of the audio signal from the bitstream." Processing may continue to block 250.

[0086] Block 250 involves "identifying multiple frequency bands associated with the frame." The process may continue to block 255. Block 255 involves "determining the power associated with the primary channel and side channels for each of the multiple frequency bands associated with the frame." In some examples, determining the power of a frequency band may involve calculating the primary channel power of the frequency band and then calculating the total power of the frequency band. The total power may be the sum of the powers for the side channels. The process may continue to block 260.

[0087] Block 260 relates to "classifying each of several frequency bands as one of silent, mono, or ambisonic based on the determined power for the corresponding frequency band." In some examples, classifying each of several frequency bands may relate to evaluating the ratio of the primary channel power to the total power. In some such examples, detecting mono mode may relate to determining when the ratio exceeds a threshold. In some such examples, the threshold may be a linear function of the primary channel power. In some examples, the primary channel power may be normalized to a value between 0 and 1. In some examples, the threshold may be clipped between a maximum and a minimum value.

[0088] In some examples, classifying each of several frequency bands may involve determining that the sum power is below a minimum noise level and declaring that frequency band as mono if the primary channel power is above a noise threshold. In some examples, classifying each of several frequency bands may involve determining that the sum power is above a minimum noise level and declaring that frequency band as mono if the normalized primary channel power is above a noise threshold. Processing can then proceed to block 265.

[0089] Block 265 relates to "marking the mode of a frame as silent, mono, or ambisonic, based on several classified frequency bands for the frame." In some examples, marking a frame may relate to marking the frame as ambisonic if one or more classified frequency bands correspond to ambisonic. In some examples, marking a frame may relate to marking the frame as mono if none of the classified frequency bands correspond to ambisonic and one or more of the classified frequency bands correspond to mono. In some examples, marking a frame may relate to marking the frame as silent if none of the classified frequency bands correspond to ambisonic and none of the classified frequency bands correspond to mono. In some examples, marking a frame may relate to setting one or more SPAR metadata parameters to 0, setting one or more DirAC metadata parameters to 0, or both. Processing can then proceed to block 270.

[0090] Block 270 includes "encoding the marked mode of a frame in the bitstream." In some examples, the bitstream may be an immersive audio and audio (IVAS) encoded bitstream. In some examples, marking the mode of a frame may involve setting one or more spatial reconstruction (SPAR) metadata parameters to 0, one or more directional audio coding (DirAC) metadata parameters to 0, or both.

[0091] Another exemplary object detection process for encoders may involve the following: ● For each frame of the encoder, a loop is performed across each MDFT frequency band. ●Calculate the following for each frequency band. ○W channel power (W_p); and ○ The sum of the power of each individual non-W channel. ● To prevent division by zero errors, check whether the sum power is close to zero. ○If the total power is close to 0, check if the W power is greater than NOISE_THRESHOLD, and if so, declare that bandwidth as mono. ○ Otherwise, restart the loop iteration process. ●Calculate the ratio of wattage power to total power. ●Calculate RATIO_THRESHOLD. ○Normalize the wattage (W) between 0 and 1. ○ Scale MAX_TRESHOLD using norm W power. ○ Clip the calculated threshold to ensure that the range is between MIN_THRESHOLD and MAX_TRESHOLD. ● Check whether the calculated ratio exceeds RATIO_THRESHOLD. ○If the ratio exceeds RATIO_THRESHOLD, declare that bandwidth as mono. ○ Otherwise, declare that bandwidth as ambisonics. ●If any of these checks fail, the frequency band will be declared silent. ● After the frequency band loop, check if any of the bands have been declared as ambisonics. ○If true, the encoder frame is non-mono; Otherwise, check whether any of the bandwidths have been declared as mono. ■If true, the encoder frame is mono; ■Otherwise, the encoder frame is non-mono. ●If the frame is a mono, increment the mono-frame counter. ○If the frame counter is >30 milliseconds, set the global mono flag to 1. ●If the frame is not mono, reset the mono frame counter to 0.

[0092] Figure 3 is a flowchart outlining several examples of a band classification process that includes a sequential iterative loop across various frequency bands for a given frame. The exemplary method 300 can be divided into blocks such as blocks 305, 310, 315, 320, 325, 327, 330, 335, 340, 345, 350, and 355. The various blocks may be described as operations, processes, methods, steps, steps, or functions. Method 300 may be performed by an apparatus or system configured to implement an encoder-mono detector. The blocks of Method 300, as with other methods described herein, are not necessarily performed in the order shown. In some implementations, one or more blocks of Method 300 may be performed concurrently. Furthermore, some implementations of Method 300 may involve more or fewer blocks than those illustrated and / or described. The blocks of Method 300 relate to a band classification process that may be performed by one or more devices, such as the device shown in Figure 7. Processing can be started in block 305.

[0093] Block 305 deals with obtaining a frequency band of the audio signal. In this example, the frequency band is the MDFT frequency band, but in other examples, the frequency band may be of a different type. The process continues in Block 310. In this example, Method 300 deals with the following: ●For each frame of the audio signal (where the audio signal represents an audio scene with a main channel and side channels) (e.g., frame loop): Each MDFT frequency band associated with a frame (e.g., a frequency band loop) is classified based on the power of the corresponding MDFT frequency band as follows: ■Calculate the power of the corresponding MDFT frequency band as follows: ●W channel power (W) in block 310 p); and ● The sum power (Sum p ) as the sum of the non-W channel powers in block 315. ■ Analyze the power calculation (W p , Sum p ) and determine whether its band is silent, mono, or ambisonics as follows: ● In block 320, determine whether the sum power (Sum p ) is below the nominal minimum noise level ε. Here, ε is a minimum value that may be close to 0. In some examples, the minimum noise level ε may be equal to NOISE_THRESHOLD. ● If the sum power (Sum p ) is determined to be below the nominal minimum noise level: ○ In block 325, determine whether the W channel power (W p ) is greater than or less than NOISE_THRESHOLD; ○ If in block 32, it is determined that the channel power is below the noise threshold (e.g., W p ≤NOISE_THRESHOLD), declare that band as silent; ○ If in block 330, it is determined that the channel power exceeds the noise threshold (e.g., W p >NOISE_THRESHOLD), declare that band as mono; and ○ Proceed to the next frequency band of that frame and return to block 310. ● If in block 320, it is determined that the sum power exceeds the nominal minimum noise level (e.g., Sum p ≥ε): ○ Calculate the following: ■ The ratio (R) of the W channel power (W p ) to the sum power (Sum p ) in block 335; ■ RATIO_THRESHOLD in block 340. In some examples, block 340 is the W channel power (W pThis may involve using a linear function with ) as input. • Normalize the W channel power between 0 and 1 (Norm_W p ) • Scale MAX_THRESHOLD using normalized wattage. • Clip the calculated threshold to ensure that the range is between MIN_THRESHOLD and MAX_TRESHOLD. ○ In block 345, determine whether the calculated ratio (R) is greater than or less than RATIO_THRESHOLD. ○In block 355, if the calculated ratio is determined to be below the ratio threshold (e.g., R ≤ RATIO_THRESHOLD), declare that bandwidth as ambisonics; ○In block 350, if the calculated ratio is determined to exceed the ratio threshold (e.g., R>RATIO_THRESHOLD), declare that bandwidth as mono; and ○For example, by returning to block 310, we move to the next frequency band of the frame. ○Some disclosed examples of Method 300 involve marking a frame as either ambisonic, mono, or silent based on a classified MDFT frequency band for the frame, by: ■ If one or more of the classified MDFT frequency bands for a frame are declared as ambisonic, mark that frame as ambisonic; ■ If none of the classified MDFT frequency bands for a frame are declared as ambisonics, and one or more of the classified MDFT frequency bands for a frame are declared as mono, mark that frame as mono; ■If none of the classified MDFT frequency bands for a frame are declared as ambisonics, and none of the classified MDFT frequency bands for a frame are declared as mono, the frame is marked as silent.

[0094] Figure 4 is a flowchart outlining several examples of a frame marking process that includes a sequential iterative loop across various frequency bands for a given frame. The exemplary method 400 can be divided into blocks such as blocks 405, 410, 415, 420, 425, 430, 435, 440, 445, and 450. The various blocks may be described as operations, processes, methods, steps, steps, or functions. Method 400 may be performed after a frame classification process, for example, as shown in Figure 2B or Figure 3. In some examples, Method 400 may be performed by an apparatus or system configured to implement an encoder-mono detector. The blocks of Method 400, as with other methods described herein, are not necessarily performed in the order shown. In some implementations, one or more blocks of Method 400 may be performed concurrently. Furthermore, some implementations of Method 400 may involve more or fewer blocks than those illustrated and / or described. The block of method 400 relates to a frame marking process that may be performed by one or more devices, such as the device shown in Figure 7. The process may be initiated in block 405.

[0095] Block 405 involves obtaining the classified frequency bands of the audio frame. The classified frequency bands may have been previously classified according to one of the disclosed classification processes. In this example, the classified frequency bands have been previously classified as ambisonic, mono, or silent. Processing may then proceed to Block 410.

[0096] Block 410 is involved in determining the classification of the frequency bands obtained in block 405. In this example, the classification results are provided to block 415. Processing may continue from block 415.

[0097] Block 415 is involved in determining whether the frequency band is classified as ambisonics. In this example, in block 415, if it is determined that the frequency band is classified as ambisonics, the process flow continues from block 415 to block 420, where it is determined that the entire frame containing the classified frequency band is not mono. Since no further analysis is required, in this example, the process flow continues to "end of loop" in block 425. However, if in block 415 it is determined that the frequency band is not classified as ambisonics, the process flow can continue from block 415 to block 430.

[0098] Block 430 is involved in determining whether the frequency band is classified as mono. According to this example, in block 430, if it is determined that the frequency band is classified as mono, the process flow can continue from block 430 to block 435, where the entire frame is tentatively marked as mono. However, if in block 430 it is determined that the frequency band is not classified as mono, the process flow can continue from block 430 to block 440.

[0099] Block 440 is involved in determining whether the classified frequency band evaluated in block 430 is the last classified frequency band of the audio frame. If not, the next classified frequency band is obtained and the process can return from block 440 to block 410. However, if in block 440 it is determined that the classified frequency band evaluated in block 430 is the last classified frequency band of the audio frame, the process can continue from block 440 to block 445.

[0100] Block 445 is involved in determining whether the entire frame is tentatively marked as a single item. According to this example, in block 445, if it is determined that the entire frame is tentatively marked as a single item, the process can continue from block 445 to block 450. In this example, in block 450, the entire frame is classified as a single item. The process flow can continue from block 450 to the "loop end" block 425. However, if in block 445 it is determined that the entire frame is not tentatively marked as a single item, the process flow can continue to the "loop end" block 425.

[0101] send In some examples, the single-item detection information identified here may be transmitted from the encoder to the decoder using implicit signaling techniques. In some such examples, existing metadata may be modified to specific values to support this technique without requiring additional data bits within the bitstream. Thus, implicit signaling eliminates potential increases in data transmission overhead such as potential increases in bandwidth, speed, or data storage / buffering requirements.

[0102] According to some examples, implicit signaling may be achieved by using existing SBA metadata and setting the SBA metadata to specific values. By using existing SBA metadata, the bit accuracy of existing tests can be maintained. In one example, six different SPAR and DIRAC metadata fields are set to 0 in the encoder and then checked in the decoder. If the six metadata fields are 0 in the decoder, the decoder processes that frame as a single item. Signaling also takes into account quantization of metadata values through the codec. Due to quantization, certain values cannot be encoded in the bitstream. The detector in the decoder takes this quantization into account to correctly detect single items.

[0103] Some examples of metadata values ​​for implicit signaling include: ● SPAR metadata ○ Prediction coefficient ○ Mutual prediction coefficient ○ Decorrelation coefficient ●DiRAC metadata ○ Energy ratio ○Azimuth ○Elevation angle

[0104] In some cases, when a mono is detected, all of these metadata values ​​are set to 0. The SPAR value relates to extracting higher-order channels from downmixed channels transmitted in the bitstream. If these values ​​are set to 0, the extraction is not completed. The DiRAC value relates to diffusion and location. If these values ​​are set to 0, the mono channel is not distributed across other channels in the ambisonics output.

[0105] Decoder / Decoder Mono Detection / Storage Device According to some examples, decoders are configured to analyze signal transmission through a bitstream (or implicit signal transmission through metadata). When mono is being transmitted, some implementations configure the decoder to set the DIRAC energy ratio to 0 and the spreading to 1 for all frequency bands and subframes within the decoder. This ensures that the W channel energy is not dispersed across other channels and that the audio signal remains as mono content.

[0106] According to some examples, the decoder is configured to look for the six values ​​mentioned above in each frequency band. In some such examples, if any of the values ​​in any band do not match the expected value of the thing, the entire block is declared as non-mono. Otherwise, in some examples, the block is declared as mono. According to some such examples, the expected values ​​for mono are all 0 except for energy_ratio, which is expected to be <0.15 due to the quantization of its value.

[0107] When a mono is detected in the decoder, in some cases the energy_ratio value is reset to 0 for all bandwidths and time slices within the block. In some such cases, the decoded metadata is not applied to all frequency bands within the decoder, and therefore the decoded metadata is extended to cover all frequency bands. This may be the reason why the energy_ratio value is reset to 0. In some such cases, the spreading vector applied to all frequency bands in the decoder is set to 1. This results in the correct behavior within the decoder for mono content contained within the ambisonics format.

[0108] Figure 5 is a flowchart outlining additional exemplary methods that may be performed by a decoder as disclosed herein. Exemplary Method 500 may be divided into blocks such as blocks 505, 510, 515, 520, and 525. Various blocks may be described as operations, processes, methods, steps, steps, or functions. The blocks of Method 500, as with other methods described herein, may not necessarily be performed in the order shown. In some implementations, one or more blocks of Method 500 may be performed concurrently. Furthermore, some implementations of Method 500 may involve more or fewer blocks than those illustrated and / or described. The blocks of Method 500 relate to audio signal decoding that may be performed by one or more devices, such as the device shown in Figure 7. Processing can begin in block 505.

[0109] Block 505 includes “Obtaining the Encoded Bitstream.” In this example, the bitstream has been previously encoded according to one of the disclosed encoding methods. In some examples, the encoded bitstream may be an Immersive Speech and Audio (IVAS) encoded bitstream. Processing may proceed to Block 510. Block 510 involves “Decoding the Encoded Bitstream and Obtaining the Downmix Channel, Spatial Metadata, and Mono-Mode Indicator.” In some examples, the mono-mode indicator may be based on the value of one or more spatial metadata parameters in the encoded bitstream, set to a value indicating mono-mode. According to some examples, the one or more spatial metadata parameters may include one or more SPAR metadata parameters, one or more DirAC metadata parameters, or both. Processing may proceed to Block 515.

[0110] Block 515 relates to "setting one or more parameters of spatial metadata to 0 when a mono-mode is detected." In some examples, setting one or more parameters of spatial metadata to 0 when a mono-mode is detected may relate to setting one or more energy ratio values ​​to 0. In some examples, method 500 may relate to setting the spreading vector for all frequency bands to 1 when a mono-mode is detected. The process can then proceed to block 520.

[0111] Block 520 involves “upmixing downmixed channels using spatial metadata.” Processing can then proceed to Block 525. Block 525 involves “rendering the upmixed channels to a desired audio format.” In some examples, Method 500 may involve providing the audio in the desired audio format to one or more loudspeakers in an audio environment. According to some examples, Method 500 may involve playing the audio in the desired audio format through one or more loudspeakers.

[0112] Exemplary system Figure 6 is a block diagram showing some examples of the IVAS codec system. As with other disclosed implementations, the number, types, and arrangement of elements shown in Figure 6 are merely examples. Other implementations may involve more elements, fewer elements, and / or different arrangements of elements.

[0113] In this example, the IVAS codec system 600 includes an exemplary IVAS encoder 605 configured to encode dual-ended detection and an exemplary IVAS decoder 655 configured to decode based on dual-ended detection. In some examples, the IVAS encoder 605 and the exemplary IVAS decoder 655 may be implemented by an instance of the electronic device architecture 700 shown in Figure 7. In some examples, the IVAS codec system 600 may be configured for encoding and decoding according to one or more implicit signaling methods described above.

[0114] In the example shown in Figure 6, the IVAS encoder 605 includes a mono detector module 610 and an encoder module 615. In these examples, the mono detector module 610 is configured to receive an audio input signal 601 and to output modified metadata 612 to the encoder module 615. The modified metadata 612 may include DirAC metadata, SPAR metadata, or a combination thereof. According to these examples, the audio data 601 is in ambisonics format. Thus, the audio data 601 is FOA or some kind of HOA. The mono detector module 610 may be configured to perform one of the disclosed mono detection methods, depending on the specific implementation. According to these examples, the mono detector module 610 is configured to analyze the power of the primary channel of the audio signal, in this example the ambisonics W channel. In these examples, the mono detector module 610 is also configured to analyze the power of the side channels of the audio signal, in this example at least the ambisonics X, Y, and Z channels. According to these examples, the mono detector module 610 is configured to detect mono modes for encoding an audio signal based on an analysis of the power of the main channel and the power of the side channels.

[0115] In these examples, the encoder module 615 is configured to receive an audio input signal 601 and modified metadata 612, and to output an encoded bitstream 620. According to these examples, the encoder module 615 is configured to calculate one or more downmix channels and spatial metadata from the audio signal for mono modes detected by the mono detector module 610. In these examples, the encoder module 615 is configured to encode the one or more downmix channels and spatial metadata into a bitstream for the detected mono modes, and to indicate mono modes in the encoded bitstream 620. According to these examples, the encoder module 615 is configured for both SPAR and DirAC encoding. In this example, the encoded bitstream 620 is an IVAS bitstream.

[0116] In the example shown in Figure 6, the IVAS decoder 655 includes a mono detector and preserver module 660 and a decoder module 655. Here, the IVAS decoder 655 is configured to receive an encoded bitstream 620 and output a rendered audio signal 670. According to this example, the IVAS decoder 655 is configured to decode the encoded bitstream 620 and obtain downmix channels, spatial metadata, and one or more mono-mode indicators.

[0117] In these embodiments, the mono detector and the storage module 660 are configured to receive the encoded bitstream 620 and output the modified metadata 662 to the decoder module 655. The modified metadata 662 may relate to DirAC metadata, SPAR metadata, or a combination thereof. According to these examples, the mono detector and the storage module 660 are configured to detect the mono mode based on the one or more mono mode indicators. In this example, the one or more mono mode indicators are based on the values of one or more spatial metadata parameters in the encoded bitstream that are set to a value indicating the mono mode. According to these examples, the one or more mono mode indicators include one or more SPAR metadata parameters, one or more DirAC metadata parameters, or both. According to some examples, the mono detector and the storage module 660 may be configured to set one or more parameters of the spatial metadata to 0 when the mono mode is detected. For example, the mono detector and the storage module 660 may be configured to set one or more energy ratio values to 0. In some examples, the mono detector and the storage module 660 may be configured to set the diffusivity value to 1.

[0118] In these examples, the decoder module 655 is configured to receive an encoded bitstream 620 and modified metadata 662, and to output a rendered audio signal 670. In this example, the decoder module 655 is configured to upmix the downmixed channels of the encoded bitstream 620 using the spatial metadata of the modified metadata 662, and to render the upmixed channels to a desired audio format to produce the rendered audio signal 670. In these examples, the decoder module 655 is configured to upmix the downmixed channels using spatial metadata and to render the upmixed channels to a desired audio format based at least in part on the modified spatial metadata received from the mono detector and storage module 660.

[0119] Exemplary system architecture Figure 7 shows a block diagram of an exemplary electronic device architecture 700 suitable for implementing the systems, devices, and methods described herein. Architecture 700 includes, but is not limited to, the server and client devices described herein. As illustrated, architecture 700 includes a central processing unit (CPU) 701 capable of executing various processes according to, for example, a program stored in read-only memory (ROM) 702, or a program loaded from, for example, a memory unit 708 into random access memory (RAM) 703. The CPU 701 may include one or more general-purpose single-chip or multi-chip processors, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or a combination thereof. RAM 703 also stores, as needed, data required when the CPU 701 executes various processes. The CPU 701, ROM 702, and RAM 703 are connected to each other via bus 804. The input / output (I / O) interface 705 is also connected to bus 704.

[0120] The following components are connected to the I / O interface 705: an input unit 706 which may include a keyboard, mouse, etc.; an output unit 707 which may include a display such as a liquid crystal display (LCD) and one or more speakers; a storage unit 708 which includes a hard disk or another suitable storage device; and a communication unit 709 which includes a network interface card such as a network card (e.g., wired or wireless).

[0121] In some implementations, the input unit 706 includes one or more microphones located at different positions (depending on the host device) that enable the capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other appropriate formats).

[0122] In some implementations, the output unit 707 includes a system with varying numbers of speakers. The output unit 707 can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats) (depending on the capabilities of the host device).

[0123] In some embodiments, the communication unit 709 is configured to communicate with other devices (for example, via a network). The drive 710 is also connected to the I / O interface 705 as needed. Removable media 711, such as magnetic disks, optical disks, magneto-optical disks, flash drives, or other suitable removable media, are mounted on the drive 710, and computer programs read therefrom are installed in the storage unit 708 as needed. Those skilled in the art will understand that although the system 700 is described as including the above-described components, in actual applications it is possible to add, remove, and / or replace some of these components, and all such modifications or changes fall within the scope of this disclosure.

[0124] 1.0 Algorithm Analysis As mentioned above, SPAR attempts to maximize perceived audio quality while minimizing the bitrate by reducing the energy of the transmitted audio data, while allowing the decoder to reconstruct the secondary statistics (i.e., covariance) of the ambisonic audio scene using the transmitted metadata. DirAC attempts to preserve the dominant sonic directionality and diffusion in the input scene. DirAC and SPAR technologies are outlined in sections 1.1 and 1.3 below, respectively.

[0125] 1.1 DirAC Technology 1.1.1 Reference papers The DirAC technology is described in Non-Patent Document 1, which is incorporated by reference into this application. [Non-Patent Document 1] V. Pulkki, “Directional Audio Coding in Spatial Sound Reproduction and Stereo Upmixing”, Laboratory of Acoustics and Audio Signal Processing, Helsinki University of Technology, Finland, 2006.

[0126] 1.2 Exemplary Implementation of DirAC Analysis in the MDFT Domain 1.2.1 DirAC analysis in the MDFT region In one exemplary implementation, the DirAC analysis block takes an ambisonics time-domain FOA channel as input and uses a modified discrete Fourier transform (MDFT) to transform the FOA channel into the frequency domain. The intensity and reference power are then calculated in the MDFT domain. r ,w i ,x r ,x i ,y r ,y i ,z r ,z iIf we consider the real and imaginary bin samples of the W, X, Y, and Z channels of the FOA component of the ambisonic input in the MDFT domain, then the intensity corresponding to the frequency bin f of channel X is calculated as follows: I(f) x =w r *x r + w i *x i [1] Similarly, the intensities corresponding to the Y and Z channels are calculated.

[0127] The reference power E at frequency bin f is calculated as follows: E(f)=w r *w r +x r *x r +y r *y r +z r *z r [2]

[0128] The direction vector dv corresponding to the X channel (or front-to-back direction) and frequency bin f is calculated as follows:

number

[0129] 1.2.2 DirAC Parameter Estimation in Banded Regions The intensity, reference power, and direction vector for each bin are then converted to the banded domain by applying the absolute response of the filter bank to the above calculated values ​​in [1], [2], and [3]. The banded intensity, reference power, and direction vector in a specific frequency band are respectively I s ,E,dv s Let's assume that s can be x, y, or z.

[0130] 1.2.2.1 Calculation of DoA angle (azimuth and elevation) The azimuth and elevation angles of the dominant sound source in a scene for a given time-frequency tile are calculated in degrees as follows:

number

[0131] 1.2.2.2 Calculation of Diffusivity and Energy Ratio For diffusion, the long-term average of E and I is calculated over N frames or M subframes. In one exemplary implementation, a frame represents 20 milliseconds of audio data, a subframe represents 5 milliseconds of audio data, and the long-term average of E and I is performed over 160 milliseconds of audio data, i.e., 8 frames or 32 subframes. The long-term average is calculated over I. slow,s , E slow Therefore, the diffuseness is given as follows:

number

[0132] DirAC metadata parameters, namely the DoA angle and diffusion parameters, are quantized and encoded by the metadata quantization and encoding block. Based on the available bitrate, DirAC selects N_dmx audio channels (also called N_dmx downmix channels) from the N channel input, where N_dmx ≤ N, and one of the N_dmx downmix channels is the W channel of the ambisonics input to be encoded by the core coder. The core coder bits and DirAC metadata bits are multiplexed into a bitstream and sent to the decoder. The decoder uses the metadata dequantization and decoding block to decode the bitstream using the core decoder and DirAC metadata parameters to reconstruct the N_dmx downmix channels. The N_dmx downmix channels and DirAC metadata parameters are input to the DirAC synthesis and rendering block. The DirAC compositing and rendering block calculates the directional component of the output spatial audio scene using the W channel and spherical harmonics for each DoA angle. The DirAC compositing and rendering block also calculates the diffuse component of the output spatial audio scene using the decorrelated version of the W channel generated using the decorrelator block and the diffuseness parameter in the DirAC metadata. Then, using N_dmx downmix channels and the directional and diffuse components, it outputs the desired audio output format.

[0133] 1.3 Exemplary Implementation of SPAR (Spatial Reconstruction) with FOA Input SPAR is a technique for efficiently encoding spatial audio inputs. SPAR receives a multi-channel input and generates spatial metadata and a downmix signal. This allows for more efficient encoding of the combination of spatial metadata and the downmix signal compared to encoding each channel of the multi-channel input separately. SPAR aims to reproduce the covariance of an N-channel multi-channel input, calculating spatial metadata and an N_dmx-channel downmix signal (N_dmx ≤ N) based on a parameterized input covariance. The spatial metadata and downmix are quantized, encoded, and sent to a decoder. The decoder decodes the bitstream, dequantizes the spatial metadata, and reconstructs the downmix signal. The decoder then uses the spatial metadata and downmix, along with zero or more decorrelators, to reconstruct the multi-channel input audio scene. An exemplary implementation of SPAR is further described in PCT Patent Application PCT / US2023 / 010415 (Spatial Coding Of Higher Order Ambisonics For A Low Latency Immersive Audio CODEC), filed on 9 January 2023.

[0134] 1.3.1 Primary Ambisonics (FOA) Input For an FOA input consisting of channels W, Y, Z, and X (according to ACN channel ordering conventions), the SPAR downmix signal can vary across 1 to 4 channels, and the spatial metadata parameters include a prediction parameter PR, a cross-prediction parameter C, and a decorrelation parameter P. These parameters are calculated from the covariance matrix of the windowed input audio signal and over a specified number of frequency bands (e.g., 12 frequency bands). An illustrative representation of SPAR parameter extraction is shown below.

[0135] 1.3.1.1 Side Signal Prediction All side signals (Y, Z, X) are predicted from the main audio signal W, and the prediction coefficients for the residual channels are calculated using equation

[11] .

number

[11] .

number

[0136] The downmix described above is also called a passive W downmix, where W does not change during the downmix process. Another method of downmixing is an active W downmix, which allows some mixing of the Y, X, and Z channels into the W channel, as follows: W'=W+F Y *Y+F Z *Z+F X *X

[12] Here, F Y This is a function of the normalized input covariance RYW.

number

[0137] 1.3.1.2 Remixing of W channel and prediction channel (Y',Z',Z') The W channel and the prediction channel (Y', Z', X') are remixed from the most acoustically relevant to the least relevant. Here, as shown in equation

[13] , the remixing involves rearranging or rearranging the channels according to some methodology.

number

[0138] Note that one embodiment of remixing may involve rearranging the input channels to W, Y', X', and Z', based on the assumption that left and right audio cues are more important than front and back cues, with up and down cues being the last to be heard.

[0139] 1.3.1.3 Calculation of Covariance After Prediction The covariances of the 4-channel prediction and the downmix after remixing are calculated as shown in equations

[14] and

[15] .

number

[0140] In the example of a WABC downmix with 1 to 4 downmix channels, d and u represent the following channels. Here, placeholder variables A, B, and C can be any combination of the X, Y, and Z channels in the FOA. [Table 1]

[0141] 1.3.1.4 Additional C coefficient These calculations determine whether it is possible to mutually predict the rest of the fully parametric channel from the transmitted residual channel. The additional C coefficient required is as follows: C=R ud (R dd +Imax(ε,tr(R dd )*0.005)) -1

[16]

[0142] Therefore, C has the shape of (1×2) for a 3-channel downmix and (2×1) for a 2-channel downmix. One embodiment of spatial noise filling does not require these C parameters, and these parameters can be set to 0. Alternative embodiments of spatial noise filling may also include the C parameter.

[0143] 1.3.1.5 Remaining energy in the parameterized channel The remaining energy of the parameterized channel that needs to be filled by the decorrelator is calculated. The residual energy Resuu in the upmix channel is the difference between the actual energy Ruu (after prediction) and the regenerated cross-prediction energy Reguu.

number

[0144] 1.4 Merging DirAC and SPAR As mentioned above, both DirAC and SPAR have different strengths and characteristics, and it is desirable to combine the complementary aspects of each technology to produce a merged system that is advantageous in one or more of the following dimensions: higher audio quality, lower bitrate, flexibility in input / output formats, and / or reduced computational complexity. Some embodiments for efficiently merging these two technologies are listed below.

[0145] 1.4.1 Frequency-based SPAR-DirAC splitting Encoding the low-frequency band with SPAR and the higher-frequency band with DirAC has been observed to improve the coding efficiency and quality at the decoder while reconstructing the spatial audio scene. Furthermore, it may be desirable to perform efficient upmixing of the SPAR-reconstructed FoA signal to HOA by encoding the low-frequency band with SPAR to reconstruct the input covariance at the output, or by encoding the higher-frequency band with DirAC with the same or finer temporal resolution.

[0146] In the encoder, the first embodiment involves: 1) using a filter bank to convert a time-domain broadband ambisonics input to a frequency-banded domain; 2) performing DirAC analysis in the high-frequency band to obtain DirAC MD parameters in the high-frequency band; 3) performing SPAR analysis in the low-frequency band to obtain SPAR MD parameters in the low-frequency band; 4) obtaining SPAR MD parameters in the high-frequency band by converting the DirAC MD parameters to SPAR MD using an MD conversion routine (D2S) (described in sections 1.5 to 2.4); 5) generating a downmix matrix from the SPAR MD and applying the downmix matrix to the input channel to obtain the downmix channel described in section 2.3; 6) quantizing and encoding the SPAR MD parameters in the low-frequency band and the DirAC MD parameters in the high-frequency band; 7) encoding the downmix channel using a core audio coder; and 8) multiplexing the MD bits and core coder bits into a bitstream and sending the bitstream to the decoder.

[0147] In the decoder, the second embodiment involves: 1) obtaining MD bits and core coder bits from the bitstream; 2) decoding the downmix channel using the core audio decoder; 3) dequantizing the MD bits by decoding the low-frequency SPAR MD parameters and high-frequency DirAC MD parameters; 4) obtaining the high-frequency band SPAR MD from the DirAC MD using a D2S conversion routine; 5) performing a filter bank analysis on the decoded downmix channel; 6) generating a SPAR upmix in the filter bank region using the SPAR MD in all frequency bands; and 7) generating a spatial audio output in the decoder. In some embodiments, as part of step 7), filter bank synthesis is performed on the SPAR upmix channel to reconstruct the ambisonic channel in the decoder. In some other embodiments, as part of step 7), DirAC analysis is performed on the upmix channel generated by the SPAR to obtain DirAC MD parameters in all frequency bands and perform a DirAC upmix to a desired output format, including but not limited to HOA2 / HOA3.

[0148] In the decoder, the third embodiment involves: 1) obtaining MD bits and core coder bits from the bitstream; 2) decoding the downmix channel using the core audio decoder; 3) decoding and dequantizing the low-frequency SPAR MD parameters and high-frequency DirAC MD parameters from the MD bits; 4) obtaining high-frequency band SPAR metadata (MD) from the DirAC MD using a D2S conversion routine and obtaining the low-frequency band DirAC MD from the SPAR MD and / or downmix covariance using an MD conversion routine from SPAR to DirRAC (S2D) (described in Section 2.4); 5) performing a filter bank analysis on the decoded downmix channel; 6) generating a SPAR upmix in the filter bank region using the SPAR MD in all frequency bands; and 7) generating a spatial audio output in the decoder. In some embodiments, as part of step 7), filter bank synthesis is performed on the SPAR upmixed channel to reconstruct the ambisonics channel in the decoder. In some other embodiments, as part of step 7), the DirAC MD parameters for all frequency bands, including the low-frequency DirAC MD obtained in step 4), are applied to the SPAR upmix to perform a DirAC upmix to a desired output format, including but not limited to HOA2 / HOA3.

[0149] 1.4.2 Channel-based SPAR-DirAC splitting (same processing across all bandwidths) In some embodiments, a subset of the ambisonic input channels may be reconstructed (residually or parametrically) via SPAR, and some channels are reconstructed by DirAC. Any further upmixes to higher orders are also handled by DirAC. SPAR reconstructs at least enough channels for DirAC analysis to be performed in the decoder. Here, generally, DirAC analysis requires FOA channels (or planar FOA channels in the case of planar). As used herein, residual coding is direct audio coding of the residuals, from which the output channels are reconstructed together with the predicted components from W. Parametric coding is coding of cross-prediction and decorrelation parameters, from which the output is reconstructed together with the predicted components from W, the cross-predicted components of the residuals, and the decorrelated version of W.

[0150] SPAR generally works with B-format representations of input and output ambisonic audio. DirAC reconstructs audio signals in A-format or Equivalent Spatial Domain (ESD) in some cases, and in B-format in others. The following description primarily deals with the latter B-format case. However, similar embodiments are possible for DirAC synthesis in A-format or ESD. For that purpose, SPAR-reconstructed B-format channels can be used to generate a relatively sparse set of DirAC prototype signals in B-format, A-format, or ESD, from which DirAC synthesis generates a denser set of upmixed signals, each of which can drive speakers in a multi-loudspeaker system. Such a multi-loudspeaker system may correspond to an actual loudspeaker setup, such as 7.1.4 or 5.1, or a virtual loudspeaker system, which is an intermediate step to immersive binaural rendering of the synthesized audio signals.

[0151] The various embodiments described above are listed in Table II below and are summarized in the table. [Table 2]

[0152] For HOA3 inputs, the channels are reconstructed according to the following options: FOA, or HOA2, or FOA + quadratic planar channels, or FOA + quadratic + cubic planar channels are reconstructed using SPAR. On the other hand, HOA2 and HOA3, or HOA3, or quadratic nonplanar and HOA3, or quadratic and cubic nonplanar channels are reconstructed using DirAC to reduce computational complexity without compromising quality.

[0153] For bitrates with FOA input where the SPAR downmix channels are less than four: 1) FOA with SPAR is implemented with a DirAC blind upmix to HOA2 / HOA3; or 2) Planar FOA with SPAR, a DirAC upmix to a full FOA, and possibly a blind upmix to HOA2 / HOA3 with DirAC.

[0154] For bitrates where there are four SPAR downmix channels in FOA: 1) FOA using SPAR is implemented along with blind upmixing to HOA2 / HOA3 using DirAC.

[0155] For planar FOA inputs, for bitrates where the SPAR downmix channel is less than 3, WY reconstruction using SPAR is implemented along with upmixing to the planar FOA, and possibly blind upmixing to planar HOA2 / HOA3 using DirAC.

[0156] For bitrates with a planar FOA input and three SPAR downmix channels, WYX reconstruction using SPAR is implemented along with blind upmixing to planar HOA2 / HOA3 using DirAC.

[0157] 1.4.3 Reconstruction of individual channels using SPAR and DirAC in part In connection with the channel-based SPAR-DirAC partitioning technique disclosed in Section 1.4.2, a further category of channels that are parametrically reconstructed in part from both SPAR and DirAC methods can be introduced. The motivation for this is to reduce the reliance on numerous decorrelators in the decoder, which can reduce the complexity metric of the mixture. In this approach, SPAR predictions and cross-predictions are used to reconstruct the majority of a particular parametrically reconstructed signal, and the missing covariances are recovered depending on the DirAC diffusivity.

[0158] 1.4.4 Alternative methods to reduce the use of decorrelators in decoders Instead of adding decorrelation proportional to the decorrelation coefficient, in this embodiment, energy matching of the mutually / predictively parametrically constructed channels is achieved by applying the gain derived from the SPAR coefficient. A particular ambisonic signal S can be parametrically reconstructed as follows:

number

number

[0159] 1.4.5 Combined frequency-based and channel-based splitting In some embodiments, sections 1.4.1 and 1.4.2 are combined to obtain the advantages of merging SPAR and DirAC by performing a combination of frequency-based and channel-based division. In one exemplary implementation, the input to the merged SPAR-DirAC system is an N-channel ambisonics signal. From these N channels, M channels are supplied to the SPAR subsystem, where M ≤ N. In some embodiments, these M channels include FOA channels. In some other embodiments, these M channels include FOA and planar HOA channels. The SPAR can then operate in any downmix configuration based on the operating bitrate, where N is the number of downmix channels. dmx 1≦N dmx The limit is ≤ M. For low frequencies, SPAR calculates SPAR parameters, including prediction, cross-prediction, and decorrelation parameters, based on the method described in Section 1.3, while for high frequencies, DirAC parameters are calculated as described in Section 1.2, and SPAR parameters are estimated from the DirAC parameters as described in Sections 1.5 to 2.4 below. In some embodiments, SPAR also calculates SPAR parameters for high frequencies for a subset of input channels based on the method described in Section 1.3.

[0160] Next, the M channels reconstructed by the SPAR on the decoder side are used by DirAC to reconstruct the representation of the original N-channel input scene.

[0161] An exemplary implementation combining frequency-based and channel-based splitting with an HOA3 input to a SPAR-DirAC merge system is given below.

[0162] 1.4.5.1 Exemplary Encoder Embodiment 1 Figure 8 is a block diagram of an encoder 800 having frequency-based and channel-based splitting between SPAR and DirAC according to one or more embodiments. In this embodiment, the SPAR operates in a 4-channel downmix mode. The input to encoder 800 is an HOA (third-order ambisonics) signal. A DirAC parameter estimator 801 is limited to high frequencies and estimates DirAC parameters calculated according to Section 1.2 based on the FOA channels in the ambisonics input. The estimated DirAC parameters are quantized and encoded 802, and the quantized DirAC MD is converted to SPAR MD 803.

[0163] SPAR analysis and metadata calculation 804 are performed based on the FOA, planar HOA2, and planar HOA3 channels at low frequencies, according to Section 1.3. The SPAR metadata is quantized and encoded 805, and the SPAR MD obtained from the quantized SPAR metadata at low frequencies and the DirAC MD at high frequencies is converted into a downmix matrix 806. The MDFT conversion 807 is applied to the FOA, planar HOA2, and planar HOA3 signals. The MDFT coefficients and downmix matrix are crossfaded and frequency-band mixed using a filter bank mixer 808 to produce a 4-channel downmix. The 4-channel downmix is ​​encoded by one or more core codecs 809 (e.g., an Enhanced Voice Services (EVS) encoder). The encoded SPAR metadata at low frequencies and the encoded DirAC metadata at high frequencies are packed together with the core codec-encoded bits to form the final bitstream 810 output by the encoder 800.

[0164] Encoder 800 is an exemplary embodiment of an encoder that combines DirAC and SPAR. In other embodiments, SPAR and DirAC are combined by frequency division only or channel division only.

[0165] 1.4.5.1.2 Exemplary Decoder Embodiment 1 Figure 9 is a block diagram of a decoder 900 having frequency-based and channel-based splitting between SPAR and DirAC, according to one or more embodiments. In this embodiment, the decoder 900 receives a bitstream 901 (810) and provides the bits encoded by the core codec to one or more core codec decoders 907 (e.g., an EVS decoder). The high-frequency DirAC MD 902 is decoded and then converted to a high-frequency SPAR MD 903 using a DirAC MD to SPAR MD conversion 913. In one embodiment, 913 in the decoder is the same as 803 in the encoder. The SPAR MD in the bitstream is decoded to reconstruct the low-frequency SPAR metadata 904. The SPAR upmix matrix 905 is generated using the low-frequency SPAR metadata 904 extracted from the bitstream 910 and the high-frequency SPAR metadata 903 converted from the high-frequency DirAC metadata. The downmix channel is reconstructed by one or more instances of the core decoder 907 and converted into a frequency-banded domain by the filter bank 908 (e.g., CLDFB filter bank, quadrature mirror filter bank (QMF), etc.).

[0166] In some embodiments, the main downmix channel is input to a decorrelator 909, and the output of the decorrelator 909, along with the upmix matrix, is input to a SPAR upmix unit 906 to reconstruct the FOA, Planar HOA, and Planar HOA channels. Decorrelation can be performed in the time domain or in a frequency-banded domain (e.g., the CLDFB domain). The decorrelator can generate a time-domain decorrelation output and then convert it to the frequency-banded domain, or convert the input to the frequency-banded domain and generate a decorrelation output in the frequency-banded domain. The output channel of 906 is fed to a DirAC parameter estimator 910, which estimates low-frequency DirAC metadata based on the reconstructed FOA signal in the frequency-banded domain. The DirAC upmixer 911 uses the low-frequency and high-frequency DirAC metadata to upmix the FOA, Planar HOA2, and Planar HOA3 channels into 16 HOA3 channels. This is the frequency-banded domain representation of the original 16-channel HOA3 input to encoder 200. The combiner 312 (e.g., CLDFB combiner) synthesizes / renders the frequency-banded domain representation of the 16-channel HOA3 into a time-domain representation for playback on various audio systems with different speaker configurations and capabilities.

[0167] 1.4.5.2.1 Exemplary Encoder Embodiment 2 Figure 10 is a block diagram of an alternative encoder 1000 having frequency-based and channel-based division between SPAR and DirAC according to one or more embodiments. In this embodiment, the input to encoder 1000 is an HOA signal. DirAC parameters are estimated (1001), quantized and encoded (1002). DirAC parameter estimation is limited to high frequencies and is performed according to Section 1.2 based on the FOA channel. SPAR analysis and metadata calculation (1004) and quantization and encoding (1005) are performed at low frequencies according to Section 1.3 based on the FOA, planar HOA, and HOA channels plus zero or more non-planar channels (e.g., non-planar channels).

[0168] For high frequencies, SPAR analysis and parameter estimation are performed for non-FoA channels according to Section 2.2.7.2 (this is not performed in system 200). In this embodiment, the SPAR operates in 4-channel downmix mode, and the SPAR FoA metadata at high frequencies is estimated based on the DirAC metadata using the method described in Section 2.2 to obtain the SPAR downmix matrix for all frequencies.

[0169] The quantized and encoded SPAR metadata is used to generate a downmix matrix 1007. An MDFT transformation 1006 is applied to the FOA, planar HOA2, and planar HOA3 signals. The MDFT coefficients and downmix matrix are frequency-band mixed with a crossfade 1008 to generate a 4-channel downmix. The 4-channel downmix is ​​encoded by one or more core codecs 1009. The SPAR metadata encoded at low frequencies for the FOA channel and at full frequencies for the HOA channel, along with the DirAC metadata encoded at high frequencies, are packed together with the core codec-encoded bits to form the final bitstream 1010 output by the encoder 1000.

[0170] The downmixed channels are encoded by one or more core codecs (e.g., EVS) 1009. For FOA channels, SPAR metadata is encoded for low frequencies and DirAC metadata is encoded for high frequencies, but for non-FOA channels, SPAR metadata is encoded for the entire frequency range and packed together with the core codec encoded bits to form the final bitstream 1010 output by encoder 1000. In this embodiment, the SPAR metadata calculation for HOA2 and HOA3 channels at high frequencies is performed according to the method described in Section 2.2.7.2. Furthermore, in this embodiment, the SPAR metadata calculation 1004 for HOA2 and HOA3 channels at high frequencies, according to the method described in Section 2.2.7.2, relies on the SPAR MD for the FOA channels at high frequencies, which is estimated from the DirAC MD at high frequencies 303.

[0171] In Embodiment 2, the conversion from DirAC MD to SPAR MD is performed only for the FOA channel, and full-band SPAR MD is used for any HOA channel handled by SPAR. In general, any number of non-planar HOA channels can be handled by SPAR. In Embodiment 2, only one non-planar HOA channel was added. Also, although these embodiments focus on four downmix channels (Ndmx=4), any number of transport channels (e.g., 1 to 16) are possible.

[0172] 1.4.5.2.2 Exemplary Decoder Embodiment 2 Figure 11 is a block diagram of an alternative decoder 1100 having frequency-based and channel-based splitting between SPAR and DirAC, all arranged according to embodiments described herein.

[0173] In this embodiment, the decoder 1100 receives the encoded bitstream 1104 and provides the bits encoded by the core codec to one or more core decoders 1105. The high-frequency DirAC MD 1102 is decoded and then converted to the high-frequency SPAR MD 1103 using the DirAC MD to SPAR MD converter 1113. In one embodiment, the DirAC MD to SPAR MD converter 1113 in the decoder is the same as the DirAC MD to SPAR MD converter 903 in the encoder. The SPAR MD 1104 corresponding to the FOA and planar HOA and zero or more non-planar HOA channels are decoded and fed into the SPAR mixture matrix 1106. Missing SPAR MD 1103 for the high-frequency FOA channels are inferred from the DirAC MD in the same way as in the encoder 900. The SPAR upmix matrix 1106 is generated using the SPAR MD 1104 extracted from the bitstream 1110 and the high-frequency SPAR MD 1103 converted from the high-frequency DirAC MD. The downmix channels, reconstructed by one or more instances of the core decoder 1105, are converted to a frequency-banded region with the help of filter bank analysis 1107, and the upmix matrix 1106 is applied to reconstruct the FOA, planar HOA2, planar HOA3 channels and zero or more non-planar channels.

[0174] The decoded downmix channel outputs from one or more core decoders 1105 are fed to a decorrelator(s) 1109, and the output of the decorrelator 1109, along with the upmix matrix, is input to a SPAR upmix unit 1108 to reconstruct the FOA, Planar HOA2, and Planar HOA3 channels. The decorrelator can be implemented in the time domain or the frequency-banded domain (e.g., CLDFB domain). The decorrelator can generate a time-domain decorrelation output and then convert it to the frequency-banded domain, or convert the input to the frequency-banded domain and generate a decorrelation output in the frequency-banded domain. The output channels of 1108 are fed to a DirAC parameter estimator 1110, which estimates the DirAC metadata at low frequencies based on the reconstructed FOA signal in the frequency-banded domain and uses the DirAC parameters at high frequencies extracted from the bitstream 1101. Alternatively, the DirAC upmixer 1108 may estimate the DirAC parameters over the entire frequency range based on the FOA signal in the frequency band domain (e.g., CLDFB domain) and ignore the high-frequency DirAC parameters from the bitstream 1101.

[0175] The DirAC upmixer 1111 uses DirAC metadata from 1110 and 1102 to convert the FOA, planar HOA2, planar HOA3, and zero or more non-planar channels into an HOA3 output, which is a frequency-bandwidth domain (e.g., CLDFB domain) representation of the original 16-channel HOA3 input to the encoder 900. The combiner 1112 (e.g., CLDFB combiner) synthesizes / renders the 16-channel HOA3 frequency-bandwidth domain representation into a time-domain representation for playback on various audio systems with different speaker configurations and capabilities. Note that the output of the decorrelator 1109 is the CLDFB domain, and this covers embodiments where CLDFB analysis follows time-domain decorrelation and CLDFB analysis with CLDFB domain decorrelation.

[0176] 1.4.6 IrAC Upmix Channel Ir When estimating higher-order channels from first-order channels using the DirAC approach, directional panning in the higher-order channels upmixed using the DirAC approach can be performed using DOA angles and spherical harmonic responses. However, the addition of diffusivity and decorrelation to these higher-order channels must be handled carefully, as it has been found that too much decorrelation can degrade audio quality, and too little decorrelation can cause spatial collapse.

[0177] The following describes embodiments for adding diffusivity to higher-order channels upmixed by the DirAC approach.

[0178] 1.4.6.1 Apply uniform decorrelation to all HOA upmixed channels In some embodiments, Nd decorrelated channels uncorrelated with respect to the W channel are calculated, where Nd is the number of HOA channels to be upmixed from the FOA channel by DirAC. ψ (diffusivity) is calculated using one of the methods described in this paper, and then the following is calculated: ψ Diffuseness factor (i)=ψ*Norm(i),

[24] Here, i is the channel index, and Norm is the corresponding normalization factor calculated according to a given ambisonic normalization, e.g., SN3D normalization. Diffuseness factor (i) is applied to the i-th decorrelated channel to obtain the diffusion component for the corresponding HOA channel. In one exemplary embodiment, if the input to the DirAC upmixer is an FOA channel (4 channels) and the HOA3 channel is upmixed from the input FOA channel, the number of decorrelated outputs required is 12 (Nd=12). The upmixed HOA channel H(i) can be expressed as follows: H(i) = energy_Ratio_factor * Resp i *W+DiffuseHack factor (i)*D i(W)

[25] Here, Resp i This is the spherical harmonic response for the corresponding channel index, and the DOA angle θ D It is calculated using θ, where θ D This can be expressed in terms of azimuth and elevation. The energy_Ratio_factor can be calculated as (1-ψ). D i (W) is the i-th decorrelated channel.

[0179] The above approach may result in excessive decorrelation, potentially causing the reconstructed scene to become more diffuse than necessary. Furthermore, generating too many decorrelator outputs and scaling them to achieve the desired level of diffuseness can be computationally expensive.

[0180] 1.4.6.2 Adding Directional Decorrelation to All Upmixed Channels In this embodiment, directional diffusion information is transmitted from the encoder to the decoder. The decoder uses this directional diffusion information to add only the desired amount of decorrelation to the upmixed HOA channels. This method is applicable when the input to the encoder is HOA, and due to bitrate and complexity limitations, only a few selected channels are reconstructed using SPAR, while the remaining channels are upmixed using DirAC. In one exemplary implementation, the encoder can calculate directional diffusion using P (decorrelation) coefficients calculated by SPAR in Section 1.3. This method uses additional information transmitted from the encoder to the decoder.

[0181] 1.4.6.3 Adding decorrelation to selected upmixed channels In this embodiment, the addition of diffusivity is limited to a small number of selected channels in order to maintain the overall diffusivity within the desired limits. This method also reduces computational complexity. The selection of channels for the addition of diffusivity can be static or dynamic based on signal characteristics.

[0182] 1.4.6.3.1 Static selection of channels In this embodiment, decorrelation is added to a select few HOA channels. These channels are selected based on perceptual importance. In one exemplary embodiment, if the FOA and planar HOA channels are reconstructed by SPAR and only the non-planar HOA channels are upmixed using DirAC to obtain HOA output in ACN-SN3D format, channel indices 6, 10, 12, and 14 (channel indices ranging from 0 to 15) may be selected to add decorrelation. This method does not require any additional information to be sent to the decoder.

[0183] 1.4.6.3.2-Channel Dynamic Selection In this embodiment, directional diffusion information is calculated in the encoder and transmitted to the decoder to select channels to which diffusion should be added during upmixing. This embodiment is applicable only when the input to the encoder is HOA. Only channels for which the amount of decorrelation required is higher than a first threshold are selected in the DirAC decoder to have decorrelation added. In one exemplary implementation, the encoder calculates directional diffusion using P (decorrelation) coefficients calculated by SPAR in Section 1.3, compares the P coefficient values ​​to a first threshold, and encodes channel indices with P coefficients higher than the first threshold. These indices are read by the decoder. If the number of channel indices exceeds a second threshold, a limited number of indices may be selected based on the P coefficient values ​​and the perceptual importance of a given channel. This embodiment requires additional information to be transmitted from the encoder to the decoder.

[0184] To perform frequency-based splitting as described in Section 1.4.1, an efficient mechanism is desired to convert DirAC metadata to SPAR in the DirAC frequency band and SPAR metadata to DirAC in the SPAR frequency band, so that DirAC and SPAR metadata can be reconstructed in all bands when needed to perform upmixing or downmixing. The following is an exemplary embodiment of converting DirAC metadata to SPA\\R and SPAR metadata to DirAC.

[0185] 1.5 Converting from DirAC to SPAR In some embodiments, an approximation of the input covariance matrix is ​​calculated based on quantized DirAC MD parameters (azimuth (Az), elevation (El), and diffusivity). In this paper, Az and El are the DOA angle θ. D It is also called [another name].

[0186] 1.5.1 Formulas In some embodiments, the model covariance block calculates the covariance matrix and predictive coefficients from the DirAC DOA and diffusivity as follows:

number

[0187] 1.5.2 Exemplary Covariance Calculation In some embodiments, the covariance is calculated as follows:

number

[0188] In the above equation, E is an approximation of the overall signal energy (see

[33] below). This is obtained by adding rough estimates of the directional energy and the diffused energy.r Assuming that r is the actual bin sample of the W channel in the MDFT domain, the energy corresponding to each bin is calculated as follows. E(f)=w r *w r

[31]

[0189] Next, by applying the filter bank response of each band, the energy is converted into frequency-banded power. The frequency-banded energy in each band is extrapolated to calculate the overall signal energy as follows. E=E*(Resp w 2 +Resp x 2 +Resp y 2 +Resp z 2 )

[32]

[0190] The diffused energy components are added as follows. E=E*(1+ψ 2 )

[33]

[0191] In some embodiments, when covariance smoothing is turned off, the calculated covariance as described above is used to calculate the SPAR coefficient as usual.

[0192] 2.0 Other Embodiments 2.1 DirAC MD calculation [[ID=4�]] 2.1.1 Improved calculation of DirAC diffusivity DirAC requires temporal smoothing to calculate the diffusivity parameter. In some embodiments, simple parameter averaging is performed over 160 ms (Equation 12 from Section 1.2.2.2).

[0193] In some embodiments, covariance smoothing and / or transient detector-ducker algorithms for SPAR can be used to improve the calculation of DirAC diffusivity parameters. For example, the covariance smoothing algorithm for SPAR described in PCT application PCT / 2020 / 044670, filed July 31, 2020, for "System and Method for Covariance Smoothing," can be adapted to weight recent audio events more heavily than even earlier events, and this can be done differently in each frequency band. This may be advantageous over a simple averaging operation. Using transient detection and ducking, diffusivity values ​​can be instantaneously reduced during short transients without interfering with the long-term smoothing process.

[0194] Because time smoothing introduces long-term time dependence and smoothing over time, in another embodiment, differential coding can be used to reduce the MD bitrate and improve frame loss resilience.

[0195] 2.1.2 Calculation of DirAC metadata in the frequency-banded covariance region Based on the DirAC analysis captured in Section 1.0, in some embodiments, instead of calculating the DirAC MD in the FFT (Fast Fourier Transform) or MDFT domain and then converting it to the frequency-banded domain, the DirAC MD can be calculated based on the frequency-banded covariance matrix of the input.

[0196] In some embodiments, as shown in Section 1.0, the SPAR metadata can be calculated based on the frequency-banded covariance of the input.

[0197] In some embodiments, by calculating both SPAR and DirAC metadata from the input covariance, a better conversion from SPAR to DirAC and from DirAC to SPAR MD in the desired band becomes possible. It is also computationally efficient. The following is an example of a method for calculating DirAC MD from the input covariance. 1. Calculate an N*N frequency banded covariance matrix. Here, N is the number of input channels. 2. Smooth the covariance matrix as described in Section 2.1.1. 3. Calculate the reference power as the trace of the covariance matrix. 4. R wx ,R wy ,R wz Calculate the intensity as. Here, R wx ,R wy ,R wz is the covariance between the W channel and the X, Y, Z channels. 5. Calculate the intensity norm

Number

Number

[0198] Similarly, the diffusivity calculation can be performed as follows based on the frequency banded covariance matrix.

[0199] For diffusivity, first calculate the reference power E and intensity I of the input signal in a given frequency band. E = R ww + R yy + R xx + R zz

[34] I x = R wx , I y=R wy , I z =R wz

[35]

[0200] As stated in Section 2.1.1, if the covariance has already been smoothed, the diffusivity can be calculated as follows:

number

[0201] In some embodiments, before calculating diffusivity, E and I are further averaged using the long-term averaging filters given below. E a =(1-f e )*E a-1 +f e *E

[39] I a =(1-f i )*I a-1 +f i *I

[40]

[0202] Here, E a and I a These are the long-term averages for energy and intensity, respectively, and these values ​​are used in place of E and I in the calculation of the diffusivity formula

[36] . Factor f in

[39] and

[40] e and f i This is an example of a smoothing factor.

[0203] 2.1.3 Improvement of Reference Power (E) Calculation In some embodiments, alternative methods can be used to calculate a baseline power that gives a better estimate of diffusivity and leads to a better estimate of the SPAR coefficient when derived from the DirAC coefficient.

[0204] Regarding diffusion, first, the reference power E and intensity I of the input signal are calculated within a given frequency band. Bar=R ww +R yy +R xx +R zz

[41] I x =R wx , I y =R wy , I z =R wz

[42] Here, R ij This is the covariance between the i-th and j-th channels.

[0205] The reference power is calculated as follows: E=max(R ww ,0.5*E)

[43]

[0206] The E calculated in

[43] provides a better estimate of diffusivity and SPAR coefficient when the W channel energy is greater than 0.5*E. Diffusivity is calculated as follows:

number

[0207] The diffusivity can be limited as follows: ψ=max(0,min(1,ψ))

[45] Energy ratio = 1 - ψ

[46]

[0208] The SPAR coefficient can be calculated from the DirAC coefficient using one of the methods described in this paper.

[0209] 2.2 Improvements to DirAC to SPAR MD conversion 2.2.1 Alternative methods for calculating covariance / SPAR MD from DirAC In some embodiments, the passive prediction coefficient is also Respi *Resp j It can be calculated as follows: Here, i and j can be w, y, y, z, which should be analogous to the direction vector dv for a given side channel. Thus, the prediction coefficient is I in the frequency-banded region where the variance of the W channel is norm If it is smaller, it will be closer to the actual SPAR prediction coefficient. In some embodiments, the variance of the W channel is I norm If it is greater, an additional parameter R is used for better estimation of the prediction coefficient. ww / I norm This can be sent to the decoder. In some embodiments, the prediction coefficient can also be calculated directly from the DirAC metadata.

[0210] 2.2.2 Quantization of DirAC metadata In some embodiments, the SPAR MD is calculated based on a quantized DirAC MD.

[0211] 2.2.3 General Reconstruction of SPAR Coefficients from DirAC Metadata for Any Downmix Configuration In some embodiments, the input covariance R is a 4x4 matrix calculated based on the DirAC parameter as follows: R ij =(1-cψ)*E*Resp i *Resp j If i != j, R ww =E*Res w *Resp w , R ii =(1-cψ)*E*Resp i 2 +Q i *ψ*E i!=w

[47] Here, i and j can be w, x, y, and z, and Resp w =Y 0,0 (θ D ), Resp x =Y 1,1 (θ D ), Resp y =Y 1,-1 (θD ), Resp z =Y 1,0 (θ D ) is a spherical harmonic function, Q i And c are constants in the range of 0 and 1. c=1 and Q i Setting = 1 results in an equation similar to the one described in Section 1.5, in which case both the encoder and decoder have prior knowledge of these constants. In some embodiments, Q i And c can be dynamically calculated based on the actual input covariance matrix and the above approximation of the input matrix from the DirAC parameters.

[0212] If the SPAR coefficients are normalized with respect to covariance, then the SPAR coefficients derived from the input covariance R are equal to the SPAR coefficients derived from E*R, where E can be the variance of the W channel, the overall signal energy, or any arbitrary constant.

[0213] In some embodiments, the normalized covariance matrix R_norm is derived based solely on the DirAC parameters. R_norm is a 4x4 covariance matrix for the FOA channel and is an approximation of the actual normalized input covariance matrix, where the actual input covariance matrix is: R in =UU T The 4x4 covariance matrix for the FOA input channels, where U=[WXYZ] T FOA Input It is given as follows. R_norm can be calculated based solely on the DirAC parameters given below: R norm ij =(1-cψ)*Resp i *Resp j If i != j R_norm ww =Resp w *Resp w And R_norm ii =(1-cψ)*Resp i 2+Q i If i != w channel index

[48]

[0214] SPAR coefficients, including prediction, cross-prediction, and decorrelation coefficients, are normalized covariance R_norm as disclosed in Section 1.3. ij It is calculated from.

[0215] 2.2.3.1 Example of reconstructing predictions and decorrelation coefficients directly from DirAC metadata for a single-channel downmix. From the normalized covariance matrix above, the SPAR coefficients can be calculated as follows, based on the calculations in Section 1.3.

[0216] The prediction coefficient is calculated as follows: PR x / y / z =(1-cψ)*Resp x / y / z

[49]

[0217] For a one-channel downmix, the decorrelation coefficient is calculated as follows: P x =sqrt((1-cψ)*Resp x 2 +Q x ψ-(1-cψ) 2 *Resp x 2 )

[50] P y =sqrt((1-cψ)*Resp y 2 +Q y ψ-(1-cψ) 2 *Resp y 2 )

[51] P z =sqrt((1-cψ)*Resp z 2 +Q z ψ-(1-cψ) 2 *Resp z 2 )

[52] Here, the decorrelation coefficient depends on the spherical harmonic response. To avoid this dependence, 2.2.4 can be used.

[0218] 2.2.4 Another variation of the general reconstruction of SPAR coefficients from DirAC metadata for any downmix configuration In this embodiment, the actual input covariance R in The 4x4 covariance matrix R, which is an approximation of , is calculated based on the DirAC parameter as follows: Here, the elements of the matrix are approximated as follows: R ij =(1-cψ)*E*Resp i *Resp j If i != j R ww =E*Res w *Resp w R ii =(1-cψ) 2 *E*Res i 2 +Q i *E*(1-(1-cψ) 2 ) When i != w channel index

[53] Here, Resp w =Y 0,0 (θ D ), Resp x =Y 1,1 (θ D ), Resp y =Y 1,-1 (θ D ), Resp z =Y 1,0 (θ D ) is a spherical harmonic function, Q i And c is a constant in the range of 0 and 1 (for example, c=1, Q i =1 / 3). In this case, both the encoder and decoder have prior knowledge of these constants. In some implementations, Q i And c can be dynamically calculated based on the actual input covariance matrix and the above-mentioned approximation of the input matrix from the DirAC parameters.

[0219] Assuming the SPAR coefficients are normalized, the SPAR coefficients can be derived from R, just as the SPAR coefficients derived from E*R are derived from R, where E can be the variance of the W channel only, the overall signal energy, or any arbitrary constant.

[0220] The elements of the normalized 4x4 covariance matrix for the FoA channel are derived based solely on the DirAC parameter. R norm ij =(1-cψ) 1 *Resp i *Resp j If i != j R norm ww =Resp w *Resp w R_norm ii =(1-cψ) 2 *Resp i 2 +Q i (1-(1-cψ)2) i!=w channel index

[54]

[0221] The SPAR coefficients, including the prediction, cross-prediction, and decorrelation coefficient, are calculated from R_norm as disclosed in Section 1.3.

[0222] 2.2.4.1 Example of reconstructing predictions and decorrelation coefficients directly from DirAC metadata for a single-channel downmix. Based on the normalized covariance described above, the SPAR coefficient can be calculated as follows, based on the calculations in Section 1.3.

[0223] The prediction coefficient can be calculated as follows: PR x / y / z =(1-cψ)*Resp x / y / z

[55]

[0224] For a one-channel downmix, the decorrelation coefficient can be calculated as follows: P x =sqrt(Q x (1-(1-cψ) 2 )

[56] P y =sqrt(Q y (1-(1-cψ) 2 )

[57] P z =sqrt(Q z (1-(1-cψ) 2)

[58] Here, the decorrelation coefficient does not depend on the spherical harmonic response, but only on diffusivity and a few constants.

[0225] Calculation of the constant "c" - Solution 1 In one exemplary implementation, to further improve the prediction coefficient, the passive W prediction coefficient is PR i =sqrt(1-ψ)*Resp i Here, i can be x, y, or z.

[59] (1-ψ) can be set such that this occurs.

[0226] Based on equation

[53] , this gives the value of c as follows:

number

[0227] In one embodiment, to improve the SPAR coefficient calculated from DirAC MD, the actual normalized input covariance R_norm is used. in The 4x4 covariance matrix R_norm, which is an approximation of , is calculated based on the DirAC parameter as follows: Here, the elements of the matrix are approximated according to

[54] and

[61] below.

number

[0228] For a one-channel downmix, the decorrelation coefficient is calculated based on equations

[62] to

[64] as follows: P x =sqrt(Q x *ψ)

[62] P y =sqrt(Q y *ψ)

[63] P z =sqrt(Q z *ψ)

[64]

[0229] In some embodiments, Q x Qy Q z The value can be set to 1 / 3.

[0230] Calculation of the constant "c" - Solution 2 In another exemplary embodiment, c can be calculated as follows:

number

[0231] Intensity normalization is,

number

[0232] Substituting this value of c into the formula for calculating the prediction coefficient,

number

[0233] This prediction coefficient

[66] is similar to the calculation of the passive prediction coefficient disclosed in Section 1.3.1.1. For this solution, the value of c can be sent to the decoder.

[0234] 2.2.5 Energy Correction for DirAC-Based Downmixes As disclosed in Section 1.5.2 and in the solutions described in Sections 2.2.3 and 2.2.4, the calculation of covariance from DirAC metadata (MD) is performed by the signal w = x + y + z

[67] We assume that the data is perfectly SN3D normalized to this extent. Here, w, x, y, and z are the variances of the W, X, Y, and Z channels, respectively.

[0235] This assumption does not apply to real-world FoA capture (e.g., in overtalk situations, diffuse background noise capture, etc.). The above method results in spatial collapse, especially when the number of downmix channels is limited to one.

[0236] To prevent spatial compression, the downmixed signal is scaled so that the upmixed signal is energy-matched to the input. Energy compensation can be applied. An exemplary implementation of energy compensation using a single-channel downmix is ​​shown below.

[0237] Actual input covariance matrix R inNxN Here, N is the number of input channels,

number

[0238] The normalized actual input covariance matrix R_norm_in NxN It is calculated as follows:

number

[0239] Normalized covariance estimation R_norm based on DirAC metadata NxN This is calculated according to one of the techniques described in sections 1.5.2, 2.2.3, and 2.2.4.

[0240] The scaling factor is obtained as follows:

number

[0241] The SPAR downmix matrix and SPAR coefficients (including predictive coefficients, cross-predictive coefficients, and decorrelation coefficients) are computed using the DirAC-estimated normalized input covariance matrix as disclosed in Section 1.3.

[0242] Downmix matrix 1xN The downmix matrix is ​​scaled by the scale calculated in equation

[70] in Section 2.2.5. The actual downmix matrix is ​​Downmix_act 1xN year, downmix_act 1xN =scale*Downmix 1xN

[71]

[0243] In one exemplary embodiment, for a 1-channel downmix, Donwmix 1xN This is given by equation

[72] as follows: Downmix 1xN =[F W ,F Y ,F z ,F x

[72] Here, F W ,F Y ,F z ,F x This is the gain used to mix the Y, Z, and X channels into the W channel, respectively, to form a downmix channel.

[0244] After scaling by the "scale" value, the downmix channels are calculated as follows: W'=scale*(F W *W+F Y *Y+F Z *Z+F X *X)

[73]

[0245] In another exemplary implementation, F W =1, F Y =F z =F x = 0, and W' = scale * W. W ,F Y ,F z ,F x Another exemplary implementation involving the calculation of is described in Section 2.3.

[0246] Metadata parameters are not modified by this scaling. The encoder encodes the metadata parameters, and the scaled downmix and bitstream are sent to the decoder.

[0247] The decoder decodes the scaled downmix channel W' and spatial parameters, including prediction and decorrelation parameters, and applies these prediction and decorrelation parameters to reconstruct the original input scene as follows:

number

[0248] This approach scales the reconstructed signal by a scaling factor calculated in equation

[70] of this section, thereby energy matching the reconstructed scene to the input without sending additional parameters to the bitstream.

[0249] 2.2.6 Extrapolation of Directional Spreading in the DirAC Band DirAC-based covariance estimation assumes uniform diffuseness in all directions. This is not true for real-world signals, such as overtalk scenarios. Adding directional information on top of the diffuseness parameter calculated by equation [7] in Section 1.2 results in the encoding of additional metadata in the bitstream. SPAR provides directional diffuseness information in the metadata, and high-bandwidth directional information can be extrapolated using low-bandwidth directional information.

[0250] In one exemplary embodiment for an FOA input with a single channel downmix, if the SPAR encodes up to a frequency range of 6 kHz and the DirAC parameters are transmitted over a frequency range of 6 to 24 kHz, the directional information in the SPAR frequency band can be extracted as follows:

number

[0251] This directional information can be used in the high-frequency band when calculating the downmix using DirAC parameters. An example of estimating the normalized covariance matrix from DirAC metadata with directional spread is as follows: R_norm is a 4x4 matrix of FOA channels calculated as follows:

number

[0252] The downmix matrix and SPAR coefficients, including prediction, cross-prediction, and decorrelation coefficients, are calculated from R_norm as disclosed in Section 1.3. Exemplary calculations of prediction and decorrelation coefficients for a one-channel downmix are given in

[55] to

[58] . The downmix matrix can be further scaled according to

[70] to better energy match the reconstructed ambisonic signal at the decoder with the ambisonic signal at the encoder input.

[0253] 2.2.7 Metadata conversion from DirAC to SPAR for HoA channels 2.2.7.1 Estimation of HoA input covariance matrix from DirAC parameters In this method, the DirAC parameter is used to estimate the input covariance matrix.

[0254] The NxN covariance R is calculated based on the DirAC parameters, where N is the number of input channels in the HOA signal, and R is an approximation of the actual input covariance matrix. In some embodiments, the covariance R can be calculated as follows: R ij =(1-cψ)*E*Resp i *Resp j If i != j, R ww =E*Res w *Resp w R ii =(1-cψ) 2 *E*Res i 2 +Q i *E*(1-(1-cψ) 2 ) When i != w channel index.

[82] Here, Resp i Q is a spherical harmonic function, i c and Q are constants in the range of 0 and 1, for example, for a first-order channel, i.e., 0 ≤ i ≤ 3, c = 1 and Q i = 1 / 3, and for the second-order channel, i.e., 4 ≤ i ≤ 8, Q i = 1 / 5, and for the third channel, i.e., 9 ≤ i ≤ 15, Q i = 1 / 7, and in this case, both the encoder and decoder have prior knowledge of these constants. In some embodiments, Q i And c are dynamically calculated based on the actual input covariance matrix and the aforementioned approximation of the input matrix from the DirAC parameters.

[0255] Assuming the SPAR coefficients are normalized, the SPAR coefficients derived from R are equal to the SPAR coefficients derived from E*R, where E can be the dispersion of the W channel only, the overall signal energy, or any arbitrary constant.

[0256] The covariance R in equation

[81] is normalized, and the elements of this NxN normalized covariance matrix R_norm are derived based solely on the DirAC parameter as follows:

number

[0257] The SPAR coefficients, including the prediction, cross-prediction, and decorrelation coefficients, are calculated from R_norm as disclosed in Section 1.3.

[0258] 2.2.7.2 Improving the spatial resolution of the DirAC to SPAR conversion by restricting DirAC covariance estimation to FOA channels only It has been observed that covariance estimation for HOA channels from DirAC parameters is not optimal when the HOA channels contain critical information. Ambience loss has been observed when estimating the entire NxN covariance matrix (or all HOA SPAR parameters) from DirAC parameters. For such HOA signals, a different approach is desirable. The following describes several embodiments for DirAC-to-SPAR conversion with improved spatial resolution.

[0259] 2.2.7.2.1 By independently calculating and encoding SPAR HoA parameters In this method, the DirAC parameter estimates the input covariance matrix for the FOA channel only, and then the SPAR parameter corresponding to the FOA channel is calculated from that estimate. This is done by the method described in sections 2.2 and 2.2.4.

[0260] SPAR parameters, including the prediction coefficient, cross-prediction coefficient, and decorrelation coefficient for the HoA channel, are calculated independently based on the actual covariance matrix of the input signal, using the method described in Section 1.3.

[0261] This method requires encoding the SPAR HoA parameters into a bitstream for all frequencies.

[0262] 2.2.7.2.2 Alternative calculation of SPAR HOA parameters based on DirAC estimated FOA This method is applicable to SPAR modes where the number of downmix channels is less than the number of input channels to the SPAR, i.e., when the SPAR has mutual prediction and / or decorrelation coefficients to encode for the HOA channels. In this method, the DirAC parameter is used to estimate the input covariance matrix for the FOA channels only, and then to compute the SPAR parameter corresponding to the FOA channels. This is done by the method described in sections 2.2 and 2.2.4.

[0263] Calculation of HOA prediction coefficients The SPAR prediction coefficients for the HOA channel are calculated independently based on the actual covariance matrix of the input signal, using the method described in Section 1.3.

[0264] Calculation of HOA mutual prediction coefficients Section 1.3 shows that the mutual prediction coefficients in SPAR MD depend on the predicted side channel or residual channel in the downmix. Furthermore, the residual channel in the FOA component of the ambisonics input depends on the SPAR MD derived from the DirAC MD in a set of frequency bands. Thus, the mutual prediction coefficients in the HOA channel may depend on the DirAC MD in the FOA channel, and it has been observed that calculating the mutual prediction coefficients in the HOA channel based on the DirAC MD in the FOA channel and the SPAR MD in the FOA and HOA channels leads to a better estimate of these coefficients. In one exemplary implementation, the prediction coefficients for the HOA channel (4 to N) are calculated from the actual input covariance matrix as described in Section 1.3. These prediction coefficients are quantized based on the quantization strategy. The DirAC-estimated FoA prediction coefficients, along with the SPAR-estimated HOA quantized prediction coefficients, are used to generate the downmix matrix as described in Section 1.3. A post-prediction covariance matrix is ​​calculated from the actual input covariance and the downmix matrix calculated above. Next, as described in Section 1.3, the cross-prediction coefficients are calculated from the post-prediction matrix.

[0265] Calculation of HOA decorrelation coefficient It has been observed that calculating the HOA decorrelation coefficient directly from the ambisonics input covariance, as described in Section 1.3, without relying on the DirAC MD in the FOA channel, improves the estimation of the decorrelation coefficient and yields the desired amount of decorrelation in the reconstructed HOA channel in the decoder. This helps reduce audio artifacts that can occur due to too much decorrelation and also avoids spatial collapse resulting from too little decorrelation. In one exemplary implementation, the prediction coefficients corresponding to all side channels are first calculated from the actual input covariance matrix, as described in Section 1.3, where the side channels in ambisonics are all channel inputs except the W channel. The calculation of the decorrelation coefficient from the prediction coefficients and covariance matrix is ​​then the same as described in Section 1.3. This method encodes the SPAR HoA parameters in the bitstream for all frequencies.

[0266] 2.3 Active W Downmix Based on DirAC Metadata 2.3.1 Based on DirAC-based covariance estimation From the DirAC metadata, the input covariance can be estimated as a DirAC metadata-based input signal (4×4) covariance matrix estimate, as given in Section 2.2.3 or 2.2.4.

number

[85] Alternatively, S can be calculated as follows, as given in Section 2.2.4: S ij =(1-cψ)*E*Resp i *Resp j If i != j S ii =(1-cψ) 2 *E*Res i 2 +Q i *E*(1-(1-cψ) 2 ).

[86]

[0267] One possible approach to performing an active downmix based on the above covariance matrix is ​​to have the following prediction matrix:

number

number

[89] are not shown because they are not important for calculating the active downmixing gain.

[0268] By setting ^u*^r and minimizing ^r, we obtain the linear equation given by the following equation.

number

[90] , E cancels out in the numerator and denominator, and g can be calculated directly from the DirAC metadata on both the encoder and decoder sides.

[0269] The actual downmix matrix and post-scaling for a 1-channel downmix are given as follows:

number

number

[0270] The scaled prediction coefficients are calculated as follows: g'=g / r

[93] Here, g'^u=[pr x ;pr y ;pr z ] is the active prediction coefficient.

[0271] The decorrelation coefficient is calculated as follows: Post_prediction [4x4] =Pred*R [4x4] *Pred' Here, Pred is the prediction matrix given in

[91] , and the decorrelation coefficient is calculated as follows: Post_prediction [4x4] It is calculated from.

number

[0272] The calculation of the active W downmix channel from the FOA input [W, Y, Z, X] is given as follows: W'=scale*r*(W+F Y *Y+F Z *Z+F X *X) F Y =F*Resp y F X =F*Resp x F Z =F*Resp z

[0273] The calculation of scale is given in

[70] , the calculation of another scale factor r is given in

[92] , W' is encoded in the core coder, DirAC MD is encoded, and these encoded bits are sent together to the decoder.

[0274] The inverse prediction matrix in the decoder is given as follows:

number

[0275] The reconstruction of the FOA channel in the decoder is as follows: W out =W'(1-f s (pr x 2 +pr y 2 +pr z 2 ))

[96] X out =pr x W'+p x D1(W')

[97] Y out =pr y W'+p y D2(W')

[98] Z out =pr z W'+p z D3(W')

[99] Here, prx, pry, and prz are prediction parameters calculated from the DirAC MD given in

[90] , px, py, and pz are decorrelation parameters calculated from the DirAC MD given in

[94] , D1(W'), D2(W'), and D3(W') are three decorrelation channels decorrelated with respect to W', and fs is the scaling constant used in

[92] .

[0276] 2.4 Metadata Conversion from SPAR to DirAC To perform upmixing to a desired output format in a decoder, it may be desirable to convert SPAR MD to DirAC MD in a set of frequency bands so that DirAC MD is available in all required frequency bands. Direct conversion from SPAR MD to DirAC MD also reduces complexity. In one exemplary implementation, it is possible to derive the direction vector dv from the prediction coefficients.

number

[0277] 2.4.1 Diffusion calculation from SPAR metadata Assuming that SPAR perfectly reconstructs the covariance (COV) matrix, the output covariance matrix can be computed in the decoder from the input (DMX + decorrelator) covariance and upmix matrices. From the output COV, the reference power and intensity are computed and averaged over N frames (e.g., 8 frames). From there, the diffusivity is computed according to equation [7].

[0278] As disclosed below, there are other embodiments for directly calculating DirAC diffusivity from SPAR metadata without calculating the output covariance matrix.

[0279] 2.4.1.1 Alternative methods for spreading 1-channel downmixing using passive W downmixing (where the W channel in the downmix is ​​the same as the W channel in the input, or simply a delayed version thereof) Let w, x, y, and z be the variances of W, X, Y, and Z. In a one-channel downmix, y can be approximated as follows: y=w(pr y 2 +pd y 2 )

[0101] Here, pr y is the prediction coefficient, pd yis the decorrelation coefficient for the Y channel. Similarly, x and z can be calculated for the X and Z channels. Then the reference power E can be calculated as (w+x+y+z). E=w(1+pr y 2 +pd y 2 +pr x 2 +pd x 2 +pr z 2 +pd z 2 )

[0102]

[0280] The intensity can be calculated as follows: I y =w*pr y , I x =w*pr x , I z =w*pr z

[0103]

[0281] Referring to equation [7], the diffusive ψ can be directly approximated from the SPAR metadata as follows:

number

[0282] 2.4.1.2 Alternative Method for Diffusivity of 1-Channel Downmix Using Active W Downmix As described in Section 2.3.1, the inverse matrix is ​​given using active W computation.

number

[0283] Let w, x, y, and z be the variances of W, X, Y, and Z. In a one-channel downmix, y can be approximated as follows: y=w(pr y 2 +pd y 2 )

[0106] Here, pry is the prediction coefficient and pdy is the decorrelation coefficient for the Y channel. Similarly, x and z can also be calculated.

[0284] Then, the reference power can be calculated as (w+x+y+z). E=w((1-f s g' 2 ) 2 +pr y 2 +pd y 2 +pr x 2 +pd x 2 +pr z 2 +pd z 2 )

[0107]

[0285] The strength can be calculated as follows: I y =w*pr y , I x =w*pr x , I z =w*pr z

[0108]

[0286] Referring to equation [7], the diffusive ψ can be directly approximated from the SPAR metadata as follows (when w is averaged separately):

number

[0287] 2.4.1.3 Alternative methods for diffusion of any passive W downmix channel configuration This method is based on the normalization of the input ambisonic signal. For example, if the FoA input is normalized using Schmidt quasi-normalization (SN3D), we assume w = x + y + z, where w, x, y, and z are the variances of the W, X, Y, and Z channels, respectively. This gives w + x + y + z = 2 * w.

[0288] Substituting the assumed dispersion and intensity from equation

[0108] in Section 2.4.1.1 into the diffusivity formula in equation [7],

number

[0289] According to exemplary embodiments of the present disclosure, the processes described above may be implemented as a computer software program or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product which includes a computer program tangibly embodied on a machine-readable medium, the computer program which includes program code for performing the method. In such embodiments, the computer program may be downloaded and mounted from a network via a communication unit 709 and / or installed from a removable medium 711, as shown in Figure 7.

[0290] In general, various exemplary embodiments of this disclosure may be implemented in hardware or special-purpose circuits (e.g., control circuits), software, logic, or any combination thereof. For example, the units described above may be performed by a control circuit (e.g., CPU 701 in combination with the other components in Figure 7), and thus the control circuit may perform the actions described in this disclosure. Some aspects may be implemented in hardware, while others may be implemented in firmware or software that can be performed by a controller, microprocessor, or other computing device (e.g., a control circuit). Various aspects of the exemplary embodiments of this disclosure are illustrated and described using block diagrams, flowcharts, or any other pictorial representation, but it will be understood that the blocks, apparatus, systems, techniques, or methods described herein may, as non-limiting examples, be implemented in hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controllers, or other computing devices, or any combination thereof.

[0291] Furthermore, the various blocks shown in the flowchart can be viewed as method steps and / or actions resulting from the operation of computer program code and / or as a group of coupled logic circuit elements constructed to perform related functions. For example, embodiments of the present disclosure include a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code configured to perform the methods described above.

[0292] In the context of this disclosure, a machine-readable medium may be any tangible medium that contains or can store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable storage medium may be non-transient and may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. More specific examples of machine-readable storage media include electrical connections having one or more wires, portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0293] Computer program code for performing the methods of this disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to a processor of a general-purpose computer, a special-purpose computer, or another programmable data processing device having a control circuit, and when executed by the processor of such computer or other programmable data processing device, the program code will perform functions / operations specified in flowcharts and / or block diagrams. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.

[0294] This paper includes many specific implementation details, but these should not be interpreted as limitations on the scope of what can be claimed, but rather as descriptions of features that may be specific to a particular embodiment. Certain features described herein in the context of a separate embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable subcombination. Furthermore, features may be described above as acting in a certain combination, and may even be initially claimed as such, but one or more features from the claimed combination may, in some cases, be excluded from the combination, and the claimed combination may be directed towards a subcombination or a variation of a subcombination. The logical flow shown in the figures does not require a specific order or sequential order shown to achieve the desired result. Furthermore, other steps may be provided from the described flow, or steps may be removed, and other components may be added to or removed from the described system. Thus, other implementations are within the scope of the following claims.

Claims

1. A method for encoding audio signals: The step of acquiring an audio signal, wherein the audio signal represents the input audio scene with a main channel and side channels; The steps include: analyzing the power of the main channel of the audio signal; The step of analyzing the power of the side channel of the aforementioned audio signal; A step of detecting a mono mode for encoding the audio signal based on an analysis of the power of the main channel and the power of the side channel; For the detected mono mode, the steps include: calculating one or more downmix channels and spatial metadata from the audio signal; For the detected mono mode, the steps include encoding the one or more downmix channels and spatial metadata in a bitstream; The step of indicating mono mode in the bitstream Methods that include...

2. The method according to claim 1, further comprising outputting the bitstream, storing the bitstream, transmitting the bitstream, or a combination thereof.

3. The method according to claim 1 or 2, wherein the detection of mono mode is based on the determination that the input audio signal has non-silent audio in the main channel and that the side channel is silent.

4. The detection of the aforementioned mono-mode is: The main channel power of the main channel of the aforementioned audio signal is calculated; This includes calculating the total power as the sum of the powers of the side channels of the audio signal, The method according to claim 1 or 2.

5. The method according to claim 4, wherein the detection of the mono-mode further includes evaluating the ratio of the primary channel power to the total power.

6. The method according to claim 5, further comprising determining when the mono-mode detection exceeds a threshold.

7. The method according to any one of claims 1 to 6, further comprising implicitly transmitting the mono-mode by setting one or more parameters of the spatial metadata to a value of 0.

8. The method according to claim 6, wherein the one or more parameters of the spatial metadata include one or more spatial reconstruction (SPAR) metadata parameters, one or more directional audio coding (DirAC) metadata parameters, or both.

9. The method according to any one of claims 1 to 8, further comprising explicitly signaling to the mono mode by setting a mono flag.

10. The method according to any one of claims 1 to 9, wherein the bitstream is an immersive audio and sound (IVAS) encoded bitstream.

11. A method for decoding audio signals: The steps include obtaining the encoded bitstream; The steps include: decoding the encoded bitstream to obtain downmix channels, spatial metadata, and mono-mode indicators; When mono mode is detected, the process involves setting one or more parameters of the spatial metadata to 0; The steps include: upmixing the downmix channel using the spatial metadata; The stage of rendering the upmixed channels to the desired audio format and Methods that include...

12. The method according to claim 11, wherein the rendering generates rendered audio data, and the method further includes transmitting the rendered audio data, storing the rendered audio data, transmitting the rendered audio data, or a combination thereof.

13. The method according to claim 11, wherein the rendering generates rendered audio data, and the method further comprises providing the rendered audio data to one or more loudspeakers for playback.

14. The method according to any one of claims 11 to 13, wherein the mono-mode indicator is based on the value of one or more spatial metadata parameters in the encoded bitstream, which is set to a value indicating mono-mode.

15. The method according to claim 14, wherein the one or more spatial metadata parameters include one or more spatial reconstruction (SPAR) metadata parameters, one or more directional audio coding (DirAC) metadata parameters, or both.

16. The method according to any one of claims 11 to 15, comprising setting one or more parameters of spatial metadata to 0 when it is detected that the mono-mode is involved in setting one or more energy ratio values ​​to 0.

17. The method according to claim 16, wherein the received energy ratio metadata value is non-zero due to artifacts in the quantization process.

18. The method according to claim 17, further comprising setting one or more diffusivity values ​​to 1.

19. The method according to any one of claims 11 to 18, wherein the encoded bitstream is an immersive audio and sound (IVAS) encoded bitstream.

20. A method for encoding an audio signal representing an input audio scene with a main channel and side channels, the method being: The steps include: obtaining frames of the audio signal from the bitstream; The steps include: identifying a plurality of frequency bands associated with the frame; A step of determining the power associated with the main channel and the side channel for each of the plurality of frequency bands associated with the frame; The step of classifying each of the aforementioned frequency bands as one of silence, mono, or ambisonics based on the determined power for the corresponding frequency band; The steps include marking the mode of the frame as one of silent, mono, or ambisonics based on a plurality of classified frequency bands for the frame; A step of encoding the marked mode of the frame in the bitstream. The method according to any one of claims 11 to 18, including

21. The method according to claim 20, further comprising outputting the bitstream, storing the bitstream, transmitting the bitstream, or a combination thereof.

22. To further mark the aforementioned frame: The method includes marking a frame as ambisonics if one or more classified frequency bands correspond to ambisonics. The method according to claim 20 or 21.

23. To further mark the aforementioned frame: The method includes marking the frame as mono if none of the classified frequency bands correspond to ambisonics, and one or more classified frequency bands correspond to mono. The method according to claim 20 or 21.

24. To further mark the aforementioned frame: This includes marking the frame as silent if none of the classified frequency bands correspond to ambisonics and none of the classified frequency bands correspond to mono. The method according to any one of claims 20 to 23.

25. Determining the power of the frequency band further involves: Calculate the main channel power in that frequency band; This includes calculating the total power of that frequency band, The method according to any one of claims 20 to 24, wherein the total power is the sum of the powers to the side channels.

26. To further classify each of the aforementioned frequency bands: This includes evaluating the ratio of the main channel power to the total power, The method according to claim 25.

27. The method according to claim 26, wherein detecting a mono-mode includes determining when the ratio exceeds a threshold.

28. The method according to claim 27, wherein the threshold is a linear function of the main channel power, the main channel power may be normalized to a value between 0 and 1, and the threshold is clipped between a maximum and a minimum value.

29. Classifying each of the aforementioned frequency bands is as follows: It is determined that the sum power is less than the minimum noise level; When the aforementioned main channel power exceeds the noise threshold, the frequency band is declared as mono. The method according to claim 25, further comprising:

30. Classifying each of the aforementioned frequency bands is as follows: It is determined that the sum power exceeds the minimum noise level; Declaring the frequency band as mono when the normalized main channel power exceeds the noise threshold. The method according to claim 25, further comprising:

31. The method according to any one of claims 20 to 30, wherein the bitstream is an immersive audio and sound (IVAS) encoded bitstream.

32. The method according to any one of claims 20 to 31, wherein marking the mode of the frame involves setting one or more spatial reconstruction (SPAR) metadata parameters to 0, setting one or more directional audio coding (DirAC) metadata parameters to 0, or both.

33. An apparatus configured to perform the method described in any one of claims 1 to 32.

34. A system configured to perform the method described in any one of claims 1 to 32.

35. One or more non-transient computer-readable media on which instructions for controlling one or more devices to perform the method according to any one of claims 1 to 32 are encoded.

36. Input / Output (I / O) systems; and One or more processors A device having one or more processors: The step of acquiring an audio signal via the I / O system, wherein the audio signal represents an input audio scene with a main channel and side channels; The steps include: analyzing the power of the main channel of the audio signal; The step of analyzing the power of the side channel of the aforementioned audio signal; A step of detecting a mono mode for encoding the audio signal based on an analysis of the power of the main channel and the power of the side channel; For the detected mono mode, the steps include: calculating one or more downmix channels and spatial metadata from the audio signal; For the detected mono mode, the steps include encoding the one or more downmix channels and spatial metadata in a bitstream; The step of indicating mono mode in the bitstream A device configured to perform the following actions.

37. The apparatus according to claim 36, wherein one or more processors are further configured to output the bitstream, store the bitstream, transmit the bitstream, or a combination thereof.

38. The apparatus according to claim 36 or 37, wherein marking the frame further includes marking the frame as ambisonics if one or more classified frequency bands correspond to ambisonics.

39. The apparatus according to any one of claims 36 to 38, further comprising marking the frame as mono if none of the classified frequency bands correspond to ambisonics, and if one or more classified frequency bands correspond to mono.

40. Input / Output (I / O) systems; and One or more processors A device having one or more processors: The steps include: acquiring frames of audio signals from a bitstream via the aforementioned I / O system; The steps include: identifying a plurality of frequency bands associated with the frame; A step of determining the power associated with the main channel and the side channel for each of the plurality of frequency bands associated with the frame; The step of classifying each of the aforementioned frequency bands as one of silence, mono, or ambisonics based on the determined power for the corresponding frequency band; The steps include marking the mode of the frame as one of silent, mono, or ambisonics based on a plurality of classified frequency bands for the frame; A step of encoding the marked mode of the frame in the bitstream. A device configured to perform the following actions.

41. The apparatus according to claim 40, wherein one or more processors are further configured to output the bitstream, store the bitstream, transmit the bitstream, or a combination thereof.

42. The apparatus according to claim 40 or 41, wherein marking the frame further includes marking the frame as ambisonics if one or more classified frequency bands correspond to ambisonics.

43. The apparatus according to any one of claims 40 to 42, further comprising marking the frame as mono if none of the classified frequency bands correspond to ambisonics, and if one or more classified frequency bands correspond to mono.

44. Input / Output (I / O) systems; and One or more processors A device having one or more processors: The steps include: obtaining the encoded bitstream via the aforementioned I / O system; The steps include: decoding the encoded bitstream to obtain downmix channels, spatial metadata, and mono-mode indicators; When mono mode is detected, the process involves setting one or more parameters of the spatial metadata to 0; The steps include: upmixing the downmix channel using the spatial metadata; The stage of rendering the upmixed channels to the desired audio format and A device configured to perform the following actions.

45. The apparatus according to claim 44, wherein the rendering generates rendered audio data, and the one or more processors are further configured to transmit the rendered audio data, store the rendered audio data, transmit the rendered audio data, or a combination thereof.

46. The apparatus according to claim 44 or 45, wherein the rendering generates rendered audio data, and the one or more processors are further configured to provide the rendered audio data to one or more loudspeakers for playback.