MASA and ISM metadata diffusion and merging

The method optimizes the encoding of combined MASA and ISM audio formats by determining energy parameters and selecting ratio parameters, addressing quality degradation and bitrate inefficiencies in immersive audio codecs.

JP2026516944APending Publication Date: 2026-05-27NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2024-02-01
Publication Date
2026-05-27

AI Technical Summary

Technical Problem

Existing immersive audio codecs face challenges in efficiently encoding and merging different audio formats, such as MASA and ISM, leading to quality degradation and bitrate inefficiencies, particularly at low bitrates.

Method used

A method and apparatus for merging MASA and ISM metadata by determining energy parameters, generating weight values, and selecting appropriate ratio parameters based on comparisons to optimize the encoding process, ensuring efficient bitrate usage and improved audio quality.

Benefits of technology

The solution enhances audio quality and bitrate efficiency by optimizing the encoding of combined audio formats, reducing complexity and improving sound scene stability across varying bitrates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026516944000001_ABST
    Figure 2026516944000001_ABST
Patent Text Reader

Abstract

An apparatus comprising means for obtaining at least one first direct-to-overall ratio parameter for a first audio stream, obtaining a first signal energy parameter for a first audio stream, generating at least one first weight value based on at least one first direct-to-overall ratio parameter and the first signal energy parameter, obtaining a second direct-to-overall ratio parameter for a second audio stream, obtaining a second signal energy parameter for a second audio stream, generating a diffusion energy compensated direct-to-overall ratio parameter, generating a second weight value partially based on at least one of the second direct-to-overall ratio parameter, the diffusion energy compensated direct-to-overall ratio parameter, and the second signal energy parameter, and selecting at least one of the first direct-to-overall ratio parameter and the diffusion energy compensated direct-to-overall ratio parameter based on a comparison of at least one first weight value and the second weight value.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to an apparatus and method for merging MASA and ISM metadata, not limited to audio encoding, but aimed at preserving diffusion parameters. [Background technology]

[0002] Parametric spatial audio capture from inputs such as microphone arrays and other sound sources is a typical and effective choice for estimating a set of parameters from the input (microphone array signal), including the direction of sound in the frequency band and the ratio of directional to omnidirectional portions of the captured sound in the frequency band. These parameters are known to well represent the perceptual spatial properties of sound captured at the microphone array's position. These parameters can be used for spatial sound synthesis, whether binaural for headphones, for loudspeakers, or for other formats such as ambisonics.

[0003] Therefore, the directionality in the frequency band, as well as the direct-to-overall and diffuse-to-overall energy ratios, are particularly effective parameterizations for spatial audio capture.

[0004] A parameter set consisting of directional parameters in the frequency band and energy ratio parameters in the frequency band (indicating sound directivity) can also be used as spatial metadata for an audio codec (which may include other parameters such as surround coherence, spread coherence, number of directions, and distance). For example, these parameters can be estimated from audio signals captured by a microphone array, and stereo or mono signals can be generated from the microphone array signals and transmitted along with the spatial metadata.

[0005] Immersive audio codecs are designed to support a wide range of operating points, from low-bitrate operation to transparency. One example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, which is designed for use on communication networks such as 3GPP 4G / 5G networks, including use in immersive services such as immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general audio. Furthermore, it is expected to support channel-based audio, object-based audio, and scene-based audio input including spatial information about sound fields and sources. This codec is also expected to operate with low latency to enable conversational services, as well as support high error robustness under various transmission conditions.

[0006] Stereo signals can be encoded using, for example, the IVAS audio core codec, or the AAC (Advanced Audio Coding) or EVS (Enhanced Voice Services) encoder. A decoder can decode the audio stream signal into a PCM (Pulse Code Modulation) signal and process the sound in the frequency band (using spatial metadata) to obtain a spatial output, such as a binaural output.

[0007] The aforementioned immersive audio codecs are particularly well-suited for encoding spatial sound captured from microphone arrays (e.g., those found in mobile phones, VR cameras, and standalone microphone arrays). However, such encoders may have other input types, such as loudspeaker signals, audio object signals, and ambisonic signals. [Overview of the Initiative]

[0008] According to a first embodiment, an apparatus is provided comprising means for: obtaining at least one first direct-to-overall ratio parameter for a first audio stream; obtaining a first signal energy parameter for a first audio stream; generating at least one first weight value based on at least one first direct-to-overall ratio parameter and the first signal energy parameter; obtaining a second direct-to-overall ratio parameter for a second audio stream; obtaining a second signal energy parameter for a second audio stream; generating a diffusion energy compensated direct-to-overall ratio parameter; generating a second weight value partially based on at least one of the second direct-to-overall ratio parameter, the diffusion energy compensated direct-to-overall ratio parameter, and the second signal energy parameter; and selecting at least one of the first direct-to-overall ratio parameter and the diffusion energy compensated direct-to-overall ratio parameter based on a comparison of at least one first weight value and the second weight value.

[0009] Means for selecting one of at least one first direct-to-whole ratio parameter and diffusion energy compensation direct-to-whole ratio parameter based on a comparison of at least one first weight value and a second weight value may further include selecting one of at least one first direct-to-whole ratio parameters when at least one first weight value is greater than the second weight value, and selecting a diffusion energy compensation direct-to-whole ratio parameter when the second weight value is greater than at least one first weight value.

[0010] A means for selecting one of at least one first direct-to-whole ratio parameters when at least one first weight value is greater than a second weight value can be that at least one first weight value is strictly greater than a second weight value.

[0011] The means for selecting the diffusion energy compensation direct versus overall ratio parameter when the second weight value is greater than at least one first weight value can be the case where the second weight value is strictly greater than at least one first weight value.

[0012] Means for generating at least one first weight value based on at least one first direct-to-overall ratio parameter and a first signal energy parameter may also be for generating at least one first weight value based on the multiplication of at least one first direct-to-overall ratio parameter and a first signal energy parameter.

[0013] At least one first direct-to-overall ratio may include at least two first direct-to-overall ratios, and means for generating at least one first weight value based on at least one first direct-to-overall ratio parameter and a first signal energy parameter may be for generating first weight values ​​associated with each of the at least two first direct-to-overall ratio parameters.

[0014] Means for generating first weight values ​​associated with each of at least two first direct-to-overall ratio parameters may be for generating first weight values ​​based on the multiplication of each of at least one first direct-to-overall ratio parameter with a first signal energy parameter.

[0015] Means for selecting one of at least one first direct-to-overall ratio parameter and one of the diffusion energy compensation direct-to-overall ratio parameters based on a comparison of at least one first weight value and a second weight value may further include selecting at least two first direct-to-overall ratio parameters when both first weight values ​​are greater than the second weight values, and selecting a diffusion energy compensation direct-to-overall ratio parameter and one of the at least two first direct-to-overall ratio parameters when the second weight value is greater than at least one of the two first weight values.

[0016] Means for generating a second weight value, partially based on at least one of a second direct-to-overall ratio parameter, a diffusion energy compensated direct-to-overall ratio parameter, and a second signal energy parameter, may be for one of the following: generating at least one second weight value based on the multiplication of at least one second direct-to-overall ratio parameter and a second signal energy parameter; generating at least one second weight value based on the multiplication of a diffusion energy compensated direct-to-overall ratio parameter and a second signal energy parameter; and generating at least one second weight value based on the average of the diffusion energy compensated direct-to-overall ratio parameter and at least one second direct-to-overall ratio parameter multiplied by the second signal energy parameter.

[0017] Means for generating diffusion energy compensated direct versus whole ratio parameters may be for generating at least one merged diffusion versus whole ratio based at least in part on a first signal energy parameter, a second signal energy parameter, and a diffusion signal energy parameter.

[0018] The means may further include obtaining at least one first diffusion-to-overall ratio parameter for a first audio stream, and obtaining a diffusion signal energy parameter, in part, based on a first signal energy parameter and a first diffusion-to-overall ratio parameter.

[0019] The means for generating diffusion energy compensation direct versus overall ratio parameters may be for generating at least one diffusion energy compensation direct versus overall ratio parameter based on at least one merged diffusion versus overall ratio.

[0020] The means may further be for obtaining, for the second audio stream, at least one second overall diffusion pair ratio parameter, and the means for generating at least one merged overall diffusion pair ratio may further be for generating at least one merged overall diffusion pair ratio based on the second overall diffusion pair ratio parameter.

[0021] The means for generating the diffusion energy compensation direct pair overall ratio parameter may further be for: generating at least one expected diffusion energy compensation direct pair overall ratio parameter based on at least one merged overall diffusion pair ratio; comparing at least one expected diffusion energy compensation direct pair overall ratio parameter with a second direct pair overall ratio parameter; selecting the diffusion energy compensation direct pair overall ratio parameter from at least one expected diffusion energy compensation direct pair overall ratio parameter and the second direct pair overall ratio parameter based on the comparison such that the diffusion energy compensation direct pair overall ratio parameter is the smaller of at least one expected diffusion energy compensation direct pair overall ratio parameter and the second direct pair overall ratio parameter.

[0022] The means may further be for encoding one selected from at least one first direct pair overall ratio parameter and the diffusion energy compensation direct pair overall ratio parameter.

[0023] The means may further be for merging the first audio stream and the second audio stream.

[0024] According to a second aspect, a method is provided that includes: obtaining at least one first direct-to-total ratio parameter for a first audio stream; obtaining a first signal energy parameter for the first audio stream; generating at least one first weight value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; obtaining a second direct-to-total ratio parameter for a second audio stream; obtaining a second signal energy parameter for the second audio stream; generating a diffusion energy compensation direct-to-total ratio parameter; generating a second weight value based at least in part on at least one of the second direct-to-total ratio parameter, the diffusion energy compensation direct-to-total ratio parameter, and the second signal energy parameter; and selecting one of the at least one first direct-to-total ratio parameter and the diffusion energy compensation direct-to-total ratio parameter based on a comparison of the at least one first weight value and the second weight value.

[0025] Selecting one of the at least one first direct-to-total ratio parameter and the diffusion energy compensation direct-to-total ratio parameter based on a comparison of the at least one first weight value and the second weight value may further include selecting one of the at least one first direct-to-total ratio parameter when the at least one first weight value is greater than the second weight value, and selecting the diffusion energy compensation direct-to-total ratio parameter when the second weight value is greater than the at least one first weight value.

[0026] Selecting one of the at least one first direct-to-total ratio parameter when the at least one first weight value is greater than the second weight value may be when the at least one first weight value is strictly greater than the second weight value.

[0027] The selection of the diffusion energy compensation direct versus total ratio parameter can be considered as the case where the second weight value is strictly greater than at least one of the first weight values, provided that the second weight value is greater than at least one of the first weight values.

[0028] Generating at least one first weight value based on at least one first direct-to-overall ratio parameter and a first signal energy parameter may include generating at least one first weight value based on the multiplication of at least one first direct-to-overall ratio parameter and a first signal energy parameter.

[0029] At least one first direct-to-overall ratio may include at least two first direct-to-overall ratios, and generating at least one first weight value based on at least one first direct-to-overall ratio parameter and a first signal energy parameter may include generating a first weight value associated with each of the at least two first direct-to-overall ratio parameters.

[0030] Generating a first weight value associated with each of at least two first direct-to-overall ratio parameters may include generating a first weight value based on the multiplication of each of at least one first direct-to-overall ratio parameter with a first signal energy parameter.

[0031] Selecting at least one first direct-to-whole ratio parameter and one of the diffusion energy compensation direct-to-whole ratio parameters based on a comparison of at least one first weight value and a second weight value may include selecting at least two first direct-to-whole ratio parameters when both first weight values ​​are greater than the second weight values, and selecting a diffusion energy compensation direct-to-whole ratio parameter and one of the at least two first direct-to-whole ratio parameters when the second weight value is greater than at least one of the two first weight values.

[0032] Generating a second weight value based in part on at least one of a second direct-to-overall ratio parameter, a diffusion energy compensated direct-to-overall ratio parameter, and a second signal energy parameter may include one of the following: generating at least one second weight value based on the multiplication of at least one second direct-to-overall ratio parameter and a second signal energy parameter; generating at least one second weight value based on the multiplication of a diffusion energy compensated direct-to-overall ratio parameter and a second signal energy parameter; and generating at least one second weight value based on the average of a diffusion energy compensated direct-to-overall ratio parameter and at least one second direct-to-overall ratio parameter multiplied by the second signal energy parameter.

[0033] Generating a diffusion energy compensated direct versus whole ratio parameter may include generating at least one merged diffusion versus whole ratio based at least partially on a first signal energy parameter, a second signal energy parameter, and a diffusion signal energy parameter.

[0034] The method may further include obtaining at least one first spread-to-overall ratio parameter for a first audio stream, and obtaining a spread signal energy parameter based in part on a first signal energy parameter and the first spread-to-overall ratio parameter.

[0035] Generating diffusion energy compensation direct versus overall ratio parameters may include generating at least one diffusion energy compensation direct versus overall ratio parameter based on at least one merged diffusion versus overall ratio.

[0036] The method may further include obtaining at least one second diffusion-to-overall ratio parameter for a second audio stream, and generating at least one merged diffusion-to-overall ratio may further include generating at least one merged diffusion-to-overall ratio based on the second diffusion-to-overall ratio parameter.

[0037] Generating a diffusion energy compensation direct-to-overall ratio parameter may further include: generating at least one expected diffusion energy compensation direct-to-overall ratio parameter based on at least one merged diffusion-to-overall ratio; comparing at least one expected diffusion energy compensation direct-to-overall ratio parameter with a second direct-to-overall ratio parameter; and selecting a diffusion energy compensation direct-to-overall ratio parameter from at least one expected diffusion energy compensation direct-to-overall ratio parameter and the second direct-to-overall ratio parameter based on the comparison, such that the diffusion energy compensation direct-to-overall ratio parameter is the smaller of the at least one expected diffusion energy compensation direct-to-overall ratio parameter and the second direct-to-overall ratio parameter.

[0038] The method may further include encoding a selected one of at least one first direct-to-whole ratio parameter and a diffusion energy compensated direct-to-whole ratio parameter.

[0039] The method may further include merging the first audio stream with the second audio stream.

[0040] According to a third aspect, a device is provided comprising at least one processor and at least one memory storing an instruction, wherein when the instruction is executed by at least one processor, the system is caused to perform at least the following: obtain at least one first direct-to-whole ratio parameter for a first audio stream; obtain a first signal energy parameter for a first audio stream; generate at least one first weight value based on at least one first direct-to-whole ratio parameter and the first signal energy parameter; obtain a second direct-to-whole ratio parameter for a second audio stream; obtain a second signal energy parameter for a second audio stream; generate a diffusion energy compensated direct-to-whole ratio parameter; generate a second weight value based in part on at least one of the second direct-to-whole ratio parameter, the diffusion energy compensated direct-to-whole ratio parameter, and the second signal energy parameter; and select at least one of the first direct-to-whole ratio parameter and the diffusion energy compensated direct-to-whole ratio parameter based on a comparison of at least one first weight value and the second weight value.

[0041] An apparatus that performs the selection of at least one first direct-to-whole ratio parameter and a diffusion energy compensation direct-to-whole ratio parameter based on a comparison of at least one first weight value and a second weight value may further perform the selection of at least one first direct-to-whole ratio parameter if at least one first weight value is greater than the second weight value, and the selection of a diffusion energy compensation direct-to-whole ratio parameter if the second weight value is greater than at least one first weight value.

[0042] An apparatus that causes the selection of at least one first direct-to-whole ratio parameter when at least one first weight value is greater than a second weight value can be one where at least one first weight value is strictly greater than a second weight value.

[0043] An apparatus that causes a device to select a diffusion energy compensation direct versus total ratio parameter when a second weight value is greater than at least one first weight value can be one where the second weight value is strictly greater than at least one first weight value.

[0044] An apparatus that performs the task of generating at least one first weight value based on at least one first direct-to-whole ratio parameter and a first signal energy parameter may be further configured to perform the task of generating at least one first weight value based on the multiplication of at least one first direct-to-whole ratio parameter and a first signal energy parameter.

[0045] At least one first direct-to-overall ratio may include at least two first direct-to-overall ratios, and an apparatus that causes to generate at least one first weight value based on at least one first direct-to-overall ratio parameter and a first signal energy parameter may further cause to generate first weight values ​​associated with each of the at least two first direct-to-overall ratio parameters.

[0046] An apparatus that performs the task of generating first weight values ​​associated with each of at least two first direct-to-whole ratio parameters may be further configured to generate first weight values ​​based on the multiplication of each of at least one first direct-to-whole ratio parameter with a first signal energy parameter.

[0047] An apparatus that performs the selection of at least one first direct-to-whole ratio parameter and one of the diffusion energy compensation direct-to-whole ratio parameters based on a comparison of at least one first weight value and a second weight value may further perform the selection of at least two first direct-to-whole ratio parameters when both first weight values ​​are greater than the second weight values, and the selection of a diffusion energy compensation direct-to-whole ratio parameter and one of the at least two first direct-to-whole ratio parameters when the second weight value is greater than at least one of the two first weight values.

[0048] An apparatus that causes a second weight value to be generated based in part on at least one of a second direct-to-whole ratio parameter, a diffusion energy compensated direct-to-whole ratio parameter, and a second signal energy parameter may further perform one of the following: generating at least one second weight value based on the multiplication of at least one second direct-to-whole ratio parameter and a second signal energy parameter; generating at least one second weight value based on the multiplication of a diffusion energy compensated direct-to-whole ratio parameter and a second signal energy parameter; and generating at least one second weight value based on the average of the diffusion energy compensated direct-to-whole ratio parameter and at least one second direct-to-whole ratio parameter multiplied by the second signal energy parameter.

[0049] An apparatus that performs the generation of a diffusion energy compensated direct versus whole ratio parameter may be further configured to perform the generation of at least one merged diffusion versus whole ratio based at least in part on a first signal energy parameter, a second signal energy parameter, and a diffusion signal energy parameter.

[0050] The device may be configured to further perform, with respect to a first audio stream, at least one first diffusion-to-overall ratio parameter, and to obtain a diffusion signal energy parameter, partially based on the first signal energy parameter and the first diffusion-to-overall ratio parameter.

[0051] An apparatus that performs the task of generating diffusion energy compensation direct versus overall ratio parameters may be further configured to generate at least one diffusion energy compensation direct versus overall ratio parameter based on at least one merged diffusion versus overall ratio.

[0052] The device may be configured to further perform obtaining at least one second diffusion-to-overall ratio parameter for a second audio stream, and the device which performs generating at least one merged diffusion-to-overall ratio may further perform generating at least one merged diffusion-to-overall ratio based on the second diffusion-to-overall ratio parameter.

[0053] An apparatus for generating a diffusion energy compensation direct-to-overall ratio parameter may further perform: generating at least one expected diffusion energy compensation direct-to-overall ratio parameter based on at least one merged diffusion-to-overall ratio; comparing at least one expected diffusion energy compensation direct-to-overall ratio parameter with a second direct-to-overall ratio parameter; and selecting a diffusion energy compensation direct-to-overall ratio parameter from at least one expected diffusion energy compensation direct-to-overall ratio parameter and the second direct-to-overall ratio parameter based on the comparison, such that the diffusion energy compensation direct-to-overall ratio parameter is the smaller of the at least one expected diffusion energy compensation direct-to-overall ratio parameter and the second direct-to-overall ratio parameter.

[0054] The apparatus may be configured to further perform encoding of at least one first direct-to-whole ratio parameter and a selected one of the diffusion energy compensated direct-to-whole ratio parameters.

[0055] The device may be configured to further perform the merging of the first audio stream and the second audio stream.

[0056] According to a fourth aspect, an apparatus is provided comprising: an acquisition circuit configured to acquire at least one first direct-to-overall ratio parameter for a first audio stream; an acquisition circuit configured to acquire a first signal energy parameter for a first audio stream; a generation circuit configured to generate at least one first weight value based on at least one first direct-to-overall ratio parameter and the first signal energy parameter; an acquisition circuit configured to acquire a second direct-to-overall ratio parameter for a second audio stream; an acquisition circuit configured to acquire a second signal energy parameter for a second audio stream; a generation circuit configured to generate a diffusion energy compensated direct-to-overall ratio parameter; a generation circuit configured to generate a second weight value based in part on at least one of a second direct-to-overall ratio parameter, a diffusion energy compensated direct-to-overall ratio parameter, and a second signal energy parameter; and a selection circuit configured to select at least one of the first direct-to-overall ratio parameter and the diffusion energy compensated direct-to-overall ratio parameter based on a comparison of at least one first weight value and a second weight value.

[0057] According to the fifth aspect, a computer program (or a computer-readable medium containing program instructions) is provided which includes instructions causing the device to perform at least the following: for a first audio stream, obtain at least one first direct-to-overall ratio parameter; for a first audio stream, obtain a first signal energy parameter; generate at least one first weight value based on at least one first direct-to-overall ratio parameter and the first signal energy parameter; for a second audio stream, obtain a second direct-to-overall ratio parameter; for a second audio stream, obtain a second signal energy parameter; generate a diffusion energy compensated direct-to-overall ratio parameter; generate a second weight value based in part on at least one of the second direct-to-overall ratio parameter, the diffusion energy compensated direct-to-overall ratio parameter, and the second signal energy parameter; and select at least one of the first direct-to-overall ratio parameter and the diffusion energy compensated direct-to-overall ratio parameter based on a comparison of at least one first weight value and the second weight value.

[0058] According to the sixth aspect, a non-temporary computer-readable medium is provided which includes program instructions causing the device to perform at least the following: for a first audio stream, obtain at least one first direct-to-overall ratio parameter; for a first audio stream, obtain a first signal energy parameter; generate at least one first weight value based on at least one first direct-to-overall ratio parameter and the first signal energy parameter; for a second audio stream, obtain a second direct-to-overall ratio parameter; for a second audio stream, obtain a second signal energy parameter; generate a diffusion energy compensated direct-to-overall ratio parameter; generate a second weight value based in part on at least one of the second direct-to-overall ratio parameter, the diffusion energy compensated direct-to-overall ratio parameter, and the second signal energy parameter; and select at least one of the first direct-to-overall ratio parameter and the diffusion energy compensated direct-to-overall ratio parameter based on a comparison of at least one first weight value and the second weight value.

[0059] According to the seventh aspect, an apparatus is provided comprising: means for obtaining at least one first direct-to-overall ratio parameter for a first audio stream; means for obtaining a first signal energy parameter for a first audio stream; means for generating at least one first weight value based on at least one first direct-to-overall ratio parameter and the first signal energy parameter; means for obtaining a second direct-to-overall ratio parameter for a second audio stream; means for obtaining a second signal energy parameter for a second audio stream; means for generating a diffusion energy compensated direct-to-overall ratio parameter; means for generating a second weight value based in part on at least one of the second direct-to-overall ratio parameter, the diffusion energy compensated direct-to-overall ratio parameter, and the second signal energy parameter; and means for selecting one of the at least one first direct-to-overall ratio parameter and the diffusion energy compensated direct-to-overall ratio parameter based on a comparison of at least one first weight value and the second weight value.

[0060] According to the eighth aspect, a computer-readable medium is provided which includes program instructions causing the apparatus to perform at least the following: for a first audio stream, obtain at least one first direct-to-whole ratio parameter; for a first audio stream, obtain a first signal energy parameter; generate at least one first weight value based on at least one first direct-to-whole ratio parameter and the first signal energy parameter; for a second audio stream, obtain a second direct-to-whole ratio parameter; for a second audio stream, obtain a second signal energy parameter; generate a diffusion energy compensated direct-to-whole ratio parameter; generate a second weight value based in part on at least one of the second direct-to-whole ratio parameter, the diffusion energy compensated direct-to-whole ratio parameter, and the second signal energy parameter; and select at least one of the first direct-to-whole ratio parameter and the diffusion energy compensated direct-to-whole ratio parameter based on a comparison of at least one first weight value and the second weight value.

[0061] An apparatus comprising means for performing the actions of the above method.

[0062] A device configured to perform the actions of the method described above.

[0063] A computer program that includes program instructions for causing a computer to perform the above method.

[0064] A computer program product stored on a medium can cause the device to perform the methods described herein.

[0065] The electronic device may include the apparatus described herein.

[0066] The chipset may include the devices described herein.

[0067] The embodiments of this application aim to address problems related to the prior art.

[0068] Next, to better understand this application, please refer to the attached drawings as examples. [Brief explanation of the drawing]

[0069] [Figure 1] This is a schematic diagram of the device used for extracting MASA metadata. [Figure 2] This is a schematic diagram of an example MASA metadata stream merger. [Figure 3] Figure 2 is a schematic diagram showing a more detailed example of the MASA stream merger. [Figure 4] This is a schematic diagram of an example OMASA stream merger for merging MASA streams and ISM streams into a combined stream. [Figure 5] This is a schematic diagram of an exemplary system of an apparatus suitable for carrying out several embodiments. [Figure 6] This is a schematic diagram of an exemplary merger suitable for use as part of the encoder shown in Figure 5, configured to perform merging according to several embodiments. [Figure 7] Figure 6 is a flowchart illustrating the operation of an exemplary metadata merger in several embodiments. [Figure 8] This is a schematic diagram of an additional exemplary metadata merger, suitable for use as part of the encoder shown in Figure 5, configured to perform merging according to several embodiments. [Figure 9] This is a flowchart illustrating the operation of a metadata merger, an additional example of the example shown in Figure 8, based on several embodiments. [Figure 10] This is a schematic diagram of another exemplary metadata merger suitable for use as part of the encoder shown in Figure 5, configured to perform merging according to several embodiments. [Figure 11]This is a schematic diagram of a multidirectional metadata merger, suitable for use as part of the encoder shown in Figure 5, configured to perform merging in several embodiments. [Figure 12] This is a diagram of an example device suitable for implementing the apparatus shown in the previous drawing. [Modes for carrying out the invention]

[0070] The following section further details suitable devices and possible mechanisms for encoding parametric spatial audio signals, including transport audio signals and spatial metadata. As mentioned earlier, immersive audio codecs (such as 3GPP IVAS) are planned to support a wide range of operating points, from low-bitrate operation to transparency. They are expected to support channel-based audio, object-based audio, and scene-based audio inputs, including spatial information about the sound field and sound source.

[0071] In the following examples, the codecs are configured to receive multiple input formats. Specifically, the codecs are configured to acquire or receive multi-audio signals (e.g., those received from a microphone array, or as a multi-channel audio format input, or as an ambisonics format input) and audio object signals (these may also be called Independent Stream with Metadata - ISM format). Furthermore, depending on the circumstances, the codecs may be configured to handle multiple input formats simultaneously.

[0072] This combined (input) format mode allows for simultaneous encoding of, for example, two different audio input formats. An example of two different audio input formats currently under consideration is a combination of the MASA format and the Audio Object Format (ISM format).

[0073] Metadata-assisted spatial audio (MASA) is an example of a parametric spatial audio format and is a suitable representation as an input format for IVAS.

[0074] This can be thought of as an audio representation consisting of "N channels + spatial metadata." This is a scene-based audio format, particularly well-suited for capturing spatial audio on practical devices such as smartphones. The idea is to describe the sound scene in terms of sound source direction, which changes with time and frequency, and, for example, energy ratios. Sound energy that is not defined (not described) by direction is described as diffuse (coming from all directions).

[0075] As mentioned above, spatial metadata associated with an audio signal may include multiple parameters (multiple directions, and associated with each direction (or direction value), such as direct-to-whole energy ratio, spread coherence, distance, etc.) for each time-frequency tile. Spatial metadata may also include other parameters, or may be associated with other parameters that are considered omnidirectional (such as surround coherence, spread-to-whole energy ratio, residual-to-whole energy ratio, etc.), but when combined with directional parameters, it can be used to define the characteristics of an audio scene. For example, a reasonable design choice that can produce a good output is one in which the spatial metadata includes one or more directions (and associated with each direction, such as direct-to-whole ratio, spread coherence, distance value, etc.) for each time-frequency subframe.

[0076] Figure 1 shows an illustrative MASA analyzer 101. The MASA analyzer 101 is configured to receive an input audio signal 100, analyze the input audio signal, and generate a transport audio signal 102 and spatial metadata 104.

[0077] An example of MASA spatial metadata is shown in the table below. These values ​​are available for each time-frequency tile. In some embodiments, a frame is subdivided into 24 frequency bands and 4 time subframes. In other embodiments, other divisions of frequency and time may be employed. Furthermore, in some embodiments (as implemented in IVAS, for example), the frame size is 20 ms (and therefore the time subframe is 5 ms). However, similarly, other frames may be employed in other embodiments. In some embodiments, the MASA analyzer is configured to determine one or two directions for each time-frequency tile (i.e., one or two direction indices, a direct-to-overall energy ratio, and a spread coherence parameter for each time-frequency tile). However, in some embodiments, the analyzer is configured to generate more than two directions for a given time-frequency tile.

[0078] [Table 1]

[0079] The MASA stream can be rendered to various outputs, including multi-channel loudspeaker signals (e.g., 5.1) and binaural signals.

[0080] Another input format supported by IVAS, as mentioned earlier, is the ISM (Independent Streams with Metadata) format. The ISM format is intended to represent individual sound sources (audio objects) within a sound scene, along with associated metadata that describes how the audio signal can be rendered. This may be in contrast to the MASA format, where the audio signal and associated metadata perceptually describe the entire sound scene. Example metadata for the ISM format may include the following parameters: Location (for example, azimuth and elevation, or azimuth, elevation, and radius), direction, Spread, Distance attenuation, and Directional pattern.

[0081] In some embodiments, some exemplary metadata parameters, such as position, orientation, and spread, may vary over time but not over frequency, while some parameters, such as distance attenuation and directivity patterns, may vary over frequency but are typically not constant over time. However, parameters may also exhibit temporal and frequency variations (or invariance) other than those exemplified above and herein.

[0082] In some embodiments, ISM format inputs can be obtained by capturing individual sound sources in a scene, for example, using close microphones or lavalier microphones (placed on or near each individual sound source). Examples of such individual sound sources include separate speakers, singers, or individual instruments in a conference call. Metadata can then be associated with these ISM format signals, for example, automatically by a position tracker or manually by a mixing professional. Alternatively, ISM format signals can be generated by mixing to create the appropriate associated metadata. An example of a fully generated sound source is a game audio engine.

[0083] As described above, research is underway to enable the IVAS codec to support combined coding of multiple audio formats. In many use cases, content captured from a sound scene may exist in multiple formats, and encoding them as a combined format allows for more efficient encoding and transmission. An example use case illustrating this situation is the "reporter scene," in which spatial capture is performed to obtain the overall sound scene, and a separate microphone is added to the reporter to optimally capture its voice, thus improving the quality of the complete capture. In this example, the corresponding capture formats could be MASA format (from the captured sound scene) and ISM format (from the reporter). Without combined coding, these formats would require separate instances of the IVAS codec, but with combined coding, only one instance is needed. This reduces complexity and allows for highly flexible optimization of bitrate usage between different parts of the combined format (audio and metadata).

[0084] This combined format can be referred to as the "Objects and MASA" format (OMASA), which combines the MASA format and the ISM format into a single combined format.

[0085] GB2217905.5 describes bitrate-dependent operating modes for OMASA coding. The following example features the lowest bitrate operating mode in which an ISM format input is merged with a MASA format input and they are encoded as a MASA format stream. However, it should be understood that embodiments may use other modes or be employed in other OMASA coding situations.

[0086] At the lowest bitrate of OMASA format coding, a format merging method is required to merge formats into a more suitable format for encoding and transmission through tight bitrate limitations. An example of a merging method is described in GB2574238. This merging method shows two parametric spatial audio format inputs, such as the MASA format being merged into a single-output parametric spatial audio format. This is shown, for example, with respect to Figures 2 and 3.

[0087] Figure 2 shows a (metadata)stream merger 201 configured to receive MASA stream 1 200 and MASA stream 2 202 and generate a (MASA)combined stream 204. Figure 3 further illustrates the exemplary (metadata)stream merger 201 in more detail with respect to metadata merging. In this example, the (metadata)stream merger 201 is configured to receive MASA stream 1 200 and MASA stream 2 202.

[0088] The illustrative (metadata) stream merger 201 includes an energy determinator (reference numeral 311 for stream 1 and reference numeral 313 for stream 2) configured to determine the energy 300 of stream 1 and the energy 302 of stream 2. The energy determinator can, for example, separately calculate the energy of the MASA transport signal at each parameter TF tile (b,s) for both MASA format streams:

[0089]

number

[0090] Here,

number

[0091] The illustrative (metadata) stream merger 201 further includes ratio determinators (reference code 301 for stream 1 and reference code 303 for stream 2). The ratio determinators can retrieve ratios from the MASA streams and pass them to the weight determinators.

[0092] Furthermore, the illustrative (metadata) stream merger 201 may include a weight decisioner (reference code 305 for stream 1 and reference code 307 for stream 2). The weight decisioner is

number

number

[0093] The illustrative (metadata) stream merger 201 further includes a weight comparator 309 configured to compare weights for each TF tile and control the metadata selector 321.

[0094] Furthermore, the exemplary (metadata) stream merger 201 includes a metadata selector 321 configured to receive metadata from the MASA stream and control from the weight comparator 309. The metadata selector 321 can be configured as follows: For TF tiles, if w1 > w2, metadata for that TF tile is selected from MASA stream 1. Otherwise, select metadata for that TF tile from MASA Stream 2.

[0095] Next, the selected metadata can be output as merged metadata in the (MASA) combined stream 204. In some situations, a > "greater than" comparison is equivalent to an ≥ "greater than or equal to" comparison. This choice applies to all the examples below. In some situations, the comparison may include a bias factor (which can be additive and / or multiplicative) configured to skew the determination in a certain direction.

[0096] This can be further extended to input formats other than the MASA input format. An example is shown in Figure 4. In Figure 4, the ISM stream and the MASA stream are merged and encoded to generate a bitstream.

[0097] Therefore, for example, Figure 4 shows a merger / encoder that includes an audio stream merger 401 configured to generate an audio stream 412 combined from the Stream 1 (MASA) input audio signal 400 and the Stream 2 (ISM) input audio signal 402.

[0098] The merger / encoder further includes an audio stream energy determinator 403, which is configured to determine the audio signal energy 408 in a similar manner to that described above with respect to the energy determinators 311 and 313, for example, as shown in Figure 3.

[0099] The merger / encoder may include an ISM-to-MASA converter 405 configured to receive an (ISM) stream 2 406 and generate a MASA stream 2 202. Since the metadata of the ISM stream contains values ​​compatible with the MASA stream, the ISM-to-MASA converter 405 is configured to generate or construct a MASA stream from the ISM stream. For a single (active) object, the MASA stream can be generated, for example, by converting the position of the ISM stream to the direction of the MASA stream, making the constructed MASA stream fully directional. In other words, assign a direct-to-whole ratio value of 1 and a spread-to-whole ratio value of 0 to the constructed MASA format. Defining the ISM as a fully directional source and assigning it a direct-to-whole ratio of 1 and a spread-to-whole ratio of 0 can be justified by the intended use where the ISM is acquired as a single audio object, such as a lavalier microphone.

[0100] In some cases, especially when multiple objects are present simultaneously in the ISM stream, converter 405 is configured to convert the ISM to a primary ambisonics (FOA) representation and analyze the source direction from the representation. The direct-to-total energy ratio can be obtained by estimating the diffuse-to-total energy ratio of this FOA signal and determining the direct-to-total energy ratio from it. Any other suitable method can be used to convert the ISM stream to a MASA stream (two specific examples are given above).

[0101] Next, the metadata of the two MASA format streams can be merged.

[0102] For example, the merger / encoder may include a metadata stream merger 407, which receives (MASA) stream 1 200 and MASA stream 2 202 and is configured to merge metadata from these inputs in a manner similar to that described by the example shown with respect to Figure 3. The combined MASA stream 410 generated by the metadata stream merger 407 can then be output to a spatial metadata encoder 409.

[0103] Next, the merger / encoder may include a spatial metadata encoder 409, which is configured to encode the combined MASA stream 410 based on a suitable MASA metadata encoding method and output the encoded metadata to a bitstream creator 413.

[0104] The merger / encoder may further include a bitstream creator 413, which is configured to receive encoded audio 414 from the audio encoder 411 and encoded metadata 416 from the spatial metadata encoder 409, and to generate a bitstream 418 that is output for storage and / or transmission.

[0105] With respect to Figure 5, an exemplary system of an apparatus suitable for carrying out several embodiments is shown. The apparatus system is configured to receive a stream 1 input audio signal 500, stream 1 spatial metadata (format 1) 504, and a stream 2 input audio signal 502, stream 2 spatial metadata (format 2) 506. These are received by a merger / encoder 501 configured to merge the inputs to produce a bitstream 510. The bitstream is then passed to a decoder 503, which then decodes the bitstream to produce an output audio signal 512.

[0106] This concept aims to improve upon the above example of merging MASA streams and streams converted from ISM to MASA. The above method can result in quality degradation, particularly in terms of time-frequency resolution, and when the assigned bitrate of the encoded and transmitted merged MASA format is limited to, for example, 24.4kbps. The quality degradation is most noticeable at the lowest bitrates with low time-frequency resolution, but losses can also occur at higher bitrates, therefore the goal is to improve quality at various bitrates.

[0107] The reason for the quality degradation is that the direct-to-whole ratio of the ISM stream is close to 1, and when combined with a considerable amount of signal energy, the presence of the ISM stream is emphasized in the resulting merged encoded MASA format. This can be perceived as a fluctuating instability in the reproduced sound scene, as ambient diffuse sound present in the MASA stream is almost completely eliminated when the ISM stream is activated. In addition, in some situations, TF tiles in the merged stream containing metadata from the ISM stream are allocated a higher bit budget than other TF tiles in metadata encoding, and coding artifacts (e.g., heavily quantized directions) are perceived due to the high direct-to-whole ratio of the ISM stream. This can also lead to a decrease in the perceived quality of the merged stream.

[0108] The concepts discussed below in relation to several embodiments and examples are apparatuses and methods that aim to merge a MASA stream with a stream converted from ISM to MASA, while preserving the overall diffusion properties of the MASA format sound source stream in the merged MASA format stream.

[0109] Accordingly, the embodiments and examples described herein relate to the merging of two parametric spatial audio (i.e., audio signals associated with spatial metadata) streams, where the first stream is a MASA format stream and the second stream is an ISM stream.

[0110] The apparatus and method propose to adjust the merging of parameters of these two spatial metadata streams so that the perceived overall quality of the merged streams is improved by preserving the sense of saturation of the MASA format scene through the preservation of its diffusion energy.

[0111] This can be carried out in the following manner in the following embodiments and examples: Get the direct versus overall ratio for both streams; Obtain the signal energy for both streams, and the diffusion energy for the MASA stream; Determine the direct versus overall ratio parameter for diffusion energy compensation: Determine the weights for both streams (where the weights can be based on the direct-to-whole ratio and signal energy estimates of the streams); Compare the weights for both streams in each TF tile; Based on the comparison, select the metadata to use for each TF tile in the merged MASA stream. The selection can be a “MASA stream” parameter for TF tiles in MASA streams where the MASA stream-related weights are greater than the ISM stream-related weights, a diffusion energy compensation ISM stream parameter where the ISM stream-related weights are greater than those of the MASA, or a selection of MASA stream and diffusion energy compensation ISM stream parameters for TF tiles based on the value of the ISM stream-related weight being greater than at least one of the multiple MASA stream-related weights.

[0112] In the following disclosure, the terms "weight" and "weight value" should be understood to be interchangeable.

[0113] Thus, the apparatus and method aim to provide improved merging. The examples presented herein involve two parametric spatial audio streams, where one stream is originally an ISM stream, i.e., one or more independent audio objects, and the other stream is a MASA format stream (or any similar parametric format). In some embodiments, the other stream format can be merged based on the examples presented herein, without significant inventive modifications. When the metadata of a data stream in a certain input format exhibits properties similar to an ISM stream, in other words, when it is highly directional and low diffuse, the stream in this input format is particularly well-suited for merging using the examples described herein.

[0114] In some embodiments, the apparatus and methods described herein can be implemented with a communication codec encoder such as 3GPP IVAS, and combined format coding can be applied at low bitrates where complete merging of streams of separate formats is required for good quality transmission.

[0115] As described herein, embodiments can further be implemented as part of a merger / encoder 501, as shown in Figure 5. Thus, in some embodiments, the device is configured to receive two streams of parametric spatial audio ("Stream 1" and "Stream 2") as inputs. Each input stream contains an audio signal + metadata ("Input audio signal for Stream 1", "Spatial metadata for Stream 1 (Format 1)", "Input audio signal for Stream 2", "Spatial metadata for Stream 2 (Format 2)").

[0116] In this example, the streams are a MASA format stream and an ISM stream. In some embodiments, the merger / encoder may include a preprocessor that converts the captured audio signal into a MASA format stream and an ISM stream. The input stream is given to an encoder that quantizes the format and encodes it into a single bitstream. This bitstream is sent to a decoder, which decodes the bitstream to produce an output audio signal, or in some embodiments, an output format (e.g., MASA format).

[0117] The illustrative merger shown in Figure 6 is suitable for performing metadata merging operations within the merger / encoder in Figure 5. In these embodiments, merging or combining audio signals and metadata is performed using different methods. Combining or merging audio signals can be performed using any suitable method. For example, input audio signals are mixed to prepare a merged audio signal, which is then encoded. Therefore, merging or combining audio signals will not be described in further detail.

[0118] Figure 6 shows (metadata) mergers or couplers in several embodiments.

[0119] The Merger is configured to receive metadata 600 for MASA stream 1, transport signals 602 for MASA stream 1, and metadata / object audio 610 for the ISM stream.

[0120] In some embodiments, the merger includes an ISM-MASA converter 607. The ISM-MASA converter 607 is configured to receive metadata / object audio 610 of the ISM stream and generate metadata 614 of the MASA stream 2 and a transport signal 612 of the ISM(MASA) stream. This can be done in any suitable way, for example, by assigning the values ​​of the direction parameter of the ISM stream to the direction parameter of the MASA format metadata for all frequency bands so that the direct-to-whole ratio and spread-to-whole ratio of the MASA format metadata are always 1 and 0 (or by converting the audio signal to an FOA representation as described above and estimating the parameters therefrom). The metadata 614 of the MASA stream 2 is passed to a direct-to-whole ratio acquirer 609. The transport signal 612 of the ISM(MASA) stream is passed to a transport signal energy determiner 611 and an audio stream merger 670.

[0121] The merger includes a direct-to-whole ratio acquirer 601, which acquires metadata 600 of MASA stream 1 and outputs the MASA ratio to the weighted decisioner 605.

number

[0122] The metadata merger includes a direct-to-whole ratio acquirer 609, which acquires metadata 614 from the MASA stream 2 and outputs ISM-based ratios to the weighted determinist 613.

number

[0123] In addition, the metadata merger includes a transport signal energy determinator 603 configured to estimate the energy of the MASA transport signal at each parameter TF tile using the following:

number

[0124] Here, X i (k,n) is the CLDFB representation of the i-th transport channel in CLDFB bin k and slot n, grouped into parameter band b and frame s. For simplicity, the index (b,s) is omitted below, but the calculation is performed for each TF tile. Output e MASA 606 is output to the weight determination unit 605, the diffusion signal energy acquirer 653, and the merged diffusion versus overall ratio determination unit 655.

[0125] The Merger includes a further transport signal energy determinator 611 and is configured to compute an estimate of the transport signal energy of the ISM transport signal 612 at each parameter TF tile:

number

[0126] The MASA further includes a spread-to-total ratio acquirer 651, which obtains the spread-to-total ratio from the metadata 600 of the MASA stream 1.

number

[0127] The Merger further includes a diffusion signal energy acquirer 653, which measures the diffusion to total ratio.

number

number

number

[0128] The merger includes a merged diffusion-to-whole ratio determinant 655 configured to calculate the merged diffusion-to-whole ratio of the merged MASA stream 1 and the ISM stream:

number

[0129] This value describes the amount of spreading energy of the signal after merging the two streams. Spreading signal energy

number

[0130] The data merger also includes a merged direct vs. whole ratio determinant 657. This is,

number

[0131] The metadata merger includes a weight determinator 605, and this weight determinator 605 is

number

[0132] The metadata merger includes a further weighting determinator 613, which is configured to use the direct versus total energy ratio and the transport signal energy derived from the ISM stream to determine comparison weights for the ISM stream. These weights are similarly,

number

[0133] Therefore, two weights w1 and w2 for each TF tile are passed to the weight comparator 615.

[0134] The metadata merger includes a weight comparator 615 configured to compare weights and generate selection control to the metadata selector 617.

[0135] The metadata merger further includes a metadata selector 617 configured to select based on selection control from a weighted comparator and output the selected metadata 622.

[0136] For example, a selector can be configured to perform the following selection operation: For TF tiles, if w1 > w2, select metadata for that TF tile from MASA stream 1 600; Otherwise, select metadata from the MASA stream 2 614 for that TF tile (in other words, metadata from the original "ISM stream") and create a new direct vs. whole ratio.

number

number

number

[0137] In addition, in some embodiments, the merger includes an audio stream merger 670 configured to receive the transport / audio signal 664 of MASA stream 2 and the transport signal 602 of MASA stream 1 from the ISM-MASA converter 607, and from these generate a transport / audio signal 674 of a combined stream that can be output to a suitable audio encoder.

[0138] The merged metadata 622 can then be passed to a suitable spatial metadata encoder, which can provide an encoded bitstream that can be multiplexed with the audio signal within a suitable bitstream creator, for example, and sent to a decoder (or stored for later decoding). The operation of the decoder will not be discussed in more detail, as the decoding and decoding of the bitstream itself can be performed using any suitable decoder. In other words, no changes are required to the decoder to support the merged MASA stream. Thus, in summary, the decoder decodes the bitstream by demultiplexing it into audio channels and metadata, which are then output as MASA format or passed to a renderer that renders them into various output audio signals. These signals can be binaural, loudspeaker, ambisonic, or other formats.

[0139] Regarding Figure 7, a flowchart illustrating an example of the operation shown in Figure 6 is provided.

[0140] Therefore, as indicated by 700, it is shown that metadata for MASA stream 1 is retrieved.

[0141] Furthermore, it is shown that the transport signal for MASA stream 1 is obtained by 702.

[0142] In addition, it has been shown that metadata / object audio of the ISM stream can be obtained using 721.

[0143] 703 demonstrates obtaining the direct-to-whole ratio from the metadata of MASA Stream 1.

[0144] 705 demonstrates that the transport signal energy is determined from the transport signal of MASA stream 1.

[0145] 709 demonstrates obtaining the diffusion-to-total ratio from the metadata of MASA stream 1.

[0146] In 711, we demonstrate that the diffusion signal energy can be obtained from the diffusion-to-total ratio and the transport signal energy.

[0147] In 707, we show that the weight of stream 1 is determined from the direct-to-whole ratio and the transport signal energy.

[0148] 723 indicates the conversion of metadata from ISM to MASA and the retrieval of audio from the ISM stream (or MASA stream 2).

[0149] The value 725 indicates that the direct-to-overall ratio is obtained from the converted MASA metadata.

[0150] In 719, we demonstrate how to determine the transport signal energy from the transport signal of the ISM stream.

[0151] In 727, we show that the weights for stream 2 are determined from the direct-to-whole ratio and the transport signal energy (for the original ISM stream).

[0152] We demonstrate in 713 that we can determine the merged diffusion versus overall ratio.

[0153] Next, in 715, we show that the merged direct to overall ratio is determined from the merged direct to overall ratio.

[0154] The weight comparison is shown as 729.

[0155] Next, metadata selection based on comparison is shown in 731.

[0156] The output of the selected metadata is shown as 735.

[0157] In addition, as shown in 741, audio stream merging is demonstrated.

[0158] Next, the combined or merged audio streams are output, as shown in 743.

[0159] This embodiment aims to give greater consideration to the overall diffuse energy of the MASA format, which is important for conveying the perceptual sense of immersion in a sound scene. This aims to improve upon known merging methods, which often completely remove (or significantly reduce) the diffuse energy for a TF tile when metadata from an ISM format stream is selected for use in that tile. Thus, the embodiment adjusts the amount of diffuse energy to account for the increased overall energy from the ISM stream. Assuming that the ISM stream energy is predominantly direct, it reduces the diffuse energy for the tile, but only by a percentage of the relative stream energy. This aims to significantly improve the quality of the merged stream, as the perception of the diffuse field is better preserved.

[0160] While these exemplary embodiments aim to achieve significant quality improvements, some alternative embodiments may further refine the criteria and / or determined ratio values ​​of the merged stream metadata in an attempt to further improve quality. These further embodiments are discussed below. It should also be noted that these improvements can be used in combination. For example, in some embodiments, it may be possible to apply the improvements of the following two examples.

[0161] In some embodiments as shown in Figure 8, the comparison weights of stream 2w2 (MASA stream formed from ISM stream) are:

number

[0162] Therefore, in the exemplary embodiment shown in Figure 8, the weight determination unit 813 is passed to the weight comparator 615.

number

[0163] In some embodiments, the weight determination 813 uses two direct versus overall ratios r ISM ,

number

[0164]

number

[0165] Figure 9 is a flowchart illustrating the operation of the example apparatus in Figure 8, and differs from Figure 7 in that the operation for determining the weights of Stream 2 differs in the elements used to generate the weights as shown in 927.

[0166] The difference with the examples shown in Figures 6 and 8 is that (slightly) different weights are used for stream determination. However, if the determination is the same, the resulting selected metadata will also be the same. In practice, using these adjusted determination criteria preserves the stability of the sound scene, and further preserves the sense of spaciousness of the scene, especially since the MASA stream is selected more frequently when the ISM stream contains a ratio close to 1.

[0167] In some embodiments, the conversion from ISM to MASA representation can be performed by assuming that the signal energy is perfectly directional and assigning a direct-to-total energy ratio of 1. While this applies to some situations, it is possible to perform this conversion using other methods, such as FOA representation. In such embodiments, the ISM may contain some diffuse energy. Another exemplary situation in which a non-negligible diffuse-to-total energy ratio may exist in the ISM stream is when two or more ISM objects are active on the same TF tile, and the direct-to-total energy ratio is significantly less than 1. Such parameterization may be desirable to perceive space more realistically in reproduction.

[0168] If the amount of diffusion energy in the ISM is not negligible, it would be useful to consider this diffusion energy. Otherwise, for TF tiles where the ISM metadata was selected but the energy of the MASA stream was less than that of the ISM stream, the ratio value may unintentionally increase. This can be perceived as a loss of breadth in the merged stream.

[0169] The consideration of diffusion energy can be employed in many forms. For example, the resulting direct-to-total energy ratio.

number

[0170]

number

[0171] Based on the example shown in Figure 6, a further example considering diffusion energy is shown in Figure 10. In this figure, the merged diffusion-total ratio determiner 1005 is modified as follows:

number

number

number

number

number

[0172] Furthermore, Figure 10 shows not only the direct versus total ratio, but also the diffusion versus total ratio passed to the merged diffusion versus total ratio determiner 1005.

number

[0173] This example is based on the example shown in Figure 6, but similar modifications to the determination of the merged diffusion to total ratio can be made with respect to the exemplary embodiment shown in Figure 8.

[0174] Figure 11 shows a further example where two directions exist simultaneously. This can be further extended to even more directional fields.

[0175] Therefore, the direct-to-overall ratio acquirer 1101 obtains two direct-to-overall energy ratios in the original MASA metadata stream.

number

number

number

[0176] The weight determination unit 1105 is,

number

number

[0177]

number

number

[0178] The weights are compared by the weight comparator 715 and then selected by the metadata selector based on the following: w 1,1 > w2 and w 1,2 > w2, use the parameters of "MASA stream 1"; w 1,1 ≦ w 1,2 < w2 or w 1,1 < w2 ≦ w 1,2 In the case of, use the parameters from "MASA direction 2" and "ISM stream". w 1,2 ≦ w 1,1 < w2 or w 1,2 < w2 ≦ w 1,1 In the case of, use the parameters from "MASA direction 1" and "ISM stream".

[0179] In some embodiments, this can be extended by configuring the diffusion signal energy acquirer 1103 to determine the diffusion signal energy of MASA using the following:

Number

Number

[0180] Furthermore, the merged diffusion pair overall ratio determiner 1105 is configured to determine the diffusion pair overall energy ratio of the combined signal

Number

Number

[0181] In addition, the merged direct pair overall ratio determiner 1107 then, depending on this new diffusion pair overall energy ratio, newly assigns the direct pair overall energy ratio using the following

number

number

number

number

[0182] Similar to the previous example, this two-directional embodiment can be extended to include the diffusion energy of the ISM object in the same manner as the embodiments described in Figures 8 and 10.

[0183] In some embodiments, metadata merging or joining is selected based on determining a measure of "object similarity" from the spatial metadata of the streams being merged. In other words, weights are determined based on the object similarity parameter.

[0184] In embodiments where both streams have low object similarity, for example, contain a non-negligible amount of diffusion energy, a merging method as shown in Figure 6 is performed.

[0185] If the object similarity between both streams is high, for example, if the diffusion energy of both is negligible, a merging method as shown in Figure 6 is performed.

[0186] If one stream has low object similarity and the other stream has high object similarity, the merging method shown in any of Figures 9 onwards, or as shown in the embodiments described above, can be implemented.

[0187] In some embodiments, the "object similarity" parameter can be determined by examining the overall diffusion energy (objects are usually highly directional, with minimal or no diffusion) and also by examining the directional variation with respect to frequency, since objects usually have no variation with respect to frequency.

[0188] In some embodiments, the extension of the exemplary embodiments described above also applies when metadata from "MASA Stream 1" is selected, resulting in a modified direct vs. total energy ratio.

number

[0189] All of the above explanations describe the case of merging two streams. If there are three or more streams to merge, the technique can be extended in a simple way, for example, by repeatedly merging streams in pairs (or cascading) until the desired number of streams remain, or by extending the above method to include the weights of all streams simultaneously and selecting the metadata of the stream with the highest weight.

[0190] In the example above, a complex-valued low-latency filter bank (CLDFB) was shown as the frequency domain representation, but other methods of time-frequency domain representation can be used, such as the short-time Fourier transform (STFT) or quadrature mirror filter bank (QMF).

[0191] Although the parametric format described above was the MASA format, the embodiment can be extended to other parametric formats such as parametric coding for ambisonics and multi-channel mixes.

[0192] An illustrative electronic device in Figure 12 can be used as one of the device components of the system described above. The device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. The device may be configured to implement, for example, an encoder and / or decoder, or any functional block described above.

[0193] In some embodiments, the device 1400 includes at least one processor or central processing unit 1407. The processor 1407 may be configured to execute various program codes, such as those described herein.

[0194] In some embodiments, the device 1400 includes at least one memory 1411. In some embodiments, at least one processor 1407 is coupled to the memory 1411. The memory 1411 can be any suitable storage means. In some embodiments, the memory 1411 includes a program code section for storing program code executable by the processor 1407. Furthermore, in some embodiments, the memory 1411 may further include a storage data section for storing data processed or to be processed, for example, according to embodiments described herein. The executable program code stored in the program code section and the data stored in the storage data section can be retrieved by the processor 1407 at any time as needed via the memory-processor coupling.

[0195] In some embodiments, device 1400 includes a user interface 1405. The user interface 1405 may be coupled to the processor 1407 in some embodiments. In some embodiments, the processor 1407 can control the operation of the user interface 1405 and receive inputs from the user interface 1405. In some embodiments, the user interface 1405 can enable a user to input commands to the device 1400, for example, via a keypad. In some embodiments, the user interface 1405 can enable a user to obtain information from the device 1400. For example, the user interface 1405 can include a display configured to display information from the device 1400 to the user. The user interface 1405 can include a touch screen or touch interface in some embodiments that enables both inputting information into the device 1400 and displaying information to the user of the device 1400. In some embodiments, the user interface 1405 may be a user interface for communication.

[0196] In some embodiments, device 1400 includes an input / output port 1409. In some embodiments, the input / output port 1409 includes a transceiver. The transceiver in such embodiments is coupled to the processor 1407 and can be configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. The transceiver or any suitable transceiver, or transmitter and / or receiver means, can be configured to communicate with other electronic devices or apparatuses via a wired connection or a wired coupling in some embodiments.

[0197] The transceiver can communicate with further devices by means of any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable wireless access architecture based on: Long Term Evolution Advanced (LTE Advanced, LTE-A) or New Radio (NR) (which may also be referred to as 5G), Universal Mobile Telecommunications System (UMTS) radio access network (UTRAN or E-UTRAN), Long Term Evolution (LTE, identical to E-UTRA), 2G network (legacy network technology), Wireless Local Area Network (WLAN or Wi-Fi), Worldwide Interoperability for Microwave Access (WiMAX), Bluetooth®, Personal Communication Service (PCS), ZigBee®, Wideband Code Division Multiple Access (WCDMA), systems using Ultra-Wideband (UWB) technology, sensor networks, Mobile Ad hoc Networks (MANETs), Cellular Internet of Things (IoT) RAN, and Internet Protocol Multimedia Subsystem (IMS), any other suitable options and / or combinations thereof.

[0198] The transceiver input / output port 1409 may be configured to receive signals.

[0199] In some embodiments, the device 1400 may be employed as at least part of a combined device. The input / output port 1409 may be coupled to headphones (which may be head-tracking headphones or non-tracking headphones) or the like, and loudspeakers.

[0200] In general, various embodiments of the present invention can be implemented in hardware or special-purpose circuits, software, logic, or a combination thereof. For example, some embodiments may be implemented in hardware, while others may be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Various embodiments of the present invention can be illustrated and described using block diagrams, flowcharts, or any other graphic representation, but it will be understood that these blocks, devices, systems, techniques, or methods described herein can be implemented, in non-limiting examples, in hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controllers or other computing devices, or any combination thereof.

[0201] Embodiments of the present invention can be implemented by computer software or hardware, or a combination of software and hardware, that can be executed by the data processor of a mobile device, such as in a processor entity. Furthermore, it should be noted that any block of the logic flow shown in the drawings can represent a program step, or an interconnected set of logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software can be stored on physical media such as memory chips or memory blocks executed within a processor, magnetic media such as hard disks and floppy disks, and optical media such as DVDs and their data variations such as CDs.

[0202] Memory may be of any type suitable for the local technical environment and can be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processor may be of any type suitable for the local technical environment and, in non-limiting examples, may be one or more of general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and processors based on multi-core processor architectures.

[0203] Embodiments of the present invention can be implemented in various components, such as integrated circuit modules. Designing integrated circuits is generally a highly automated process. Complex and powerful software tools can be used to translate logic-level designs into semiconductor circuit designs prepared for etching onto semiconductor substrates.

[0204] Programs such as those offered by Synopsys, Inc. in Mountain View, California, and Cadence Design, Inc. in San Jose, California, use well-established design rules and libraries of pre-stored design modules to automatically route conductors and place components on a semiconductor chip. Once the semiconductor circuit design is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be sent to a semiconductor manufacturing facility, or "fab," for production.

[0205] As used in this application, the term "circuit" may refer to one or more or all of the following: (a) Hardware-only circuit implementations (such as implementations of analog and / or digital circuits only), and (b) Combinations of hardware circuits and software (as needed), for example: (i) combination of analog and / or digital hardware circuits and software / firmware, (ii) A part of a hardware processor with software (including a digital signal processor), software, and memory that cooperate to perform various functions on a device such as a mobile phone or server. A hardware circuit and / or processor, such as a microprocessor or a part of a microprocessor, that requires software (e.g., firmware) to operate, but whose presence is not required when the software is not needed for operation.

[0206] This definition of circuit applies to all use of the term in this application, including in the claims. Further as used in this application, the term circuit also includes, by definition, a hardware circuit or processor (or more processors) or a part of a hardware circuit or processor, as well as any software and / or firmware associated with it (or them). The term circuit also includes, for example, a baseband integrated circuit or processor-integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or network device, where applicable to a particular claim element.

[0207] As used herein, the term “non-transient” refers to the limitations of the medium itself (i.e., tangible, rather than signal-based) as opposed to the limitations of data storage persistence (e.g., RAM vs. ROM).

[0208] Where used herein, "at least one of the following: <list of two or more elements>" and "at least one of the <list of two or more elements>", and similar phrasing, when a list of two or more elements is connected by "and" or "or", this means at least one of the elements, or at least two or more of the elements, or at least all of the elements.

[0209] The foregoing description, by illustrative and non-limiting examples, has provided a complete and useful description of exemplary embodiments of the present invention. However, various modifications and adaptations will become apparent to those skilled in the art in light of the foregoing description, when read in conjunction with the accompanying drawings and claims. However, all such and similar modifications of the teachings of the present invention are still included within the scope of the invention as defined in the claims.

Claims

1. For the first audio stream, obtain at least one first direct-to-whole ratio parameter. To obtain a first signal energy parameter for the first audio stream, To generate at least one first weight value based on the at least one first direct-to-whole ratio parameter and the first signal energy parameter, For the second audio stream, obtain the second direct-to-overall ratio parameter. To obtain a second signal energy parameter for the second audio stream, To generate a direct versus overall ratio parameter for diffusion energy compensation, Partially The second direct-to-overall ratio parameter mentioned above, The aforementioned diffusion energy compensation direct versus overall ratio parameter, and The second signal energy parameter To generate a second weight value based on at least one of the following, and Based on a comparison of the at least one first weight value and the second weight value, one of the at least one first direct-to-whole ratio parameter and the diffusion energy compensation direct-to-whole ratio parameter is selected. A device equipped with means for that purpose.

2. The means for selecting one of the at least one first weight value and the second weight value, based on a comparison of the first weight value and the second weight value, further, If the at least one first weight value is greater than the second weight value, select one of the at least one first direct-to-overall ratio parameters, and The diffusion energy compensation direct versus overall ratio parameter is selected when the second weight value is greater than the at least one first weight value. The apparatus according to claim 1, which is for the purpose of

3. The apparatus according to claim 2, wherein, when the at least one first weight value is greater than the second weight value, the means for selecting one of the at least one first direct-to-whole ratio parameters is such that the at least one first weight value is strictly greater than the second weight value.

4. The apparatus according to claim 2, wherein the means for selecting the diffusion energy compensation direct versus overall ratio parameter is such that the second weight value is strictly greater than the at least one first weight value when the second weight value is greater than the at least one first weight value.

5. The apparatus according to any one of claims 1 to 4, wherein the means for generating the at least one first weight value based on the at least one first direct-to-whole ratio parameter and the first signal energy parameter is for generating the at least one first weight value based on the multiplication of the at least one first direct-to-whole ratio parameter and the first signal energy parameter.

6. The apparatus according to any one of claims 1 to 5, wherein the at least one first direct-to-overall ratio comprises at least two first direct-to-overall ratios, and the means for generating the at least one first weight value based on the at least one first direct-to-overall ratio parameter and the first signal energy parameter is for generating a first weight value associated with each of the at least two first direct-to-overall ratio parameters.

7. The apparatus according to claim 6, wherein the means for generating a first weight value associated with each of the at least two first direct-to-overall ratio parameters is for generating a first weight value based on the multiplication of each of the at least one first direct-to-overall ratio parameter with the first signal energy parameter.

8. The means for selecting one of the at least one first weight value and the second weight value, based on a comparison of the first weight value and the second weight value, further, Selecting the at least two first direct-to-overall ratio parameters when both of the first weight values ​​are greater than the second weight values, and If the second weight value is greater than one of the at least two first weight values, select the diffusion energy compensation direct-to-overall ratio parameter and one of the at least two first direct-to-overall ratio parameters. The apparatus according to claim 6 or 7, which is for the purpose of

9. The means for generating a second weight value, partially based on at least one of the second direct-to-whole ratio parameter, the diffusion energy compensated direct-to-whole ratio parameter, and the second signal energy parameter, To generate the at least one second weight value based on the multiplication of the at least one second direct-to-whole ratio parameter and the second signal energy parameter, To generate the at least one second weight value based on the multiplication of the diffusion energy compensation direct versus overall ratio parameter and the second signal energy parameter, and The at least one second weight value is generated based on the average of the diffusion energy compensation direct-to-overall ratio parameter and the at least one second direct-to-overall ratio parameter obtained by multiplying the second signal energy parameter. The apparatus according to any one of claims 1 to 8, which is for one of the following.

10. The apparatus according to any one of claims 1 to 9, wherein the means for generating a diffusion energy compensated direct-to-whole ratio parameter is for generating at least one merged diffusion-to-whole ratio based at least in part on the first signal energy parameter, the second signal energy parameter and the diffusion signal energy parameter.

11. The means further, With respect to the first audio stream, obtain at least one first diffusion-to-whole ratio parameter, and The diffusion signal energy parameter is obtained in part based on the first signal energy parameter and the first diffusion-to-total ratio parameter. The apparatus according to claim 10, which is for the purpose of

12. The apparatus according to claim 10 or 11, wherein the means for generating a diffusion energy compensated direct-to-whole ratio parameter further comprises generating the at least one diffusion energy compensated direct-to-whole ratio parameter based on the at least one merged diffusion-to-whole ratio.

13. The apparatus according to claim 10 or 11, wherein the means further comprises obtaining at least one second diffusion-to-overall ratio parameter for the second audio stream, and the means for generating at least one merged diffusion-to-overall ratio further comprises generating the at least one merged diffusion-to-overall ratio based on the second diffusion-to-overall ratio parameter.

14. The means for generating the diffusion energy compensation direct versus overall ratio parameter further, To generate at least one expected diffusion energy compensation direct versus total ratio parameter based on the at least one merged diffusion versus total ratio, Comparing the at least one expected diffusion energy compensation direct-to-overall ratio parameter with the second direct-to-overall ratio parameter, Based on the comparison, the diffusion energy compensation direct-to-overall ratio parameter is selected from the at least one expected diffusion energy compensation direct-to-overall ratio parameter and the second direct-to-overall ratio parameter such that the diffusion energy compensation direct-to-overall ratio parameter is the smaller of the at least one expected diffusion energy compensation direct-to-overall ratio parameter and the second direct-to-overall ratio parameter. The apparatus according to claim 10, which is for the purpose of

15. The apparatus according to any one of claims 1 to 14, wherein the means further comprises encoding the selected one of the at least one first direct-to-whole ratio parameter and the diffusion energy compensated direct-to-whole ratio parameter.

16. The apparatus according to any one of claims 1 to 15, wherein the means further comprises merging the first audio stream and the second audio stream.

17. For the first audio stream, obtain at least one first direct-to-whole ratio parameter, With respect to the first audio stream, the first signal energy parameter is obtained, To generate at least one first weight value based on the at least one first direct-to-whole ratio parameter and the first signal energy parameter, For the second audio stream, obtain the second direct-to-overall ratio parameter, With respect to the second audio stream, the second signal energy parameter is obtained, To generate a direct versus overall ratio parameter for diffusion energy compensation, Partially The second direct-to-overall ratio parameter mentioned above, The aforementioned diffusion energy compensation direct versus overall ratio parameter, and The second signal energy parameter A second weight value is generated based on at least one of the following: Based on a comparison of the at least one first weight value and the second weight value, one of the at least one first direct-to-overall ratio parameter and the diffusion energy compensation direct-to-overall ratio parameter is selected. Methods that include...

18. Based on a comparison of the at least one first weight value and the second weight value, one of the at least one first direct-to-whole ratio parameter and the diffusion energy compensation direct-to-whole ratio parameter is selected. If the at least one first weight value is greater than the second weight value, select one of the at least one first direct-to-overall ratio parameters, The diffusion energy compensation direct versus overall ratio parameter is selected when the second weight value is greater than the at least one first weight value. The method according to claim 17, further comprising:

19. The method according to claim 18, wherein the selection of one of the at least one first direct-to-whole ratio parameters when the at least one first weight value is greater than the second weight value is strictly defined as the at least one first weight value being greater than the second weight value.

20. The method according to claim 18, wherein the selection of the diffusion energy compensation direct versus total ratio parameter occurs when the second weight value is greater than the at least one first weight value, and the second weight value is strictly greater than the at least one first weight value.

21. The method according to any one of claims 17 to 20, wherein generating the at least one first weight value based on the at least one first direct-to-whole ratio parameter and the first signal energy parameter further comprises generating the at least one first weight value based on the multiplication of the at least one first direct-to-whole ratio parameter and the first signal energy parameter.

22. The method according to any one of claims 17 to 21, wherein the at least one first direct-to-overall ratio comprises at least two first direct-to-overall ratios, and generating the at least one first weight value based on the at least one first direct-to-overall ratio parameter and the first signal energy parameter further comprises generating a first weight value associated with each of the at least two first direct-to-overall ratio parameters.

23. The method according to claim 22, wherein generating a first weight value associated with each of the at least two first direct-to-overall ratio parameters further comprises generating a first weight value based on the multiplication of each of the at least one first direct-to-overall ratio parameter with the first signal energy parameter.

24. The method according to claim 22 or 23, wherein selecting one of the at least one first direct-to-whole ratio parameter and the diffusion energy compensation direct-to-whole ratio parameter based on a comparison of the at least one first weight value and the second weight value further includes selecting the at least two first direct-to-whole ratio parameters when both of the first weight values ​​are greater than the second weight values, and selecting the diffusion energy compensation direct-to-whole ratio parameter and one of the at least two first direct-to-whole ratio parameters when the second weight value is greater than one of the at least two first weight values.

25. A second weight value can be generated partially based on at least one of the second direct-to-whole ratio parameter, the diffusion energy compensated direct-to-whole ratio parameter, and the second signal energy parameter. To generate the at least one second weight value based on the multiplication of the at least one second direct-to-whole ratio parameter and the second signal energy parameter, To generate the at least one second weight value based on the multiplication of the diffusion energy compensation direct versus overall ratio parameter and the second signal energy parameter, and The at least one second weight value is generated based on the average of the diffusion energy compensation direct-to-overall ratio parameter and the at least one second direct-to-overall ratio parameter obtained by multiplying the second signal energy parameter. The method according to any one of claims 17 to 24, further comprising one of the following.

26. The method according to any one of claims 17 to 25, further comprising generating a diffusion energy compensated direct versus whole ratio parameter, which in turn generates at least one merged diffusion versus whole ratio based at least in part on the first signal energy parameter, the second signal energy parameter, and the diffusion signal energy parameter.

27. With respect to the first audio stream, obtain at least one first diffusion-to-whole ratio parameter, and The diffusion signal energy parameter is obtained in part based on the first signal energy parameter and the first diffusion-to-total ratio parameter. The method according to claim 26, further comprising:

28. The method according to claim 26 or 27, wherein generating a diffusion energy compensated direct-to-whole ratio parameter further comprises generating the at least one diffusion energy compensated direct-to-whole ratio parameter based on the at least one merged diffusion-to-whole ratio.

29. The method according to claim 26 or 27, further comprising obtaining at least one second diffusion-to-overall ratio parameter for the second audio stream, and generating at least one merged diffusion-to-overall ratio, further comprising generating the at least one merged diffusion-to-overall ratio based on the second diffusion-to-overall ratio parameter.

30. The generation of the diffusion energy compensation direct versus overall ratio parameter is, To generate at least one expected diffusion energy compensation direct versus total ratio parameter based on the at least one merged diffusion versus total ratio, Comparing the at least one expected diffusion energy compensation direct-to-overall ratio parameter with the second direct-to-overall ratio parameter, Based on the comparison, the diffusion energy compensation direct-to-overall ratio parameter is selected from the at least one expected diffusion energy compensation direct-to-overall ratio parameter and the second direct-to-overall ratio parameter such that the diffusion energy compensation direct-to-overall ratio parameter is the smaller of the at least one expected diffusion energy compensation direct-to-overall ratio parameter and the second direct-to-overall ratio parameter. The method according to claim 26, further comprising:

31. The method according to any one of claims 17 to 30, further comprising encoding the selected one of the at least one first direct-to-whole ratio parameter and the diffusion energy compensated direct-to-whole ratio parameter.

32. The method according to any one of claims 17 to 31, further comprising merging the first audio stream and the second audio stream.