Diffusion retention merging of MASA and ISM metadata

By obtaining and comparing the weighted values ​​of the direct-to-total ratio parameters and merging MASA and ISM metadata, the problem of low coding efficiency of immersive audio codecs when processing spatial sounds captured by microphone arrays is solved, achieving efficient low-bitrate transmission and high-quality rendering.

CN120752699APending Publication Date: 2025-10-03NOKIA TECHNOLOGIES OY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480014214.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-23
Filing Date
2024-02-01
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing immersive audio codecs have difficulty effectively merging MASA and ISM metadata when processing spatial sounds captured by microphone arrays, resulting in low coding efficiency and the inability to efficiently transmit and render high-quality spatial audio at low bit rates.

Method used

By obtaining the direct-to-total ratio parameters and signal energy parameters, generating weighted values ​​and comparing them, selecting appropriate direct-to-total ratio parameters, generating direct-to-total ratio parameters that compensate for diffuse energy, realizing the merging of MASA and ISM metadata, and optimizing the encoding process.

Benefits of technology

Improved coding efficiency enables efficient transmission and rendering of high-quality spatial audio at low bit rates, reduces coding complexity, and improves audio signal quality and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120752699A_ABST
    Figure CN120752699A_ABST
Patent Text Reader

Abstract

An apparatus comprises means for: acquiring at least one first direct to total ratio parameter for a first audio stream; acquiring a first signal energy parameter for the first audio stream; generating at least one first weighting value based on the at least one first direct to total ratio parameter and the first signal energy parameter; acquiring a second direct to total ratio parameter for the second audio stream; acquiring a second signal energy parameter for the second audio stream; generating a direct to total ratio parameter that compensates for diffused energy; generating a second weighting value based in part on at least one of the second direct to total ratio parameter, the direct to total ratio parameter of the compensated diffused energy, and the second signal energy parameter; and selecting one of the at least one first direct to total ratio parameter and a direct to total ratio parameter that compensates for diffused energy based on a comparison of the at least one first weighted value and the second weighted value. (Figure 1) 20
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to apparatus and methods for merging MASA and ISM metadata with the aim of preserving diffuse parameters, not just for audio coding. Background Art

[0002] Parametric spatial audio capture from inputs (such as microphone arrays and other sources) is a typical and effective choice for estimating a set of parameters, such as the direction of the sound in a frequency band and the ratio between the directional and non-directional parts of the captured sound in the frequency band, from the input (microphone array signal). These parameters are known to describe the perceived spatial characteristics of the captured sound at the location of the microphone array very well. These parameters can be used accordingly to synthesize spatial sound for binaural headphones, for speakers, or other formats such as Ambisonics.

[0003] Therefore, the directionality in the frequency band and the direct to total energy ratio and the diffuse to total energy ratio are a particularly effective parameterization for spatial audio capture.

[0004] A parameter set consisting of a directional parameter in a frequency band and an energy ratio parameter in a frequency band (indicating the directionality of the sound) can also be used as spatial metadata for the audio codec (which may also include other parameters such as surround coherence, spread coherence, number of directions, distance, etc.). For example, these parameters can be estimated from an audio signal captured by a microphone array, and, for example, a stereo or mono signal can be generated from the microphone array signal to be conveyed using the spatial metadata.

[0005] Immersive audio codecs are being implemented to support multiple operating points, from low bitrate operation to transparency. An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, which is designed to be suitable for use over communication networks such as 3GPP 4G / 5G networks, including use in immersive services such as immersive voice and audio for virtual reality (VR). This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general audio. In addition, it is expected to support channel-based audio input, object-based audio input, and scene-based audio input, including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable conversational services and support high error robustness under various transmission conditions.

[0006] For example, a stereo signal can be encoded using the IVAS audio core codec or with an AAC (Advanced Audio Coding) or EVS (Enhanced Voice Services) encoder. The decoder can decode the audio signal into a PCM (Pulse Code Modulation) signal and process the sound in frequency bands (using spatial metadata) to obtain a spatial output, such as a binaural output.

[0007] The immersive audio codecs described above are particularly well-suited for encoding spatial sound captured from microphone arrays (e.g., in mobile phones, VR cameras, or standalone microphone arrays). However, such encoders can have other input types, such as loudspeaker signals, audio object signals, and ambisonic signals. Summary of the Invention

[0008] According to a first aspect, an apparatus is provided, comprising components for obtaining at least one first direct-to-total ratio parameter for a first audio stream; obtaining a first signal energy parameter for the first audio stream; generating at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; obtaining a second direct-to-total ratio parameter for a second audio stream; obtaining a second signal energy parameter for the second audio stream; generating a diffuse-energy compensated direct-to-total ratio parameter; generating a second weighting value based in part on at least one of the second direct-to-total ratio parameter, the diffuse-energy compensated direct-to-total ratio parameter, and the second signal energy parameter; and selecting one of the at least one first direct-to-total ratio parameter and the diffuse-energy compensated direct-to-total ratio parameter based on a comparison of the at least one first weighting value and the second weighting value.

[0009] The component for selecting one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating diffuse energy based on a comparison of at least one first weighted value and a second weighted value can also be used to: select one of the at least one first direct-to-total ratio parameter when the at least one first weighted value is greater than the second weighted value; and select the direct-to-total ratio parameter for compensating diffuse energy when the second weighted value is greater than the at least one first weighted value.

[0010] The means for selecting one of the at least one first direct-to-total ratio parameters when the at least one first weighting value is greater than the second weighting value may be when the at least one first weighting value is strictly greater than the second weighting value.

[0011] The means for selecting the direct to total ratio parameter to compensate for diffuse energy when the second weighting value is greater than the at least one first weighting value may be when the second weighting value is strictly greater than the at least one first weighting value.

[0012] The means for generating at least one first weighting value based on at least one first direct-to-total ratio parameter and the first signal energy parameter may be configured to generate at least one first weighting value based on a product of at least one first direct-to-total ratio parameter and the first signal energy parameter.

[0013] At least one first direct-to-total ratio may include at least two first direct-to-total ratios, wherein the component for generating at least one first weighted value based on at least one first direct-to-total ratio parameter and a first signal energy parameter may be used to generate a first weighted value associated with each of the at least two first direct-to-total ratio parameters.

[0014] The means for generating a first weighting value associated with each of the at least two first direct-to-total ratio parameters may be configured to generate each first weighting value based on a product of each of the at least one first direct-to-total ratio parameter and a first signal energy parameter.

[0015] The component for selecting one of the at least one first direct-to-total ratio parameters and the direct-to-total ratio parameter for compensating for diffuse energy based on a comparison of at least one first weighted value and a second weighted value can also be used to: select at least two first direct-to-total ratio parameters when both first weighted values ​​are greater than the second weighted value; and select the direct-to-total ratio parameter for compensating for diffuse energy and one of the at least two first direct-to-total ratio parameters when the second weighted value is greater than one of the at least two first weighted values.

[0016] The component for generating a second weighted value based in part on at least one of a second direct-to-total ratio parameter, a direct-to-total ratio parameter for compensating diffuse energy, and a second signal energy parameter can be used for one of the following: generating at least one second weighted value based on the product of at least one second direct-to-total ratio parameter and a second signal energy parameter; generating at least one second weighted value based on the product of the direct-to-total ratio parameter for compensating diffuse energy and the second signal energy parameter; and generating at least one second weighted value based on the product of an average of the direct-to-total ratio parameter for compensating diffuse energy and at least one second direct-to-total ratio parameter and the second signal energy parameter.

[0017] The means for generating a diffuse energy compensated direct to total ratio parameter may be configured to generate at least one combined diffuse to total ratio based at least in part on the first signal energy parameter, the second signal energy parameter, and the diffuse signal energy parameter.

[0018] The component may also be configured to: obtain at least one first diffuse-to-total ratio parameter for the first audio stream; and obtain a diffuse signal energy parameter based in part on the first signal energy parameter and the first diffuse-to-total ratio parameter.

[0019] The means for generating a direct to total ratio parameter for compensating diffuse energy may also be configured to generate at least one direct to total ratio parameter for compensating diffuse energy based on the at least one combined diffuse to total ratio.

[0020] The means may also be configured to obtain at least one second diffuse-to-total ratio parameter for the second audio stream, wherein the means for generating at least one combined diffuse-to-total ratio may also be configured to generate at least one combined diffuse-to-total ratio based on the second diffuse-to-total ratio parameter.

[0021] The component for generating a direct-to-total ratio parameter for compensating diffuse energy can also be used to: generate at least one expected direct-to-total ratio parameter for compensating diffuse energy based on at least one combined diffuse-to-total ratio; compare the at least one expected direct-to-total ratio parameter for compensating diffuse energy with a second direct-to-total ratio parameter; and based on the comparison, select the direct-to-total ratio parameter for compensating diffuse energy from the at least one expected direct-to-total ratio parameter for compensating diffuse energy and the second direct-to-total ratio parameter, so that the direct-to-total ratio parameter for compensating diffuse energy is the smaller of the at least one expected direct-to-total ratio parameter for compensating diffuse energy and the second direct-to-total ratio parameter.

[0022] The component may also be configured to encode a selected first direct-to-total ratio parameter of the at least one first direct-to-total ratio parameter and a direct-to-total ratio parameter compensating for diffuse energy.

[0023] The component may also be used to merge the first audio stream and the second audio stream.

[0024] According to a second aspect, a method is provided, which includes: obtaining at least one first direct-to-total ratio parameter for a first audio stream; obtaining a first signal energy parameter for the first audio stream; generating at least one first weighting value based on at least one first direct-to-total ratio parameter and the first signal energy parameter; obtaining a second direct-to-total ratio parameter for a second audio stream; obtaining a second signal energy parameter for the second audio stream; generating a direct-to-total ratio parameter for compensating for diffuse energy; generating a second weighting value based in part on at least one of the second direct-to-total ratio parameter, the direct-to-total ratio parameter for compensating for diffuse energy and the second signal energy parameter; and selecting one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating for diffuse energy based on a comparison of the at least one first weighting value and the second weighting value.

[0025] Based on the comparison of at least one first weighted value and a second weighted value, selecting one of the at least one first direct-to-total ratio parameters and the direct-to-total ratio parameter for compensating diffuse energy may also include: when at least one first weighted value is greater than the second weighted value, selecting one of the at least one first direct-to-total ratio parameters; and when the second weighted value is greater than at least one first weighted value, selecting the direct-to-total ratio parameter for compensating diffuse energy.

[0026] Selecting one of the at least one first direct-to-total ratio parameters when the at least one first weighting value is greater than the second weighting value may be when the at least one first weighting value is strictly greater than the second weighting value.

[0027] Selecting the direct to total ratio parameter to compensate for diffuse energy when the second weighting value is greater than the at least one first weighting value may be when the second weighting value is strictly greater than the at least one first weighting value.

[0028] Generating at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter may include generating at least one first weighting value based on a product of the at least one first direct-to-total ratio parameter and the first signal energy parameter.

[0029] The at least one first direct-to-total ratio may include at least two first direct-to-total ratios, wherein generating at least one first weighted value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter may include generating a first weighted value associated with each of the at least two first direct-to-total ratio parameters.

[0030] Generating a first weighting value associated with each of the at least two first direct-to-total ratio parameters may include generating each first weighting value based on a product of each of the at least one first direct-to-total ratio parameter and a first signal energy parameter.

[0031] Based on the comparison of at least one first weighted value and a second weighted value, selecting one of at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating diffuse energy may include: when both first weighted values ​​are greater than the second weighted value, selecting at least two first direct-to-total ratio parameters; and when the second weighted value is greater than one of the at least two first weighted values, selecting the direct-to-total ratio parameter for compensating diffuse energy and one of the at least two first direct-to-total ratio parameters.

[0032] Generating a second weighted value based in part on at least one of a second direct-to-total ratio parameter, a direct-to-total ratio parameter for compensating diffuse energy, and a second signal energy parameter may include one of the following: generating at least one second weighted value based on the product of at least one second direct-to-total ratio parameter and a second signal energy parameter; generating at least one second weighted value based on the product of the direct-to-total ratio parameter for compensating diffuse energy and the second signal energy parameter; and generating at least one second weighted value based on the product of the average of the direct-to-total ratio parameter for compensating diffuse energy and at least one second direct-to-total ratio parameter and the second signal energy parameter.

[0033] Generating a direct-to-total ratio parameter that compensates for diffuse energy may include generating at least one combined diffuse-to-total ratio based at least in part on the first signal energy parameter, the second signal energy parameter, and the diffuse signal energy parameter.

[0034] The method may further include obtaining at least one first diffuse-to-total ratio parameter for the first audio stream; and obtaining a diffuse signal energy parameter based in part on the first signal energy parameter and the first diffuse-to-total ratio parameter.

[0035] Generating a direct to total ratio parameter that compensates for diffuse energy may include generating at least one direct to total ratio parameter that compensates for diffuse energy based on the at least one combined diffuse to total ratio.

[0036] The method may further include obtaining at least one second diffuse-to-total ratio parameter for the second audio stream, wherein generating at least one merged diffuse-to-total ratio may further include generating at least one merged diffuse-to-total ratio based on the second diffuse-to-total ratio parameter.

[0037] Generating a direct-to-total ratio parameter for compensating diffuse energy may also include: generating at least one expected direct-to-total ratio parameter for compensating diffuse energy based on at least one combined diffuse-to-total ratio; comparing the at least one expected direct-to-total ratio parameter for compensating diffuse energy with a second direct-to-total ratio parameter; and based on the comparison, selecting a direct-to-total ratio parameter for compensating diffuse energy from the at least one expected direct-to-total ratio parameter for compensating diffuse energy and the second direct-to-total ratio parameter, so that the direct-to-total ratio parameter for compensating diffuse energy is the smaller of the at least one expected direct-to-total ratio parameter for compensating diffuse energy and the second direct-to-total ratio parameter.

[0038] The method may further include encoding a selected first direct-to-total ratio parameter of the at least one first direct-to-total ratio parameter and a direct-to-total ratio parameter that compensates for diffuse energy.

[0039] The method may further include merging the first audio stream and the second audio stream.

[0040] According to a third aspect, a device is provided, which includes at least one processor and at least one memory storing instructions, which instructions, when executed by the at least one processor, cause the system to at least perform: obtaining at least one first direct-to-total ratio parameter for a first audio stream; obtaining a first signal energy parameter for the first audio stream; generating at least one first weighting value based on at least one first direct-to-total ratio parameter and the first signal energy parameter; obtaining a second direct-to-total ratio parameter for a second audio stream; obtaining a second signal energy parameter for the second audio stream; generating a direct-to-total ratio parameter for compensating for diffuse energy; generating a second weighting value based in part on at least one of the second direct-to-total ratio parameter, the direct-to-total ratio parameter for compensating for diffuse energy, and the second signal energy parameter; and selecting one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating for diffuse energy based on a comparison of the at least one first weighting value and the second weighting value.

[0041] The device that is caused to perform the selection of one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating for diffuse energy based on the comparison of at least one first weighted value and a second weighted value can also be caused to perform: when the at least one first weighted value is greater than the second weighted value, selecting one of the at least one first direct-to-total ratio parameter; and when the second weighted value is greater than the at least one first weighted value, selecting the direct-to-total ratio parameter for compensating for diffuse energy.

[0042] The means caused to perform selecting one of the at least one first direct-to-total ratio parameter when the at least one first weighting value is greater than the second weighting value may be when the at least one first weighting value is strictly greater than the second weighting value.

[0043] The direct to total ratio parameter caused to select the compensation diffuse energy when the second weighting value is greater than the at least one first weighting value may be when the second weighting value is strictly greater than the at least one first weighting value.

[0044] The device configured to generate at least one first weighted value based on at least one first direct-to-total ratio parameter and a first signal energy parameter may also be configured to generate at least one first weighted value based on a product of at least one first direct-to-total ratio parameter and the first signal energy parameter.

[0045] At least one first direct-to-total ratio may include at least two first direct-to-total ratios, wherein the device that is caused to generate at least one first weighted value based on at least one first direct-to-total ratio parameter and a first signal energy parameter may also be caused to perform: generating a first weighted value associated with each of the at least two first direct-to-total ratio parameters.

[0046] The device that is caused to generate a first weighted value associated with each first direct-to-total ratio parameter of at least two first direct-to-total ratio parameters can also be caused to generate each first weighted value based on the product of each first direct-to-total ratio parameter of at least one first direct-to-total ratio parameter and the first signal energy parameter.

[0047] The device that is caused to perform the selection of one of the at least one first direct-to-total ratio parameters and the direct-to-total ratio parameter for compensating for diffuse energy based on the comparison of at least one first weighted value and a second weighted value can also be caused to perform: when both of the first weighted values ​​are greater than the second weighted value, selecting at least two first direct-to-total ratio parameters; and when the second weighted value is greater than one of the at least two first weighted values, selecting the direct-to-total ratio parameter for compensating for diffuse energy and one of the at least two first direct-to-total ratio parameters.

[0048] The device that is caused to perform the generation of a second weighted value based in part on at least one of the second direct-to-total ratio parameter, the direct-to-total ratio parameter for compensating diffuse energy and the second signal energy parameter can also be caused to perform one of the following items: generating at least one second weighted value based on the product of at least one second direct-to-total ratio parameter and the second signal energy parameter; generating at least one second weighted value based on the product of the direct-to-total ratio parameter for compensating diffuse energy and the second signal energy parameter; and generating at least one second weighted value based on the product of the average of the direct-to-total ratio parameter for compensating diffuse energy and at least one second direct-to-total ratio parameter and the second signal energy parameter.

[0049] The means caused to perform generating the diffuse energy compensated direct to total ratio parameter may also be caused to perform generating at least one combined diffuse to total ratio based at least in part on the first signal energy parameter, the second signal energy parameter, and the diffuse signal energy parameter.

[0050] The apparatus may be further caused to perform: obtaining at least one first diffuse-to-total ratio parameter for the first audio stream; and obtaining a diffuse signal energy parameter based in part on the first signal energy parameter and the first diffuse-to-total ratio parameter.

[0051] The means caused to perform generating a direct to total ratio parameter to compensate for diffuse energy may be further caused to perform generating at least one direct to total ratio parameter to compensate for diffuse energy based on the at least one combined diffuse to total ratio.

[0052] The apparatus may further be caused to perform: obtaining at least one second diffuse-to-total ratio parameter for the second audio stream, wherein the means for performing the generation of at least one combined diffuse-to-total ratio may further be caused to perform: generating at least one combined diffuse-to-total ratio based on the second diffuse-to-total ratio parameter.

[0053] The device that is configured to generate a direct-to-total ratio parameter for compensating diffuse energy can also be configured to: generate at least one direct-to-total ratio parameter for compensating diffuse energy based on at least one combined diffuse-to-total ratio; compare the at least one direct-to-total ratio parameter for compensating diffuse energy with a second direct-to-total ratio parameter; and based on the comparison, select the direct-to-total ratio parameter for compensating diffuse energy from the at least one direct-to-total ratio parameter for compensating diffuse energy and the second direct-to-total ratio parameter, so that the direct-to-total ratio parameter for compensating diffuse energy is the smaller of the at least one direct-to-total ratio parameter for compensating expected diffuse energy and the second direct-to-total ratio parameter.

[0054] The apparatus may be further configured to perform encoding a selected first direct-to-total ratio parameter of the at least one first direct-to-total ratio parameter and a direct-to-total ratio parameter that compensates for diffuse energy.

[0055] The apparatus may be further caused to perform: merging the first audio stream and the second audio stream.

[0056] According to a fourth aspect, a device is provided, comprising: an acquisition circuit system configured to acquire at least one first direct-to-total ratio parameter for a first audio stream; an acquisition circuit system configured to acquire a first signal energy parameter for the first audio stream; a generation circuit system configured to generate at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; an acquisition circuit system configured to acquire a second direct-to-total ratio parameter for a second audio stream; an acquisition circuit system configured to acquire a second signal energy parameter for the second audio stream; a generation circuit system configured to generate a direct-to-total ratio parameter for compensating for diffuse energy; a generation circuit system configured to generate a second weighting value based in part on at least one of the second direct-to-total ratio parameter, the direct-to-total ratio parameter for compensating for diffuse energy, and the second signal energy parameter; and a selection circuit system configured to select one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating for diffuse energy based on a comparison of the at least one first weighting value and the second weighting value.

[0057] According to a fifth aspect, a computer program [or a computer-readable medium comprising program instructions] is provided, the instructions being configured to cause an apparatus to at least: obtain at least one first direct-to-total ratio parameter for a first audio stream; obtain a first signal energy parameter for the first audio stream; generate at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; obtain a second direct-to-total ratio parameter for a second audio stream; obtain a second signal energy parameter for the second audio stream; generate a direct-to-total ratio parameter for compensating for diffuse energy; generate a second weighting value based in part on at least one of the second direct-to-total ratio parameter, the direct-to-total ratio parameter for compensating for diffuse energy, and the second signal energy parameter; and select, based on a comparison of the at least one first weighting value and the second weighting value, one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating for diffuse energy.

[0058] According to a sixth aspect, a non-transitory computer-readable medium is provided, which includes program instructions for causing a device to at least execute: obtaining at least one first direct-to-total ratio parameter for a first audio stream; obtaining a first signal energy parameter for the first audio stream; generating at least one first weighting value based on at least one first direct-to-total ratio parameter and the first signal energy parameter; obtaining a second direct-to-total ratio parameter for a second audio stream; obtaining a second signal energy parameter for the second audio stream; generating a direct-to-total ratio parameter for compensating for diffuse energy; generating a second weighting value based in part on at least one of the second direct-to-total ratio parameter, the direct-to-total ratio parameter for compensating for diffuse energy, and the second signal energy parameter; and selecting one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating for diffuse energy based on a comparison of the at least one first weighting value and the second weighting value.

[0059] According to the seventh aspect, a device is provided, which includes: a component for obtaining at least one first direct-to-total ratio parameter for a first audio stream; a component for obtaining a first signal energy parameter for the first audio stream; a component for generating at least one first weighting value based on at least one first direct-to-total ratio parameter and the first signal energy parameter; a component for obtaining a second direct-to-total ratio parameter for a second audio stream; a component for obtaining a second signal energy parameter for the second audio stream; a component for generating a direct-to-total ratio parameter for compensating for diffuse energy; a component for generating a second weighting value based in part on at least one of the second direct-to-total ratio parameter, the direct-to-total ratio parameter for compensating for diffuse energy and the second signal energy parameter; and a component for selecting one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating for diffuse energy based on a comparison of at least one first weighting value and the second weighting value.

[0060] According to an eighth aspect, a computer-readable medium is provided, which includes program instructions for causing an apparatus to at least perform the following: obtaining at least one first direct-to-total ratio parameter for a first audio stream; obtaining a first signal energy parameter for the first audio stream; generating at least one first weighting value based on at least one first direct-to-total ratio parameter and the first signal energy parameter; obtaining a second direct-to-total ratio parameter for a second audio stream; obtaining a second signal energy parameter for the second audio stream; generating a direct-to-total ratio parameter for compensating for diffuse energy; generating a second weighting value based in part on at least one of the second direct-to-total ratio parameter, the direct-to-total ratio parameter for compensating for diffuse energy, and the second signal energy parameter; and selecting one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating for diffuse energy based on a comparison of the at least one first weighting value and the second weighting value.

[0061] An apparatus comprising means for performing the actions of the above method.

[0062] An apparatus configured to perform the actions of the above method.

[0063] A computer program comprising program instructions for causing a computer to execute the above method.

[0064] A computer program product stored on a medium may cause an apparatus to perform the methods described herein.

[0065] An electronic device may comprise an apparatus as described herein.

[0066] A chipset may include apparatus as described herein.

[0067] The embodiments of the present application are intended to solve the problems associated with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings, in which:

[0069] Figure 1 Schematically illustrates an apparatus for MASA metadata extraction;

[0070] Figure 2 An example MASA metadata stream merger is schematically shown;

[0071] Figure 3 Schematically shown in more detail as Figure 2 Example MASA stream combiner shown;

[0072] Figure 4 schematically illustrates an example OMASA stream merger for merging an OMASA stream and an ISM stream into a combined stream;

[0073] Figure 5 schematically illustrates an example system of apparatus suitable for implementing some embodiments;

[0074] Figure 6 Schematically illustrates a device suitable for use as a Figure 5 An example combiner of a portion of an encoder is shown, the combiner being configured to implement combining.

[0075] Figure 7 According to some embodiments, Figure 6 A flowchart of the operation of an example metadata merger is shown;

[0076] Figure 8 Schematically illustrates a device suitable for use as a Figure 5Additional example metadata merger of a portion of the illustrated encoder, the merger configured to implement the merger;

[0077] Figure 9 According to some embodiments, Figure 8 The example shown is a flowchart of the operation of an example metadata merger;

[0078] Figure 10 Schematically illustrates a device suitable for use as a Figure 5 Another example metadata merger of a portion of the illustrated encoder, the merger configured to implement the merger;

[0079] Figure 11 Schematically illustrates a device suitable for use as a Figure 5 a multi-directional metadata merger as part of the encoder shown, the merger being configured to implement the merger; and

[0080] Figure 12 There is shown an example apparatus that is suitable for implementing the means shown in the preceding Figures. DETAILED DESCRIPTION

[0081] Suitable devices and possible mechanisms for encoding a parametric spatial audio signal comprising a transport audio signal and spatial metadata are described in more detail below. As mentioned above, immersive audio codecs (such as 3GPP IVAS) are being planned that support multiple operating points from low bitrate operation to transparency. It is expected to support channel-based audio input, object-based audio input, and scene-based audio input, including spatial information about the sound field and sound sources.

[0082] In the following, example codecs are configured to receive multiple input formats. In particular, the codec is configured to acquire or receive multiple audio signals (e.g., received from a microphone array, or as a multi-channel audio format input, or as an Ambisonic format input) and audio object signals (these may also be referred to as independent streams with metadata - ISM format). Furthermore, in some cases, the codec is configured to process more than one input format at a time.

[0083] For example, such a combined (input) format mode may enable simultaneous encoding of two different audio input formats.One example of two different audio input formats currently under consideration is a combination of the MASA format and the Audio Object Format (ISM format).

[0084] Metadata Assisted Spatial Audio (MASA) is an example of a parametric spatial audio format and representation suitable as an input format for IVAS.

[0085] It can be thought of as an audio representation consisting of "N channels + spatial metadata." It is a scene-based audio format particularly well-suited for spatial audio capture on practical devices such as smartphones. The idea is to describe the sound scene based on the direction of sound sources and, for example, energy ratios, which vary in time and frequency. Sound energy that is not defined (described) by direction is described as diffuse (coming from all directions).

[0086] As described above, the spatial metadata associated with the audio signal may include multiple parameters for each time-frequency block (such as multiple directions, and the direct-to-total energy ratio, propagation coherence, distance, etc. associated with each direction (or direction value). The spatial metadata may also include other parameters, or may be associated with other parameters, which are considered non-directional (such as surround coherence, diffuse-to-total energy ratio, residual-to-total energy ratio), but when combined with the directional parameters, these parameters can be used to define the characteristics of the audio scene. For example, a reasonable design choice that can produce high-quality output is to determine that the spatial metadata includes one or more directions for each time-frequency subframe (and the direct-to-total ratio, propagation coherence, distance value, etc. associated with each direction).

[0087] about Figure 1 , Figure 1 An example MASA analyzer 101 is shown. The MASA analyzer 101 is configured to receive input audio signal(s) 100 and analyze the input audio signals to generate transmission audio signal(s) 102 and spatial metadata 104.

[0088] An example of MASA spatial metadata is shown in the table below. These values ​​can be used for each time-frequency block. In some implementations, the frame is subdivided into 24 frequency bands and 4 time subframes. In other implementations, other divisions of frequency and time can be used. In addition, in some implementations, the frame size (e.g., as implemented in IVAS) is 20ms (and therefore, the time subframe is 5ms). However, similarly, other frame lengths can be used in other embodiments. In some embodiments, the MASA analyzer is configured to determine 1 or 2 directions for each time-frequency block (i.e., each time-frequency block has 1 or 2 direction indices, direct to total energy ratios, and propagation coherence parameters). However, in some embodiments, the analyzer is configured to generate more than 2 directions for a time-frequency block.

[0089]

[0090]

[0091]

[0092] The MASA stream can be rendered to various outputs, such as a multi-channel speaker signal (eg, 5.1) or a binaural signal.

[0093] As mentioned above, another input format supported by IVAS is the ISM (Independent Stream with Metadata) format. The ISM format is intended to represent individual sound sources (audio objects) in a sound scene with associated metadata that describes how the audio signal rendering is implemented. This can be contrasted with the MASA format, in which the audio signal and associated metadata perceptually describe the entire sound scene. Example metadata in the ISM format may include the following parameters:

[0094] Position (e.g., azimuth and elevation; or azimuth, elevation, and radius);

[0095] orientation;

[0096] reach / dissemination;

[0097] distance decay; and

[0098] Directional mode.

[0099] In some embodiments, some of the example metadata parameters (such as position, orientation, and range / propagation) may vary with time but not with frequency, while some of the parameters (e.g., range attenuation and directivity pattern) may vary with frequency but generally not with time. However, in addition to the examples described above and herein, parameters may also have both time and frequency variations (or invariance).

[0100] In some embodiments, ISM format input can be obtained by capturing individual sources in a scene using, for example, close-range or lavalier microphones (located on or near the individual sources). Examples of such individual sources are individual speakers in a conference call, singers, or individual instruments. Metadata can then be associated with these ISM format signals, for example, automatically using a position tracker or manually by a mixing professional. Alternatively, an ISM format signal can be generated by mixing and creating appropriate associated metadata. An example of a fully generated sound source is in a game audio engine.

[0101] As mentioned above, ongoing research is underway to enable the IVAS codec to support the combined encoding of multiple audio formats. In many use cases, the content of multiple formats can be captured from the sound scene, and by encoding it into the combined format, it can be more efficiently encoded and sent. An example use case of this situation is a "reporter scene", in which spatial capture is achieved to obtain the overall sound scene, and a separate microphone is added for the reporter to optimally capture his voice, thereby improving the quality of complete capture. In this example, the corresponding capture format can be a MASA format (from capturing the sound scene) and an ISM format (from the reporter). If there is no combined encoding, these formats require independent IVAS codec instances, and using combined encoding, only one instance is required. This saves complexity, and allows the bit rate between the different parts (audio and metadata) of the combined format to be used and very flexibly optimized.

[0102] This combined format may be referred to as the "Object and MASA" format (OMASA), which combines the MASA format and the ISM format into a single combined format.

[0103] GB2217905.5 describes the bit rate-dependent operating modes of OMASA encoding. The following examples feature the lowest bit rate operating modes, in which ISM format input is merged with MASA format input and encoded into a MASA format stream. However, it should be understood that embodiments may use other modes or may be used in other OMASA encoding situations.

[0104] At the lowest bitrates encoded in the OMASA format, a format merging method is needed to merge the formats into a form that is more suitable for encoding and transmission with strict bitrate constraints. An example merging method has been described in GB2574238. This merging method shows two parameterized spatial audio format inputs, such as the MASA format, being merged into a single output parameterized spatial audio format. For example, Figure 2 and Figure 3 shown.

[0105] Figure 2 A (metadata) stream merger 201 is shown, which is configured to receive MASA stream 1 200 and MASA stream 2 202 and generate a (MASA) combined stream 204 . Figure 3 An example (metadata) stream merger 201 is also shown in more detail, as well as information about metadata merging. In this example, the (metadata) stream merger 201 is configured to receive MASA stream 1 200 and MASA stream 2 202 .

[0106] The example (metadata) stream merger 201 includes an energy determiner (reference numeral 311 for stream 1 and reference numeral 313 for stream 2) configured to determine stream 1 energy 300 and stream 2 energy 302. For example, the energy determiner may calculate the energy of the MASA transmission signal in each parameter TF block (b, s) using the following formula for each of the two MASA format streams:

[0107]

[0108] here, is the CLDFB domain representation of the mth MASA stream in the i-th transport channel in the CLDFB interval k and time slot n, grouped into parameter band b and subframe s. Subframe grouping is used because these values ​​are typically calculated every 5ms (and the time resolution of the CLDFB representation is 1.25ms), but in some embodiments they can be determined per frame, in other words, every 20ms. For clarity, the indices (b,s) are omitted in the following text, but the operation is performed for each TF block. This calculation is performed for the two streams to be merged and the two streams are represented by stream 1 energy 300 and stream 2 energy 302.

[0109] The example (metadata) stream merger 201 also includes a ratio determiner (reference numeral 301 for stream 1 and reference numeral 303 for stream 2). The ratio determiner may obtain the ratio from the MASA stream and pass the ratio to the weight determiner.

[0110] Furthermore, the example (metadata) stream merger 201 may comprise a weight determiner (reference numeral 305 for stream 1 and reference numeral 307 for stream 2). The weight determiner is configured to form a comparison weight for the two MASA format streams as the direct signal energy of the MASA, i.e. in is the direct to total energy ratio of the mth MASA format stream metadata. This generates two weights w1 304 and w2 306 for each TF block. These weights are then passed to the weight comparator 309.

[0111] The example (metadata) stream merger 201 further comprises a weight comparator 309 configured to compare the weight of each TF block and control the metadata selector 321 .

[0112] Furthermore, the example (metadata) stream merger 201 comprises a metadata selector 321 configured to receive metadata from the MASA stream and to receive control from the weight comparator 309. The metadata selector 321 may be configured such that:

[0113] If w1>w2 of a TF block, metadata is selected from MASA stream 1 for the TF block.

[0114] Otherwise, metadata is selected for the TF block from MASA stream 2.

[0115] The selected metadata can then be output as merged metadata within the (MASA) combined stream 204. In some cases, the comparison "greater than" > is a comparison "greater than or equal to" ≥. This option can be applied in the following examples. In some cases, the comparison can include a bias factor (which can be an additive and / or multiplicative factor) that is configured to bias the decision in a certain direction.

[0116] This can also be extended to other input formats besides the MASA input format. Figure 4 An example is shown in , where an ISM stream and a MASA stream are combined and encoded to generate a bitstream.

[0117] So, for example, Figure 4 A merger / encoder is shown comprising an audio stream merger 401 configured to generate a combined audio stream 412 from a stream 1 (MASA) input audio signal 400 and a stream 2 (ISM) input audio signal 402 .

[0118] The merger / encoder further comprises an audio stream energy determiner 403 configured to determine the audio signal energy 408, for example in a manner similar to that described above with respect to the example Figure 3 The energy determiners 311, 313 are shown in the manner described.

[0119] The merger / encoder may include an ISM to MASA converter 405 configured to receive an (ISM) stream 2 406 and generate a MASA stream 2 202. Since the metadata of the ISM stream contains values ​​compatible with the values ​​of the MASA stream, the ISM to MASA converter 405 is configured to generate or construct a MASA stream based on the ISM stream. In the case of a single (active) object, for example, a MASA stream can be generated by converting the position of the ISM stream into the direction of the MASA stream and making the constructed MASA stream fully directional. In other words, a direct to total ratio of 1 and a diffuse to total ratio of 0 are assigned to the constructed MASA format. The definition of an ISM is a fully directional sound source, and a direct to total ratio of 1 and a diffuse to total ratio of 0 are assigned to it, which can be demonstrated by the intended application of acquiring the ISM as a single audio object (e.g., a lavalier microphone).

[0120] In some cases, particularly when there are multiple simultaneous objects in the ISM stream, the converter 405 is configured to convert the ISM to a first-order ambient acoustic (FOA) representation and analyze the source directions from the representation. The direct to total energy ratio can be obtained by estimating the diffuse to total energy ratio of the FOA signal and determining the direct to total energy ratio therefrom. Any other suitable method can be used to convert the ISM stream to a MASA stream (two specific examples are provided above).

[0121] The two MASA format stream metadata can then be merged.

[0122] For example, the merger / encoder may include a metadata stream merger 407 configured to receive (MASA) stream 1 200 and MASA stream 2 202 and to combine these inputs in a manner similar to that described with respect to Figure 3 The combined MASA stream 410 generated by the metadata stream merger 407 may then be output to the spatial metadata encoder 409 .

[0123] The merger / encoder may then include a spatial metadata encoder 409 configured to encode the combined MASA stream 410 based on a suitable MASA metadata encoding method and output the encoded metadata to a bitstream creator 413 .

[0124] The merger / encoder may also include a bitstream creator 413 that is configured to receive encoded audio 414 from the audio encoder 411 and encoded metadata 416 from the spatial metadata encoder 409 and generate a bitstream 418 to be output for storage and / or transmission.

[0125] about Figure 5 , shows an example system of apparatus suitable for implementing some embodiments. The system of apparatus is configured to receive (one or more) stream 1 input audio signals 500, stream 1 spatial metadata (format 1) 504 and (one or more) stream 2 input audio signals 502, stream 2 spatial metadata (format 2) 506. These are received by a merger / encoder 501, which is configured to merge the inputs and generate a bitstream 510. The bitstream is then passed to a decoder 503, which decodes the bitstream and generates (one or more) output audio signals 512.

[0126] In general, the concept is to improve upon the above-described examples of merging MASA streams and MASA streams, as well as ISM to MASA stream conversion. Specifically, when the time-frequency resolution and allocated bitrate of the combined MASA format being encoded and transmitted are limited, such as 24.4 kbps, the above-described methods may suffer from a loss in quality. While the loss in quality is more pronounced at the lowest bitrates, where the time-frequency resolution is lower, it can also be present at higher bitrates, so the goal is to improve quality across a range of bitrates.

[0127] The reason for the loss in quality is that the direct-to-total ratio of the ISM stream is close to 1, coupled with the typically large amount of signal energy, emphasizes the presence of the ISM stream in the resulting merged and encoded MASA format. This can be perceived as a disconcerting instability in the reproduced sound scene, as the surround diffuse sound present in the MASA stream almost completely disappears when the ISM stream is active. Furthermore, in some cases, due to the high direct-to-total ratio of the ISM stream, coding artifacts (e.g., heavily quantized directions) can be perceived, since the TF blocks in the merged stream containing metadata from the ISM stream will be allocated a higher bit budget than other TF blocks in the metadata encoding. This can also lead to a decrease in the perceived quality of the merged stream.

[0128] The concepts discussed below with respect to some embodiments and examples are apparatus and methods that aim to preserve the overall diffuse characteristics of a MASA format source stream in a merged MASA format stream while enabling merging of MASA streams and ISM conversion to a MASA stream.

[0129] Thus, the embodiments and examples shown herein relate to merging two parametric spatial audio (ie, audio signal(s) associated with spatial metadata) streams, where the first stream is a MASA format stream and the second stream is an ISM stream.

[0130] The apparatus and method propose a merging of parameters that adjust the two spatial metadata streams so as to preserve the MASA format scene envelope by preserving its diffuse energy, thereby improving the perceived overall quality of the merged stream.

[0131] This can be achieved in the following embodiments and examples by:

[0132] Get the direct-to-total ratio of two streams;

[0133] Get the signal energy of the two streams and the diffuse energy of the MASA stream;

[0134] Determine the direct to total ratio parameter to compensate for diffuse energy:

[0135] Determine the weights of the two streams (where the weights may be based on the direct-to-total ratio of the streams and the signal energy estimate);

[0136] Compare the weights of the two streams in each TF block; and

[0137] Based on the comparison, metadata for each TF block of the merged MASA stream is selected. The selection can be "MASA stream" parameters for the TF block in the MASA stream, where the MASA stream related weight is greater than the ISM stream related weight; compensated diffuse energy ISM stream parameters, where the ISM stream related weight is greater than the MASA stream related weight; or selecting MASA stream and compensated diffuse energy ISM stream parameters for the TF block based on the value of the ISM stream related weight being greater than at least one of more than one MASA stream related weights.

[0138] In the following disclosure, it is understood that the terms weight(s) and weighted value are interchangeable.

[0139] Thus, the apparatus and method are intended to provide improved merging. Although the examples given here are of two parametric spatial audio streams when one of the streams is originally an ISM stream, i.e., one or more independent audio objects, and the other stream is a MASA format stream (or any similar parametric format), in some embodiments, other stream formats can be merged based on the examples presented herein without requiring significant creative adjustments. If the metadata of an input format data stream shows similar properties to an ISM stream, in other words, high directionality and low diffuseness, then this input format stream is particularly suitable for merging using the examples described herein.

[0140] In some embodiments, the apparatus and methods described herein may be implemented within a communication codec encoder such as 3GPP IVAS, and the combined format encoding may be applied at low bit rates where full merging of separate format streams is required to achieve high quality transmission.

[0141] Furthermore, the embodiments described herein may be implemented as part of a combiner / encoder 501, such as Figure 5 Thus, in some embodiments, the apparatus is configured to receive two parametric spatial audio streams ("stream 1" and "stream 2") as input. Each of the input streams contains (one or more) audio signals + metadata ("stream 1 input audio signals" 500, "stream 1 spatial metadata (format 1) 504", "stream 2 input audio signals" 502, "stream 2 spatial metadata (format 2)" 506).

[0142] In this example, the streams are MASA format streams and ISM streams. In some embodiments, the merger / encoder may include a preprocessor that converts the captured audio signal into a MASA format stream and an ISM stream. The input stream is provided to the encoder, which quantizes and encodes the format into a single bit stream. The bit stream is sent to a decoder, which in turn decodes the bit stream and generates (one or more) output audio signals, or in some embodiments, generates an output format (e.g., MASA format).

[0143] Figure 6 The example combiner shown is suitable for Figure 5 The metadata merging operation is implemented in the combiner / encoder of the embodiment. In these embodiments, the merging or combining of the audio signals and metadata is achieved using different methods. The combining or combining of the audio signals can be achieved using any suitable method. For example, by mixing the input audio signals together to provide a combined audio signal, which is then encoded. Therefore, the merging or combining of the audio signals will not be described in further detail.

[0144] Figure 6 A (metadata) merger or combiner according to some embodiments is shown in .

[0145] The merger is configured to receive MASA stream 1 metadata 600 , MASA stream 1 transport signal(s) 602 , and ISM stream metadata / object audio 610 .

[0146] In some embodiments, the merger includes an ISM to MASA converter 607. The ISM to MASA converter 607 is configured to receive the ISM stream metadata / object audio 610 and generate MASA stream 2 metadata 614 and (one or more) ISM (MASA) stream transmission signals 612. This can be achieved in any suitable manner, for example, by assigning the values ​​of the directional parameters of the ISM stream to the directional parameters of the MASA format metadata for all bands and making the direct to total ratio and diffuse to total ratio of the MASA format metadata always 1 and 0 (or converting the audio signal to a FOA representation and estimating the parameters therefrom, as described above). The MASA stream 2 metadata 614 is passed to the direct to total ratio obtainer 609. The (one or more) ISM (MASA) stream transmission signals 612 are passed to the transmission signal energy determiner 611 and the audio stream merger 670.

[0147] The combiner includes a direct and total ratio obtainer 601 configured to obtain the MASA stream 1 metadata 600 and extract or otherwise obtain the MASA ratio 604 , the ratio is output to a weight determiner 605 .

[0148] The metadata merger includes a direct and total ratio obtainer 609 configured to obtain the MASA stream 2 metadata 614 and extract or otherwise obtain the ISM based ratio. 616 , the ratio is output to the weight determiner 613 .

[0149] Furthermore, the metadata merger comprises a transmission signal energy determiner 603 configured to estimate the energy of the MASA transmission signal in each parameter TF block using the following formula:

[0150]

[0151] Here, X i (k,n) is the CLDFB domain representation of the i-th transmission channel in CLDFB interval k and time slot n, which are grouped into parameter bands b and frames s. Hereafter, the index (b,s) is omitted for clarity, but the operation is performed for each TF block. Output e MASA 606 is output to the weight determiner 605 , the diffuse signal energy harvester 653 and the combined diffuse to total ratio determiner 655 .

[0152] The combiner comprises a further transmission signal energy determiner 611 configured to calculate a transmission signal energy estimate for the ISM transmission signal 612 in each parameter TF block using the following formula:

[0153]

[0154] Here, Y i (k,n) is the CLDFB domain representation of the ith channel of the ISM transmission signal. (For clarity, the indices (b,s) are omitted below, but the operation is performed for each TF block.) The energy e of the ISM transmission signal ISM 618 is output to the weight determiner 613 and the combined diffuse to total ratio determiner 655 .

[0155] The merger further comprises a diffuse to total ratio obtainer 651 configured to obtain or extract the diffuse to total ratio from the MASA stream 1 metadata 600. 654.

[0156] The combiner further includes a diffuse signal energy acquirer 653, which is configured to acquire the diffuse and total ratio. 654 and the energy e of the MASA transmitted signal in each parameter TF block MASA 606, and generate a diffuse energy value 656. In some embodiments, the diffuse signal energy is determined based on:

[0157] The combiner comprises a combined diffuse to total ratio determiner 655 configured to calculate a combined diffuse to total ratio of the combined MASA stream 1 and the ISM stream using the following formula:

[0158]

[0159] This value describes the amount of diffuse energy in the signal after merging the two streams. Diffuse Signal Energy 658 is passed to the combined direct to total ratio determiner 657.

[0160] The data merger also includes a merged direct to total ratio determiner 657. This can be done by 660 is implemented, and this value may be used as a modified direct to total energy ratio parameter for the block in the MASA stream 2 in which the metadata is selected.

[0161] The metadata merger comprises a weight determiner 605 configured to determine a comparison weight for the MASA formatted stream using the direct to total energy ratio and the transmitted signal energy, as the direct signal energy of the MASA, which may be 608, of which r MASA is the direct to total energy ratio of the MASA format stream metadata.

[0162] The metadata merger comprises a further weight determiner 613 configured to determine a comparison weight for the ISM stream using the direct to total energy ratio and the transmission signal energy originating from the ISM stream. The weight may similarly be

[0163] Therefore, the two weights w1 and w2 of each TF block are passed to the weight comparator 615.

[0164] The metadata merger includes a weight comparator 615 configured to compare weights and generate a selection control to a metadata selector 617 .

[0165] The metadata merger also includes a metadata selector 617 configured to select and output selected metadata 622 based on a selection control from the weight comparator.

[0166] For example, a selector can be configured to implement the following selection operations:

[0167] If w1>w2 of the TF block, metadata is selected for the TF block from MASA stream 1 600;

[0168] Otherwise, metadata is selected for this TF block from MASA stream 2 614 (in other words, metadata from the original "ISM stream") and the new direct-to-total ratio is used. 660. In other words, the new direct-to-total ratio determined from the combined direct-to-total ratio determiner 657 is selected. 660 (and effectively replaces the original direct-to-total ratio of the ISM flow 616 possible options).

[0169] Furthermore, in some embodiments, the merger includes an audio stream merger 670 that is configured to receive the MASA stream 2 transport / audio signal 664 from the ISM to MASA converter 607 and to receive (one or more) MASA stream 1 transport signals 602 and to generate therefrom a combined stream transport / audio signal 674 that can be output to a suitable audio encoder.

[0170] The combined metadata 622 can then be passed to a suitable spatial metadata encoder, and the result can be multiplexed with the audio signal, for example, within a suitable bitstream creator to provide an encoded bitstream that can be sent to a decoder (or stored for later decoding). The operation of the decoder is not discussed in further detail because the decoding of the bitstream and the decoder itself can be implemented using any suitable decoder. In other words, no changes are needed to support the combined MASA stream. Therefore, in summary, the decoder demultiplexes and decodes the bitstream into audio channels and metadata, which are then output in MASA format or passed to a renderer, which renders it into various output audio signals. These signals can be binaural, speaker, ambisonic, or any other format.

[0171] about Figure 7 , showing Figure 6 An example flow chart of the operations shown.

[0172] Thus, as shown at 700, obtaining MASA stream 1 metadata is shown.

[0173] 702 further illustrates obtaining the MASA stream 1 transmission signal(s).

[0174] Additionally, 721 illustrates obtaining ISM stream metadata / object audio.

[0175] 703 shows the direct to total ratio obtained from MASA stream 1 metadata.

[0176] 705 shows determining the transmit signal energy based on the MASA stream 1 transmit signal(s).

[0177] 709 shows the diffuse to total ratio obtained from the MASA stream 1 metadata.

[0178] 711 shows the diffuse signal energy being derived from the diffuse to total ratio and the transmitted signal energy.

[0179] 707 shows that the weight of stream 1 is determined according to the direct-to-total ratio and the transmission signal energy.

[0180] 723 shows the conversion of ISM to MASA metadata and obtaining the ISM stream (or MASA stream 2) audio.

[0181] 725 shows obtaining the direct to total ratio from the converted MASA metadata.

[0182] 719 shows determining the transmission signal energy based on the ISM stream transmission signal(s).

[0183] 727 shows the determination of stream 2 weight based on the direct to total ratio and the transmitted signal energy (for the original ISM stream).

[0184] 713 shows determining the combined diffuse to total ratio.

[0185] Next, 715 illustrates determining a combined direct-to-total ratio based on the combined-to-total ratio.

[0186] 729 shows a comparison of weights.

[0187] Then, 731 shows selecting metadata based on the comparison.

[0188] 735 shows the output of the selected metadata.

[0189] In addition, 741 also shows the merging of audio streams.

[0190] Then, 743 shows outputting the combined or merged audio stream.

[0191] This implementation aims to better take into account the overall diffuse energy of the MASA format, which is very important for conveying the perceptual envelope of the sound scene. This aims to improve upon known merging methods, which, when metadata from an ISM format stream is selected for a TF block, typically also completely eliminate (or severely reduce) the diffuse energy of that block. Therefore, an embodiment adjusts the amount of diffuse energy by taking into account the increased total energy from the ISM stream. Since it is assumed that the ISM stream energy is mainly direct, it will reduce the diffuse energy of the block, but only in proportion to the relative stream energy. This aims to significantly improve the quality of the merged stream, as the diffuse field perception is better preserved.

[0192] While these example embodiments are intended to provide significant quality improvements, in some alternative embodiments, the decision criteria and / or determined ratios for the merged stream metadata can be further adjusted to attempt to further improve quality. These other embodiments are discussed further below. It should also be noted that these improvements can also be used together. For example, in some embodiments, the following two example improvements can be applied.

[0193] In some embodiments, as Figure 8 As shown, the comparison weight w2 of flow 2 (MASA flow formed by ISM flow) is determined based on the combined direct and total energy ratio as follows: The difference from the previously described embodiment is that slightly different weights are used for the flow decisions, but when the decisions are the same, the resulting metadata will also be the same.

[0194] therefore, Figure 8 The example embodiment shown is similar to Figure 6 The embodiment shown differs in that the weight determiner 813 is based on 820 generates a comparison weight for flow 2, which is passed to the weight comparator 615.

[0195] In some embodiments, the weight determiner 813 is configured to determine two direct to total ratios r ISM 、 This is the average value of Figure 8 As shown by the dashed elements of the MASA stream 2 metadata 614, the MASA stream 2 metadata 614 is processed by the direct to total ratio determiner 609, which converts the ISM direct to total ratio r ISM is output to the weight determiner 813. Therefore, the weight 830 based on 'average' can be determined as:

[0196]

[0197] Figure 9 A flow chart is shown which shows Figure 8 Example device operation, which is related to Figure 7 The difference is that the operation of determining the weight of stream 2 differs in the elements used to generate the weight, as shown in 927.

[0198] and Figure 6 and Figure 8The examples shown differ in that (slightly) different weights are used for the stream decisions. However, when the decisions are the same, the resulting selected metadata will also be the same. In practice, using these adjusted decision criteria can maintain sound scene stability and further preserve the spatial experience in the scene, as the MASA stream is selected more often, especially in cases where the ISM stream contains a ratio close to 1.

[0199] In some embodiments, the conversion from ISM to MASA representation can be achieved by assuming that the signal energy is completely directional to assign a direct to total energy ratio of 1. While this is correct in some cases, other methods of conversion can be used, such as using FOA representation. In such embodiments, the ISM may contain a certain amount of diffuse energy. Another example case where there may be a non-negligible diffuse to total energy ratio in the ISM stream is when two or more ISM objects are active in the same TF block, resulting in a direct to total energy ratio that can be much less than 1. Such parameterization may be desirable for a more realistic perception of space in the reproduction.

[0200] If the amount of diffuse energy in the ISM is not negligible, then it may be useful to take this diffuse energy into account. Otherwise, for TF blocks where the use of ISM metadata is chosen, but the diffuse energy of the MASA stream is less than that of the ISM stream, the ratio may be unintentionally increased. This may be perceived as a loss of spatial sense in the merged stream.

[0201] The interpretation of diffuse energy can take many forms. For example, if ISM metadata is selected for the TF block, the direct to total energy ratio is obtained. can be constrained to be no greater than the original direct to total energy ratio, i.e.:

[0202]

[0203] Another example, based on Figure 6 The example shown, but considering diffuse energy, such as Figure 10 As shown, the combined diffuse and total ratio determiner 1005 is modified by the following formula:

[0204]

[0205] in describes the amount of diffuse energy in the ISM stream. Then, 1008 is passed to the combined direct to total energy ratio determiner 707 to determine the direct to total energy ratio of the combined stream. Then, the modified direct energy is compared with the total energy 1010 is passed to the metadata selector 709.

[0206] In addition, Figure 10 a direct-to-total / diffuse-to-total ratio acquirer 1009 is shown, which is configured to acquire not only the direct-to-total ratio but also the diffuse-to-total ratio 1016, and this ratio is passed to the combined diffuse-to-total ratio determiner 1005.

[0207] Although this example is based on Figure 6 the example shown, for Figure 8 the example embodiments shown, similar modifications can be made to the combined diffuse-to-total ratio determination.

[0208] Figure 11 Another example is shown, where there are two simultaneous directions. This can be further extended to more directional fields.

[0209] Thus, the direct-to-total ratio acquirer 1101 is configured to acquire two direct-to-total energy ratios in the original MASA metadata stream and one for each direction, which are also designated as 1104.

[0210] The weight determiner 1105 is configured to determine the absolute direct energy for the two directions calculated using the following formula: and These are used as additional weight terms instead of the previous w1: and

[0211] The weights are compared in the weight comparator 715 and then selected by the metadata selector based on the following:

[0212] If w 1,1 > w2 and w 1,2 > w2, then the "MASA stream 1" parameters are used;

[0213] If w 1,1 ≤ w 1,2 < w or w 1,1 < w2 ≤ w 1,2 , then the parameters from "MASA direction 2" and "ISM stream" are used.

[0214] If w <00:00015>≤ w 1,1 < w2 or w 1,2 < w2 ≤ w 1,1 , then the parameters from "MASA direction 1" and "ISM stream" are used.

[0215] In some embodiments, this can be accomplished by configuring the diffuse signal energy harvester 1103 to determine the diffuse signal energy of the MASA using the following equation: 1106 to expand:

[0216]

[0217] Furthermore, the combined diffuse to total energy ratio determiner 1111 is configured to determine the combined signal using the following formula The diffuse to total energy ratio of 1113:

[0218]

[0219] Furthermore, the combined direct to total energy ratio determiner 1107 is then configured to scale the newly assigned direct to total energy ratio according to the new diffuse to total energy ratio using the following equation:

[0220]

[0221] as well as

[0222]

[0223] Similar to the previous example, this bidirectional embodiment can also be implemented in a similar way to Figure 8 and Figure 10 The approach of the embodiments described in

[0045] is extended to include the diffuse energy of ISM objects.

[0224] In some embodiments, the merging or combining of metadata is selected based on a metric of "object similarity" determined from the spatial metadata of the streams to be merged. In other words, the weights are determined based on the object similarity parameter.

[0225] In such embodiments where both streams have low object similarity, e.g., they contain a non-negligible amount of diffuse energy, then implementations such as Figure 6 The merging method shown.

[0226] If both streams have high object similarity, for example, the amount of diffuse energy of both streams is negligible, then the implementation is as follows Figure 6 The merging method shown.

[0227] If one stream has low object similarity and one stream has high object similarity, then one can achieve Figure 9 Any one of the methods shown in, or as combined in the embodiments described above.

[0228] In some embodiments, an "object similarity" parameter may be determined by examining the overall diffuse energy (objects are typically highly directional with minimal or no diffuse component), and examining directional variations over frequency, since objects typically have no frequency variations.

[0229] In some embodiments, where metadata from "MASA stream 1" is selected, extensions of the above example embodiments may also be made by using a modified direct to total energy ratio In practice, the direct-to-total ratio will be the same whatever is chosen, while the other parameters will be chosen from one stream or the other.

[0230] The entire description above describes the case of merging two streams. In the case of merging three or more streams, the method can be extended in a straightforward manner, for example, by repeatedly merging streams in pairs (or cascades) until the desired number of streams remains, or by extending the above method to include weights for all streams simultaneously and selecting the metadata for the stream with the highest weight.

[0231] In the above example, the complex-valued low-delay filter bank (CLDFB) is used as the frequency domain representation, and other time-frequency domain representation methods such as short-time Fourier transform (STFT) or quadrature mirror filter bank (QMF) can be used.

[0232] The parameter format described above is the MASA format, but embodiments can be extended to other parameter formats, such as Ambisonics or parameter encoding of multi-channel mixes.

[0233] about Figure 12 , the example electronic device can be used as any device component in the device components of the above-mentioned system. The device can be any suitable electronic device or device. For example, in some embodiments, the device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback device, etc. For example, the device can be configured to implement the encoder and / or decoder or any functional blocks described above.

[0234] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 may be configured to execute various program codes, such as the methods described herein.

[0235] In some embodiments, device 1400 includes at least one memory 1411. In some embodiments, at least one processor 1407 is coupled to memory 1411. Memory 1411 can be any suitable storage device. In some embodiments, memory 1411 includes a program code portion for storing program code that can be implemented on processor 1407. In addition, in some embodiments, memory 1411 can also include a data storage portion for storing data, such as data that has been processed or is to be processed according to the embodiments described herein. The implemented program code stored in the program code portion and the data stored in the data storage portion can be retrieved by processor 1407 via the memory-processor coupling when needed.

[0236] In some embodiments, device 1400 includes a user interface 1405. In some embodiments, user interface 1405 can be coupled to processor 1407. In some embodiments, processor 1407 can control the operation of user interface 1405 and receive input from user interface 1405. In some embodiments, user interface 1405 can enable a user to enter commands to device 1400, for example, via a keypad. In some embodiments, user interface 1405 can enable a user to obtain information from device 1400. For example, user interface 1405 can include a display configured to display information from device 1400 to the user. In some embodiments, user interface 1405 can include a touch screen or touch interface that can enable information to be entered into device 1400 and further display information to a user of device 1400. In some embodiments, user interface 1405 can be a user interface for communication.

[0237] In some embodiments, device 1400 includes input / output port 1409. In some embodiments, input / output port 1409 includes a transceiver. In such embodiments, the transceiver can be coupled to processor 1407 and configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver components can be configured to communicate with other electronic devices or devices via a wired or wired coupling.

[0238] The transceiver can communicate with other devices through any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable radio access architecture based on: Advanced Long Term Evolution (LTE-Advanced, LTE-A) or New Radio (NR) (or may be referred to as 5G), Universal Mobile Telecommunications System (UMTS) Radio Access Network (UTRAN or E-UTRAN), Long Term Evolution (LTE, the same as E-UTRA), 2G network (legacy network technology), Wireless Local Area Network (WLAN or Wi-Fi), Worldwide Interoperability for Microwave Access (WiMAX), Personal Communications Service (PCS), Wideband Code Division Multiple Access (WCDMA), systems using Ultra-Wideband (UWB) technology, sensor networks, Mobile Ad Hoc Networks (MANET), Cellular Internet of Things (IoT) RAN and Internet Protocol Multimedia Subsystem (IMS), any other suitable options and / or any combination thereof.

[0239] The transceiver input / output port 1409 may be configured to receive signals.

[0240] In some embodiments, the apparatus 1400 may be used as at least part of a composite device.The input / output port 1409 may be coupled to headphones (which may be headphones or non-headphones) or the like and speakers.

[0241] In general, various embodiments of the present invention may be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software, which may be executed by a controller, microprocessor, or other computing device, although the invention is not limited thereto. Although various aspects of the present invention may be illustrated and described using block diagrams, flow charts, or some other graphical representation, it is well understood that, as non-limiting examples, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.

[0242] Embodiments of the present invention can be implemented by computer software that can be executed by the data processor of mobile device, such as in processor entity, or by hardware, or by the combination of software and hardware.In addition, in this respect, it should be noted that any block of the logic flow shown in the figure can represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, boxes and functions.Software can be stored on physical media such as memory chips or memory blocks implemented in processors, on magnetic media such as hard disks or floppy disks, and on optical media such as DVDs and data variants thereof, CDs.

[0243] The memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technical environment and may include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture, as non-limiting examples.

[0244] Embodiments of the present invention may be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0245] Programs such as those offered by Synopsys, Inc. of Mountain View, Calif., and Cadence Design, Inc. of San Jose, Calif., use well-established design rules and a library of pre-stored design modules to automatically route conductors and position components on a semiconductor chip. Once the design of a semiconductor circuit is complete, the final design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be sent to a semiconductor fabrication facility, or "fab," for manufacturing.

[0246] As used in this application, the term "circuitry" may refer to one or more or all of the following:

[0247] (a) hardware circuit implementation only (such as implementation in analog and / or digital circuitry only) and

[0248] (b) a combination of hardware circuitry and software such as (as applicable):

[0249] (i) a combination of (one or more) analog and / or digital hardware circuits and software / firmware, and

[0250] (ii) any portion of hardware processor(s) (including digital signal processor(s)) with software, software and memory(s) that work together to cause a device such as a mobile phone or server to perform various functions), and

[0251] Hardware circuit(s) and / or processor(s), such as microprocessor(s) or portion(s) of microprocessor(s), that require software (e.g., firmware) to operate, but where the software is not required to operate, the software may not be present.

[0252] This definition of circuitry applies to all uses of the term in this application, including in any claims. As another example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or a portion of a hardware circuit or processor and its accompanying software and / or firmware. For example, the term circuitry also covers a baseband integrated circuit or processor integrated circuit for a mobile device, or a similar integrated circuit in a server, cellular network device, or other computing or networking device, if applicable to the particular claim element.

[0253] The term "non-transitory" as used herein is a restriction to the medium itself (ie, tangible, not a signal), not to the persistence of the data storage (eg, RAM vs. ROM).

[0254] As used herein, “at least one of: ” and “<at least one of a list of two or more elements>” and similar expressions (where a list of two or more elements is connected by “and” or “or”) refer to at least any one of these elements, or at least any two or more of these elements, or at least all of these elements.

[0255] The foregoing description has provided by way of exemplary and non-limiting examples a complete and informative description of exemplary embodiments of the present invention. However, various modifications and adaptations will be apparent to those skilled in the relevant arts in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. Nevertheless, all such and similar modifications of the teachings of this invention will still fall within the scope of the invention as defined in the appended claims.

Claims

1. An apparatus comprising components for: For the first audio stream, obtaining at least one first direct-to-total ratio parameter; For the first audio stream, obtaining a first signal energy parameter; generating at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; For the second audio stream, obtaining a second direct-to-total ratio parameter; For the second audio stream, obtaining a second signal energy parameter; Generate direct-to-total ratio parameters that compensate for diffuse energy; The second weighted value is generated based in part on at least one of: the second direct-to-total ratio parameter; Direct to total ratio parameter of the compensated diffuse energy; and the second signal energy parameter; as well as Based on a comparison of the at least one first weighting value and the second weighting value, one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating diffuse energy are selected.

2. The apparatus of claim 1 , wherein the means for selecting one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating diffuse energy based on a comparison of the at least one first weighting value and the second weighting value is further configured to: When the at least one first weighted value is greater than the second weighted value, selecting one of the at least one first direct-to-total ratio parameters; and When the second weighted value is greater than the at least one first weighted value, the direct to total ratio parameter of the compensated diffuse energy is selected.

3. The apparatus of claim 2 , wherein the means for selecting one of the at least one first direct-to-total ratio parameters when the at least one first weighted value is greater than the second weighted value is when the at least one first weighted value is strictly greater than the second weighted value.

4. The apparatus of claim 2 , wherein the means for selecting the direct-to-total ratio parameter for compensating diffuse energy when the second weighting value is greater than the at least one first weighting value is when the second weighting value is strictly greater than the at least one first weighting value.

5. An apparatus according to any one of claims 1 to 4, wherein the component for generating the at least one first weighted value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter is used to: generate the at least one first weighted value based on the product of the at least one first direct-to-total ratio parameter and the first signal energy parameter.

6. An apparatus according to any one of claims 1 to 5, wherein the at least one first direct-to-total ratio includes at least two first direct-to-total ratios, and wherein the component for generating the at least one first weighted value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter is used to: generate a first weighted value associated with each of the at least two first direct-to-total ratio parameters.

7. An apparatus according to claim 6, wherein the component for generating a first weighting value associated with each of the at least two first direct-to-total ratio parameters is used to: generate each first weighting value based on the product of each of the at least one first direct-to-total ratio parameter and the first signal energy parameter.

8. The apparatus according to claim 6 , wherein the means for selecting one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating diffuse energy based on a comparison of the at least one first weighting value and the second weighting value is further configured to: When both of the first weighted values ​​are greater than the second weighted value, selecting the at least two first direct-to-total ratio parameters; and When the second weighted value is greater than one of the at least two first weighted values, the direct-to-total ratio parameter for compensating diffuse energy and one of the at least two first direct-to-total ratio parameters are selected.

9. The apparatus according to any one of claims 1 to 8, wherein the means for generating a second weighting value based in part on at least one of the second direct-to-total ratio parameter, the direct-to-total ratio parameter for compensating for diffuse energy, and the second signal energy parameter is for one of the following: generating the at least one second weighted value based on a product of the at least one second direct-to-total ratio parameter and the second signal energy parameter; generating the at least one second weighted value based on a product of a direct-to-total ratio parameter of the compensated diffuse energy and the second signal energy parameter; as well as The at least one second weighting value is generated based on a product of an average value of the direct-to-total ratio parameter of the compensated diffuse energy and the at least one second direct-to-total ratio parameter and the second signal energy parameter.

10. The apparatus of any one of claims 1 to 9, wherein the means for generating a direct-to-total ratio parameter to compensate for diffuse energy is configured to generate at least one combined diffuse-to-total ratio based at least in part on the first signal energy parameter, the second signal energy parameter, and a diffuse signal energy parameter.

11. The apparatus according to claim 10, wherein the component is further configured to: For the first audio stream, obtaining at least one first diffuse-to-total ratio parameter; and The diffuse signal energy parameter is obtained based in part on the first signal energy parameter and the first diffuse-to-total ratio parameter.

12. The apparatus according to any one of claims 10 or 11, wherein the means for generating a direct to total ratio parameter for compensating diffuse energy is further configured to generate the at least one direct to total ratio parameter for compensating diffuse energy based on the at least one combined diffuse to total ratio.

13. The apparatus according to claim 10 , wherein the means is further configured to obtain, for the second audio stream, at least one second diffuse-to-total ratio parameter, and wherein the means for generating at least one combined diffuse-to-total ratio is further configured to generate the at least one combined diffuse-to-total ratio based on the second diffuse-to-total ratio parameter.

14. The apparatus of claim 10, wherein the means for generating the direct to total ratio parameter for compensating diffuse energy is further configured to: generating at least one direct-to-total ratio parameter of an expected compensated diffuse energy based on the at least one combined diffuse-to-total ratio; comparing the direct-to-total ratio parameter of the at least one pre-compensation period diffuse energy with the second direct-to-total ratio parameter; Based on the comparison, the direct-to-total ratio parameter of the compensated diffuse energy is selected from the direct-to-total ratio parameter of the at least one expected compensated diffuse energy and the second direct-to-total ratio parameter, so that the direct-to-total ratio parameter of the compensated diffuse energy is the smaller of the direct-to-total ratio parameter of the at least one expected compensated diffuse energy and the second direct-to-total ratio parameter.

15. The apparatus according to any one of claims 1 to 14, wherein the component is further configured to encode a selected first direct-to-total ratio parameter of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating diffuse energy.

16. The apparatus according to any one of claims 1 to 15, wherein the component is further configured to merge the first audio stream and the second audio stream.

17. A method comprising: For the first audio stream, obtaining at least one first direct-to-total ratio parameter; For the first audio stream, obtaining a first signal energy parameter; generating at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter; For the second audio stream, obtaining a second direct-to-total ratio parameter; For the second audio stream, obtaining a second signal energy parameter; Generate direct-to-total ratio parameters that compensate for diffuse energy; The second weighted value is generated based in part on at least one of: the second direct-to-total ratio parameter; Direct to total ratio parameter of the compensated diffuse energy; and the second signal energy parameter; as well as Based on a comparison of the at least one first weighting value and the second weighting value, one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating diffuse energy are selected.

18. The method of claim 17 , wherein selecting one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating diffuse energy based on a comparison of the at least one first weighting value and the second weighting value further comprises: When the at least one first weighted value is greater than the second weighted value, selecting one of the at least one first direct-to-total ratio parameters; as well as When the second weighted value is greater than the at least one first weighted value, the direct to total ratio parameter of the compensated diffuse energy is selected.

19. The method of claim 18, wherein when the at least one first weighted value is greater than the second weighted value, selecting one of the at least one first direct-to-total ratio parameters is when the at least one first weighted value is strictly greater than the second weighted value.

20. The method of claim 18, wherein when the second weighting value is greater than the at least one first weighting value, selecting the direct to total ratio parameter for compensating diffuse energy is when the second weighting value is strictly greater than the at least one first weighting value.

21. The method according to any one of claims 17 to 20, wherein generating the at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter further comprises: The at least one first weighting value is generated based on a product of the at least one first direct-to-total ratio parameter and the first signal energy parameter.

22. The method according to any one of claims 17 to 21, wherein the at least one first direct-to-total ratio comprises at least two first direct-to-total ratios, and generating the at least one first weighting value based on the at least one first direct-to-total ratio parameter and the first signal energy parameter further comprises: A first weighting value associated with each of the at least two first direct-to-total ratio parameters is generated.

23. The method of claim 22, wherein generating a first weighted value associated with each of the at least two first direct-to-total ratio parameters further comprises: Each first weighting value is generated based on a product of each first direct-to-total ratio parameter of the at least one first direct-to-total ratio parameter and the first signal energy parameter.

24. The method according to any one of claims 22 or 23, wherein selecting one of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating diffuse energy based on the comparison of the at least one first weighting value and the second weighting value further comprises: When both of the first weighted values ​​are greater than the second weighted value, selecting the at least two first direct-to-total ratio parameters; as well as When the second weighted value is greater than one of the at least two first weighted values, the direct-to-total ratio parameter for compensating diffuse energy and one of the at least two first direct-to-total ratio parameters are selected.

25. The method of any one of claims 17 to 24, wherein generating a second weighting value based in part on at least one of the second direct-to-total ratio parameter, the direct-to-total ratio parameter for compensating diffuse energy, and the second signal energy parameter further comprises one of the following: generating the at least one second weighted value based on a product of the at least one second direct-to-total ratio parameter and the second signal energy parameter; generating the at least one second weighted value based on a product of a direct-to-total ratio parameter of the compensated diffuse energy and the second signal energy parameter; as well as The at least one second weighting value is generated based on a product of an average value of the direct-to-total ratio parameter of the compensated diffuse energy and the at least one second direct-to-total ratio parameter and the second signal energy parameter.

26. The method of any one of claims 17 to 25, wherein generating a direct to total ratio parameter to compensate for diffuse energy further comprises: At least one combined diffuse-to-total ratio is generated based at least in part on the first signal energy parameter, the second signal energy parameter, and a diffuse signal energy parameter.

27. The method according to claim 26, further comprising: For the first audio stream, obtaining at least one first diffuse-to-total ratio parameter; as well as The diffuse signal energy parameter is obtained based in part on the first signal energy parameter and the first diffuse-to-total ratio parameter.

28. The method of any one of claims 26 or 27, wherein generating a direct to total ratio parameter to compensate for diffuse energy further comprises: The at least one direct-to-total ratio parameter compensating for diffuse energy is generated based on the at least one combined diffuse-to-total ratio.

29. The method according to any one of claims 26 or 27, further comprising: At least one second diffuse-to-total ratio parameter is obtained for the second audio stream, wherein generating at least one merged diffuse-to-total ratio further comprises: generating the at least one merged diffuse-to-total ratio based on the second diffuse-to-total ratio parameter.

30. The method of claim 26, wherein generating the direct to total ratio parameter for compensating diffuse energy further comprises: generating at least one direct-to-total ratio parameter of an expected compensated diffuse energy based on the at least one combined diffuse-to-total ratio; comparing the at least one direct-to-total ratio parameter of the expected compensated diffuse energy with the second direct-to-total ratio parameter; Based on the comparison, the direct-to-total ratio parameter of the compensated diffuse energy is selected from the direct-to-total ratio parameter of the at least one expected compensated diffuse energy and the second direct-to-total ratio parameter, so that the direct-to-total ratio parameter of the compensated diffuse energy is the smaller of the direct-to-total ratio parameter of the at least one expected compensated diffuse energy and the second direct-to-total ratio parameter.

31. The method according to any one of claims 17 to 30, further comprising: A selected first direct-to-total ratio parameter of the at least one first direct-to-total ratio parameter and the direct-to-total ratio parameter for compensating diffuse energy are encoded.

32. The method according to any one of claims 17 to 31, further comprising: The first audio stream and the second audio stream are merged.

Citation Information

Patent Citations

  • Spatial audio parameter merging

    GB2574238A