Spatial sound rendering

By analyzing the environmental energy distribution and directional information of spatial audio signals, the synthesized output audio signals are used to control the energy distribution of the direct and diffuse portions, thus solving the problem of uneven sound environment energy reproduction in existing technologies and achieving a more accurate spatial sound rendering effect.

CN115209337BActive Publication Date: 2026-03-27NOKIA TECHNOLOGIES OY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-03-25
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately reproduce the environmental energy distribution of sound in spatial sound rendering, especially when the sound scene is uneven, resulting in poor reproduction.

Method used

By receiving spatial audio signals and analyzing their environmental energy distribution and directional information, the output audio signal is synthesized using directional parameters and environmental energy distribution parameters. The energy distribution of the direct and diffuse parts is controlled to form a directional mode filter signal, which is then processed on the basis of frequency band to achieve accurate spatial rendering.

Benefits of technology

It achieves accurate sound reproduction under different sound field conditions, and can control the uniform distribution of environmental energy or the original sound field distribution during the rendering process, thereby improving the quality of spatial audio reproduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115209337B_ABST
    Figure CN115209337B_ABST
Patent Text Reader

Abstract

An apparatus for spatial audio signal decoding, the apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: receive at least one associated audio signal, the at least one associated audio signal being based on a spatial audio signal; spatial metadata associated with the at least one associated audio signal, the spatial metadata comprising at least one parameter representing an ambient energy distribution of the spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal; synthesize at least one output audio signal from the at least one associated audio signal based on the at least one directional parameter and the at least one parameter, wherein the at least one parameter controls an ambient energy distribution of the at least one output signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of patent application number 201980035666.1 filed on 25 March 2019 with the title “Spatial sound rendering”. TECHNICAL FIELD

[0002] The present application relates to apparatuses and methods for spatial sound rendering. This includes, but is not limited to, spatial sound rendering for multi-channel loudspeaker setups. BACKGROUND

[0003] Parametric spatial audio processing is a field of audio signal processing in which a set of parameters is used to describe the spatial aspects of a sound. For example, in parametric spatial audio capture from a microphone array, it is a typical and effective choice to estimate from the microphone array signal a set of parameters such as the direction of the sound in a frequency band, and a ratio parameter representing the relative energy of directional and non-directional parts of the captured sound in a frequency band. It is well known that these parameters describe well the perceived spatial characteristics of the captured sound at the location of the microphone array. These parameters can be used accordingly for the synthesis of spatial sound, for binaural headphones, for loudspeakers or other formats such as Ambisonics.

[0004] Thus, direction in a frequency band and direct-to-total energy ratio are particularly effective parameters for the parametrization of spatial audio capture.

[0005] A set of parameters consisting of a direction parameter in a frequency band and an energy ratio parameter in a frequency band, indicating the proportion of directional sound energy, can also be used as spatial metadata for an audio codec. For example, these parameters can be estimated from an audio signal captured from a microphone array, and for example a stereo signal can be generated from the microphone array signal to be transmitted together with the spatial metadata. The stereo signal can for example be encoded with an AAC encoder. A decoder can decode the audio signal to a PCM signal and process the sound in the frequency band (using the spatial metadata) to obtain a spatial output, for example a binaural output.

[0006] The parametric encoder input format can be one or several input formats. An example input format is the first order ambisonics (FOA) format. The analysis of FOA input for spatial metadata extraction is documented in scientific literature related to directional audio coding (DirAC) and harmonic planewave expansion (Harpex). This is because there exist professional microphone arrays that are capable of directly providing FOA signals (or, specifically, a variant, B-format signals), and analysis of such input has been implemented. SUMMARY

[0007] An apparatus comprising at least one processor and at least one memory including computer program code is provided, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to receive at least one associated audio signal, the at least one associated audio signal being based on a spatial audio signal; spatial metadata associated with the at least one associated audio signal, the spatial metadata comprising at least one parameter representing an ambient energy distribution of the spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal; synthesize at least one output audio signal from the at least one associated audio signal based on the at least one directional parameter and the at least one parameter, wherein the at least one parameter controls an ambient energy distribution of the at least one output signal.

[0008] The apparatus caused to synthesize at least one output audio signal from the at least one associated audio signal based on the at least one directional parameter and the at least one parameter, wherein the at least one parameter controls an ambient energy distribution of the at least one output signal, can be further caused to: divide the at least one associated audio signal into a direct portion and a diffuse portion based on the spatial metadata; synthesize a direct audio signal based on the direct portion of the at least one associated audio signal and the at least one directional parameter; determine a diffuse portion gain based on the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal; synthesize a diffuse audio signal based on the diffuse portion of the at least one associated audio signal and the diffuse portion gain; and combine the direct audio signal and the diffuse audio signal to generate the at least one output audio signal.

[0009] The apparatus caused to synthesize a diffuse audio signal based on the diffuse portion of the at least one associated audio signal can be caused to decorrelate the at least one associated audio signal.

[0010] The apparatus caused to determine the diffuse portion gain based on the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal can be caused to: determine a direction to which a set of prototype output signals points; for each of the set of prototype output signals, determine whether a direction of the prototype output signal is within a sector defined by the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal; set a gain associated with prototype output signals within the sector to be on average greater than a gain associated with prototype output signals outside the sector.

[0011] The means caused to set the gain associated with prototype output signals within the sector to be on average greater than the gain associated with prototype output signals outside the sector can be caused to set the gain associated with prototype output signals within the sector to be 1 ; to set the gain associated with prototype output signals outside the sector to be 0; and to normalize the sum of squares of the gains to unity.

[0012] The means caused to receive spatial metadata comprising at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representative of directional information of the spatial audio signal can be caused to perform at least one of: analyzing the at least one spatial audio signal to determine the at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal; and receiving the at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal.

[0013] The at least one directional parameter representative of directional information of the spatial audio signal can comprise at least one of: at least one directional parameter representative of a direction of arrival; a diffuse parameter associated with the at least one directional parameter; and an energy ratio parameter associated with the at least one directional parameter.

[0014] The at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal can comprise at least one of: a first parameter comprising at least one azimuth and / or at least one elevation associated with the at least one spatial sector having a locally maximum average ambient energy; at least one other parameter based on an angular extent of the at least one spatial sector having the locally maximum average ambient energy.

[0015] The at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal can be a parameter representative on a per frequency band basis.

[0016] According to a second aspect, there is provided an apparatus for spatial audio signal processing, the apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: receive at least one spatial audio signal; determine at least one associated audio signal from the at least one spatial audio signal; determine spatial metadata associated with the at least one associated audio signal, wherein the spatial metadata comprises at least one parameter representing an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal; transmit and / or store: the associated audio signal and the spatial metadata comprising the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal.

[0017] The apparatus caused to determine the spatial metadata comprising the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal can be further caused to: form directional pattern filter signals to several spatial directions defined by azimuth and / or elevation based on the at least one spatial audio signal; determine a weighted time average of ambient energy per spatial sector based on the directional pattern filter signals; determine at least one spatial sector having a local maximum average ambient energy and generate a first parameter comprising at least one azimuth and / or at least one elevation associated with the at least one spatial sector having the local maximum average ambient energy; determine a range angle of the local maximum average ambient energy based on a comparison of average ambient energies of neighboring spatial sectors and generate at least one further parameter based on the range angle of the at least one spatial sector having the local maximum average ambient energy.

[0018] The apparatus caused to form directional pattern filter signals to several spatial directions defined by azimuth and / or elevation based on the at least one spatial audio signal can be caused to: form a virtual cardioid signal defined by the azimuth and / or the elevation.

[0019] The apparatus caused to determine spatial metadata associated with the at least one spatial audio signal, wherein the spatial metadata comprises at least one parameter representing an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal, can be caused to: determine spatial metadata on a per frequency band basis.

[0020] According to a third aspect, a method for spatial audio signal decoding is provided, the method comprising: receiving at least one associated audio signal, the at least one associated audio signal being based on a spatial audio signal; spatial metadata associated with the at least one associated audio signal, the spatial metadata comprising at least one parameter representing an ambient energy distribution of the spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal; synthesizing at least one output audio signal from the at least one associated audio signal based on the at least one directional parameter and the at least one parameter, wherein the at least one parameter controls an ambient energy distribution of the at least one output signal.

[0021] Synthesizing at least one output audio signal from the at least one associated audio signal based on the at least one directional parameter and the at least one parameter, wherein the at least one parameter controls an ambient energy distribution of the at least one output signal, can further comprise: dividing the at least one associated audio signal into a direct portion and a diffuse portion based on the spatial metadata; synthesizing a direct audio signal based on the direct portion of the at least one associated audio signal and the at least one directional parameter; determining a diffuse portion gain based on the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal; synthesizing a diffuse audio signal based on the diffuse portion of the at least one associated audio signal and the diffuse portion gain; and combining the direct audio signal and the diffuse audio signal to generate the at least one output audio signal.

[0022] Synthesizing a diffuse audio signal based on the diffuse portion of the at least one associated audio signal can comprise de-correlating the at least one associated audio signal.

[0023] Determining the diffuse portion gain based on the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal can comprise: determining a direction to which a set of prototype output signals is directed; for each of the set of prototype output signals, determining whether the direction of the prototype output signal is within a sector defined by the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal; setting a gain associated with the prototype output signals within the sector to be on average greater than a gain associated with the prototype output signals outside the sector.

[0024] Setting a gain associated with the prototype output signals within the sector to be on average greater than a gain associated with the prototype output signals outside the sector can comprise: setting a gain associated with the prototype output signals within the sector to be 1 ; setting a gain associated with the prototype output signals outside the sector to be 0; and normalizing a sum of squares of the gains to a unit value.

[0025] Receiving spatial metadata comprising at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representative of directional information of the spatial audio signal can comprise at least one of: analyzing the at least one spatial audio signal to determine the at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal; and receiving the at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal.

[0026] The at least one directional parameter representative of directional information of the spatial audio signal can comprise at least one of: at least one direction parameter representative of a direction of arrival; a diffuse parameter associated with the at least one direction parameter; and an energy ratio parameter associated with the at least one direction parameter.

[0027] The at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal can comprise at least one of: a first parameter comprising at least one azimuth and / or at least one elevation associated with the at least one spatial sector having a locally maximum average ambient energy; at least one other parameter based on an angular extent of the at least one spatial sector having the locally maximum average ambient energy.

[0028] The at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal can be a parameter representative on a per frequency band basis.

[0029] According to a fourth aspect, a method for spatial audio signal processing is provided, the method comprising: receiving at least one spatial audio signal; determining at least one associated audio signal from the at least one spatial audio signal; determining spatial metadata associated with the at least one associated audio signal, wherein the spatial metadata comprises at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representative of directional information of the spatial audio signal; transmitting and / or storing: the associated audio signal and the spatial metadata comprising the at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representative of directional information of the spatial audio signal.

[0030] Determining the spatial metadata comprising the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal can further comprise: forming directional pattern filtered signals to several spatial directions defined by an azimuth angle and / or an elevation angle based on the at least one spatial audio signal; determining a weighted time average of the ambient energy per spatial sector based on the directional pattern filtered signals; determining at least one spatial sector having a local maximum average ambient energy and generating a first parameter comprising at least one azimuth angle and / or at least one elevation angle associated with the at least one spatial sector having the local maximum average ambient energy; determining a range angle of the local maximum average ambient energy based on a comparison of the average ambient energy of neighboring spatial sectors with the local maximum average ambient energy and generating at least one further parameter based on the range angle of the at least one spatial sector having the local maximum average ambient energy.

[0031] Forming directional pattern filtered signals to several spatial directions defined by an azimuth angle and / or an elevation angle based on the at least one spatial audio signal can comprise forming a virtual cardioid signal defined by the azimuth angle and / or the elevation angle.

[0032] Determining spatial metadata associated with the at least one spatial audio signal, wherein the spatial metadata comprises at least one parameter representing an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal, can comprise determining spatial metadata on a per frequency band basis.

[0033] According to a fifth aspect, there is provided an apparatus comprising means for: receiving at least one associated audio signal, the at least one associated audio signal being based on a spatial audio signal; spatial metadata associated with the at least one associated audio signal, the spatial metadata comprising at least one parameter representing an ambient energy distribution of the spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal; synthesizing at least one output audio signal from the at least one associated audio signal based on the at least one directional parameter and the at least one parameter, wherein the at least one parameter controls an ambient energy distribution of the at least one output signal.

[0034] The module for synthesizing at least one output audio signal from the at least one associated audio signal based on the at least one directional parameter and the at least one parameter, wherein the at least one parameter controls an ambient energy distribution of the at least one output signal, can further be configured for: dividing the at least one associated audio signal into a direct portion and a diffuse portion based on the spatial metadata; synthesizing a direct audio signal based on the direct portion of the at least one associated audio signal and the at least one directional parameter; determining a diffuse portion gain based on the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal; synthesizing a diffuse audio signal based on the diffuse portion of the at least one associated audio signal and the diffuse portion gain; and combining the direct audio signal and the diffuse audio signal to generate the at least one output audio signal.

[0035] The module for synthesizing a diffuse audio signal based on the diffuse portion of the at least one associated audio signal can be configured for de-correlating the at least one associated audio signal.

[0036] The module for determining the diffuse portion gain based on the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal can be configured for: determining a direction to which a set of prototype output signals points; for each of the set of prototype output signals, determining whether the direction of the prototype output signal is within a sector defined by the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal; setting a gain associated with a prototype output signal within the sector to be on average greater than a gain associated with a prototype output signal outside the sector.

[0037] The module for setting a gain associated with a prototype output signal within the sector to be on average greater than a gain associated with a prototype output signal outside the sector can be configured for: setting a gain associated with a prototype output signal within the sector to be 1; setting a gain associated with a prototype output signal outside the sector to be 0; and normalizing a sum of squares of the gains to a unit value.

[0038] The module for receiving spatial metadata comprising at least one parameter representing an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal can be configured for at least one of: analyzing the at least one spatial audio signal to determine the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal; and receiving the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal.

[0039] The at least one directional parameter representing directional information of the spatial audio signal can comprise at least one of: at least one direction parameter representing a direction of arrival; a diffuse parameter associated with the at least one direction parameter; and an energy ratio parameter associated with the at least one direction parameter.

[0040] The at least one parameter representing an ambient energy distribution of the at least one spatial audio signal can comprise at least one of: a first parameter comprising at least one azimuth and / or at least one elevation associated with the at least one spatial sector having a locally maximum average ambient energy; at least one further parameter based on an angular extent of the at least one spatial sector having the locally maximum average ambient energy.

[0041] The at least one parameter representing an ambient energy distribution of the at least one spatial audio signal can be a parameter represented on a per frequency band basis.

[0042] According to a sixth aspect, there is provided an apparatus for spatial audio signal processing, the apparatus comprising means for: receiving at least one spatial audio signal; determining at least one associated audio signal from the at least one spatial audio signal; determining spatial metadata associated with the at least one associated audio signal, wherein the spatial metadata comprises at least one parameter representing an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal; transmitting and / or storing: the associated audio signal and the spatial metadata comprising the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal.

[0043] The module for determining spatial metadata comprising at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representative of directional information of the spatial audio signal can further be configured to form directional pattern filter signals to several spatial directions defined by an azimuth angle and / or an elevation angle based on the at least one spatial audio signal, to determine a weighted time average of ambient energy per spatial sector based on the directional pattern filter signals, to determine at least one spatial sector with a local maximum average ambient energy and to generate a first parameter comprising at least one azimuth angle and / or at least one elevation angle associated with the at least one spatial sector with the local maximum average ambient energy, to determine a range angle of the local maximum average ambient energy based on a comparison of average ambient energies of neighboring spatial sectors and to generate at least one further parameter based on the range angle of the at least one spatial sector with the local maximum average ambient energy.

[0044] The module for forming directional pattern filter signals to several spatial directions defined by an azimuth angle and / or an elevation angle based on the at least one spatial audio signal can be configured to form a virtual cardioid signal defined by the azimuth angle and / or the elevation angle.

[0045] The module for determining spatial metadata associated with the at least one spatial audio signal, wherein the spatial metadata comprises at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representative of directional information of the spatial audio signal, can be configured to determine spatial metadata on a per frequency band basis.

[0046] According to a seventh aspect, there is provided an apparatus comprising: receiving circuitry configured to receive at least one associated audio signal, the at least one associated audio signal being based on a spatial audio signal; spatial metadata associated with the at least one associated audio signal, the spatial metadata comprising at least one parameter representative of an ambient energy distribution of the spatial audio signal and at least one directional parameter indicative of directional information of the spatial audio signal; synthesis circuitry configured to synthesize at least one output audio signal from the at least one associated audio signal based on the at least one directional parameter and the at least one parameter, wherein the at least one parameter controls an ambient energy distribution of the at least one output signal.

[0047] According to an eighth aspect, there is provided an apparatus for spatial audio signal processing, the apparatus comprising: receiving circuitry configured to receive at least one spatial audio signal; determining circuitry configured to determine at least one associated audio signal from the at least one spatial audio signal; and determining circuitry configured to determine spatial metadata associated with the at least one associated audio signal, wherein the spatial metadata comprises at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal, and at least one directional parameter representative of directional information of the spatial audio signal; transmitting and / or storing circuitry configured to transmit and / or store: the associated audio signal and the spatial metadata, the spatial metadata comprising the at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal, and at least one directional parameter representative of directional information of the spatial audio signal.

[0048] According to a ninth aspect, there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus to perform at least the following: receiving at least one associated audio signal, the at least one associated audio signal being based on a spatial audio signal; spatial metadata associated with the at least one associated audio signal, the spatial metadata comprising at least one parameter representative of an ambient energy distribution of the spatial audio signal, and at least one directional parameter representative of directional information of the spatial audio signal; synthesizing at least one output audio signal from the at least one associated audio signal based on the at least one directional parameter and the at least one parameter, wherein the at least one parameter controls an ambient energy distribution of the at least one output signal.

[0049] According to a tenth aspect, there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus to perform at least the following: receiving at least one spatial audio signal; determining at least one associated audio signal from the at least one spatial audio signal; determining spatial metadata associated with the at least one associated audio signal, wherein the spatial metadata comprises at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal, and at least one directional parameter representative of directional information of the spatial audio signal; transmitting and / or storing: the associated audio signal and the spatial metadata, the spatial metadata comprising the at least one parameter representative of an ambient energy distribution of the at least one spatial audio signal, and at least one directional parameter representative of directional information of the spatial audio signal.

[0050] According to an eleventh aspect, there is provided a non-transitory computer- readable medium comprising program instructions for causing an apparatus to perform at least the following: receiving at least one associated audio signal, the at least one associated audio signal being based on a spatial audio signal; spatial metadata associated with the at least one associated audio signal, the spatial metadata comprising at least one parameter representing an ambient energy distribution of the spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal; synthesizing at least one output audio signal from the at least one associated audio signal based on the at least one directional parameter and the at least one parameter, wherein the at least one parameter controls an ambient energy distribution of the at least one output signal.

[0051] According to a twelfth aspect, there is provided a non-transitory computer- readable medium comprising program instructions for causing an apparatus to perform at least the following: receiving at least one spatial audio signal; determining at least one associated audio signal from the at least one spatial audio signal; determining spatial metadata associated with the at least one associated audio signal, wherein the spatial metadata comprises at least one parameter representing an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal; transmitting and / or storing the associated audio signal and the spatial metadata, the spatial metadata comprising the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal.

[0052] According to a thirteenth aspect, there is provided a non-transitory computer- readable medium comprising program instructions for causing an apparatus to perform at least the following: receiving at least one associated audio signal, the at least one associated audio signal being based on a spatial audio signal; spatial metadata associated with the at least one associated audio signal, the spatial metadata comprising at least one parameter representing an ambient energy distribution of the spatial audio signal and at least one directional parameter indicating directional information of the spatial audio signal; synthesizing at least one output audio signal from the at least one associated audio signal based on the at least one directional parameter and the at least one parameter, wherein the at least one parameter controls an ambient energy distribution of the at least one output signal.

[0053] According to a fourteenth aspect, there is provided a computer readable medium comprising program instructions for causing an apparatus to perform at least the following: receiving at least one spatial audio signal; determining at least one associated audio signal from the at least one spatial audio signal; determining spatial metadata associated with the at least one associated audio signal, wherein the spatial metadata comprises at least one parameter representing an ambient energy distribution of the at least one spatial audio signal, and at least one directional parameter representing directional information of the spatial audio signal; a transmitting and / or storing circuitry is configured to transmit and / or store: the associated audio signal and the spatial metadata comprising the at least one parameter representing an ambient energy distribution of the at least one spatial audio signal and at least one directional parameter representing directional information of the spatial audio signal.

[0054] A non-transitory computer readable medium comprising program instructions for causing an apparatus to perform the above method. An apparatus configured to perform the actions of the above method.

[0055] A computer program comprising program instructions for causing a computer to perform the above method.

[0056] A computer program product stored on a medium can cause an apparatus to perform the methods described herein.

[0057] An electronic device can comprise an apparatus as described herein.

[0058] A chipset can comprise an apparatus as described herein.

[0059] Embodiments of the present application aim to address problems associated with the prior art. BRIEF DESCRIPTION OF DRAWINGS

[0060] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings in which:

[0061] Figure 1 An example spatial capture and synthesizer according to some embodiments is schematically illustrated;

[0062] Figure 2 A flowchart illustrating a method of operating an example spatial capture and synthesizer according to some embodiments is shown;

[0063] Figure 3 A flowchart illustrating an example method of determining operating an example spatial synthesizer according to some embodiments is shown;

[0064] Figure 4 An example of an ambient energy distribution parameter definition according to some embodiments is shown;

[0065] Figure 5An example spatial synthesizer according to some embodiments is shown schematically;

[0066] Figure 6 A flowchart of an example method of operating an example spatial synthesizer according to some embodiments is shown;

[0067] Figure 7 A flowchart of an example method of determining diffuse stream gain based on ambient energy distribution parameters is shown;

[0068] Figure 8 A further example spatial capture and synthesizer according to some embodiments is shown schematically; and

[0069] Figure 9 An example apparatus suitable for implementing the illustrated devices is shown schematically. DETAILED DESCRIPTION

[0070] Suitable apparatus and possible mechanisms for providing efficient spatial processing and rendering based on a range of audio input formats are described in further detail below.

[0071] Spatial metadata consisting of direct-to-total energy ratios (or diffuse ratios) parameters in directions and frequency bands is particularly suitable for expressing the perceptual characteristics of natural sound fields.

[0072] However, sound scenes can be of various types, and in some cases the sound field has a non-uniform ambient energy distribution (e.g. ambient only or predominantly at certain axes or spatial regions). Concepts as discussed in embodiments herein describe apparatus and methods that accurately reproduce the spatial distribution of diffuse / ambient sound energy at the reproduced sound when compared to the original spatial sound.

[0073] In some embodiments this can be selectable, and thus the effect can be controlled during rendering to determine whether it is intended to reproduce a uniform distribution of ambient energy or to reproduce the distribution of ambient energy of the original sound scene. In different embodiments, reproducing a uniform distribution of ambient energy can refer to the ambient energy being distributed uniformly to different output channels, or to the ambient energy being distributed in a spatially balanced manner.

[0074] Concepts to be discussed in further detail below are to add an ambient energy distribution metadata field or parameter in the bitstream, and to utilize this field or parameter during rendering to enable reproduction of spatial audio such that it more closely represents the original sound field.

[0075] Thus, the embodiments described in the following relate to audio encoding and decoding using parametrization related to a soundfield (direction and ratio in frequency bands), and wherein the embodiments aim at improving the reproduction quality of a soundfield encoded with the aforementioned parametrization. Furthermore, the embodiments describe cases of improving the ambience quality by conveying an ambience energy distribution parameter together with the directional parameters, and reproducing the sound based on the directional parameters and the ambience energy distribution parameter, such that the ambience energy distribution parameter influences the diffuse stream synthesis using direction and ratio in frequency bands.

[0076] In particular, the embodiments discussed in the following are configured to use the ambience energy distribution parameter to modify the diffuse stream synthesis, such that the energy distribution of the soundfield is better reproduced.

[0077] In some embodiments, the ambience energy distribution parameter comprises at least a direction and a range or width associated with the analyzed ambience energy distribution.

[0078] In some embodiments, the input / processing can be implemented for first order ambisonic (FOA) input and for higher order ambisonic (HOA) input. In embodiments using HOA input instead of forming a virtual cardioid signal as described below with respect to FOA input, the method can replace the virtual cardioid signal c(k, n) with a signal having a side-oriented mode (or predominantly a side-oriented mode) formed from zeroth order to second order or higher HOA components or any suitable means to generate a signal with a side-oriented mode from the HOA signal.

[0079] With respect to Figure 1 , an example spatial capture and synthesizer is shown according to some embodiments. The spatial capture and synthesizer is shown in this example receiving a spatial audio signal 100 as input. The spatial audio signal 100 can be any suitable audio signal format, e.g., a microphone audio signal captured by a plurality of microphones or a microphone comprising a microphone array, a synthesized audio signal, a loudspeaker channel format audio signal, or a first order ambisonic (FOA) format or a variant thereof (e.g., a B-format signal) or higher order ambisonic (HOA).

[0080] In some embodiments, a converter (e.g., loudspeaker or microphone input to FOA converter) 101 is configured to receive the input audio signal 100 and convert it to a suitable FOA format signal 102.

[0081] In some embodiments, the converter 101 is configured to generate the FOA signal from a loudspeaker mix based on knowledge of the positions of the channels in the input audio signal. In other words, the w i (t), x i (t), y i (t), z i(t) components can be generated from the loudspeaker signals s i and ele i i (t) by:

[0082]

[0083] The w, x, y, z signals are generated for each loudspeaker (or object) signal s i with its own azimuth and elevation direction.

[0084] The output signal combining all these signals can be calculated as In other words, each loudspeaker or channel signal is combined into a total FOA signal.

[0085] In some embodiments, the converter 101 is configured to generate FOA signals from the microphone array signals according to any suitable method. The converter can use a linear method to obtain FOA signals from the microphone signals, in other words, applying a filter matrix or complex gain matrix in the frequency band to obtain FOA signals from the microphone array signals. The converter can be configured to extract features from the audio signals and process the signals differently depending on these features. The embodiments described herein describe an adaptive processing at least in some frequency bands and / or spherical harmonic orders and / or spatial dimensions. Thus, in contrast to traditional ambisonics, there is no linear correspondence between the output and the input. In some embodiments, the output of the converter is in the time-frequency domain. In other words, in some embodiments, the converter 101 is configured to apply a suitable time-frequency transform. In some embodiments, the input spatial audio 100 is in the time-frequency domain or can be passed through a suitable transform or filter bank.

[0086] In some embodiments, the converter uses a matrix of designed linear filters for the microphone signals to obtain the spherical harmonic components. An equivalent alternative method is to transform the microphone signals to the time-frequency domain and use a designed mixing matrix for each frequency band to obtain the spherical harmonic signals in the time-frequency domain. Another conversion method is one in which spatial audio capture (SPAC) techniques represent methods for spatial audio capture from a microphone array and output ambisonics format based on dynamic SPAC analysis. Spatial audio capture (SPAC) refers here to techniques using adaptive time-frequency analysis and processing to provide high perceptual quality spatial audio reproduction from any device equipped with a microphone array. SPAC capture is done in the horizontal plane requires at least 3 microphones, while for 3D capture at least 4 microphones are required. The SPAC methods are adaptive, in other words, they use non-linear methods to improve spatial accuracy from current conventional linear capture techniques.

[0087] ​In this document, the term SPAC is used as a broad term encompassing any adaptive array signal processing technique providing spatial audio capture. The method within the scope applies analysis and processing to the frequency band signals as this is a domain that is meaningful to the spatial hearing perception. Spatial metadata is dynamically analyzed in the frequency band, e.g. direction of arriving sound and / or determining directional or non-directional ratio or energy parameters of the recorded sound. The metadata is applied in the reproduction phase in order to dynamically synthesize the spatial sound to headphones or loudspeakers or to ambisonics (e.g. FOA) output with high spatial accuracy. For example, a plane wave arriving at the array can be reproduced as a point source at the receiver end.

[0088] One method of spatial audio capture (SPAC) reproduction is Directional Audio Coding (DirAC), which is a method using soundfield intensity and energy analysis to provide spatial metadata enabling high quality adaptive spatial audio synthesis for loudspeakers or headphones. Another example is Harmonic Plane Wave Expansion (Harpex), which is a method that can analyze two plane waves simultaneously, which can further improve spatial accuracy under certain soundfield conditions. Another method is a method primarily used for mobile phone spatial audio capture, which uses delay and coherence analysis between microphones to obtain spatial metadata, and variants thereof for devices containing more microphones. Although two variants are described in the following examples, any suitable method applied to obtain spatial metadata can be used.

[0089] The spatial analyzer 103 can be configured to receive the FOA signal 102 and generate suitable spatial parameters, e.g. directions 106 and ratios 108. The spatial analyzer 103 can for example be a computer or a mobile phone (running suitable software), or alternatively a specific device using for example a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). In some embodiments where the converter 101 employs a spatial audio capture technique to convert the input audio signal format to a FOA format signal, then the spatial analyzer 103 can comprise the converter 101, or the converter can comprise the spatial analyzer 103.

[0090] An example of a suitable spatial analysis method is Directional Audio Coding (DirAC). The DirAC method can estimate the direction and diffuse ratio (equivalent to direct-to-total ratio parameter information) from a first order ambisonics (FOA) signal.

[0091] In some embodiments, the DirAC method transforms the FOA signal into frequency bands using a suitable time-to-frequency domain transform, e.g. using a short-time Fourier transform (STFT), resulting in time-frequency signals w(k, n), x(k, n), y(k, n), z(k, n), where k is the frequency bin index and n is the time index. In such examples, the DirAC method can estimate the strength vectors by

[0092]

[0093] where Re denotes the real part and the asterisk * denotes the complex conjugate. The strength represents the direction of the propagating sound energy, and the direction parameter can thus be determined by the opposite direction of the strength vector. In some embodiments, the strength vectors can be averaged over several time and / or frequency indices before determination of the direction parameter.

[0094] Furthermore, in some embodiments, the DirAC method can determine the diffuse (pseudo-ambience) based on the FOA components. In SN3D-normalization of diffuse sound, the sum of the energy of all ambience components in a first order is equal. For example, if the zeroth order W has 1 unit of energy, then each first order X Y Z has 1 / 3 unit (summing to 1) of energy. And so on to higher orders.

[0095] The diffuse can thus be determined as

[0096]

[0097] The diffuse is a ratio value, which is 1 when the sound is completely ambient and 0 when the sound is completely directional. In some embodiments, all parameters in the equation are typically averaged over time and / or frequency. In certain systems, the expectation operator E[] can be replaced by the averaging operator.

[0098] In some embodiments, the direction parameter and the diffuse parameter can be analyzed from FOA components that have been acquired in two different ways. In particular, in this embodiment, the direction parameter can be analyzed from the signals as described above. The diffuse can be analyzed from another set of FOA signals denoted as and described in more detail below. As one particular example, consider the conversion of 5.0 inputs with loudspeakers (azimuth angles of 0, + / -30 and + / -110 (all elevations are zero, cos(ele i ) = 1, sin(ele i ) = 0) to FOA components. The FOA components used for the direction parameter analysis are acquired in the following way:

[0099]

[0100] Diffusion can be analyzed from another set of FOA signals obtained as follows:

[0101]

[0102] wherein is the modified virtual speaker position. The modified virtual speaker positions for diffusion analysis are obtained such that the virtual speakers are positioned with a uniform spacing when the FOA signals are created. The benefit of this uniform spaced positioning of the virtual speakers for diffusion analysis is that incoherent sound arrives uniformly from different directions around the virtual microphone and the time average of the intensity vectors adds up to close to zero. In the case of 5.0, the modified virtual speaker positions are 0, + / - 72, + / - 144 degrees. Thus, the virtual speakers have a constant spacing of 72 degrees.

[0103] Similar modified virtual speaker positions can be created for other speaker configurations to ensure constant spacing between adjacent speakers. In embodiments of the invention, the spacing of the modified virtual speakers is obtained by dividing the entire 360 degrees by the number of speakers in the horizontal plane. The modified virtual speaker positions are then obtained by positioning the virtual speakers with the obtained spacing starting from the center speaker or other suitable starting speaker.

[0104] In some embodiments, an alternative ratio parameter, e.g. direct-to-total energy ratio, can be determined, which can be obtained by:

[0105] r(k, n) = 1 - ψ(k, n)

[0106] When averaging, the diffusion (and direction) parameters can be determined in frequency bands that combine several frequency bins k, e.g. approximating the Bark frequency resolution.

[0107] As mentioned above, DirAC is one possible spatial analysis method option to determine the direction and ratio metadata. The spatial audio parameters, also referred to as spatial metadata or metadata, can be determined according to any suitable method. For example, by simulating a microphone array and using a spatial audio capture (SPAC) algorithm. Furthermore, the spatial metadata can include (but is not limited to): direction and direct-to-total energy ratio; direction and diffusion; inter-channel level difference, inter-channel phase difference, and inter-channel coherence. In some embodiments, these parameters are determined in the time-frequency domain. It should be noted that other parameterizations than the ones mentioned above can also be used. In general, the spatial audio parameterization typically describes how the sound is distributed in space, generally (e.g. using direction) or relatively (e.g. as a level difference between certain channels).

[0108] The transport signal generator 105 is further configured to receive the FOA signal 102 and generate suitable transport audio signals 110. The transport audio signals can also be referred to as associated audio signals, and are spatial audio signals that contain directional information of a soundfield and are input into the system. It should be understood that the soundfield in this context can refer to a captured natural soundfield with directional information, or to a surround sound scene with directional information created with known mixing and audio processing means. The transport signal generator 105 can be configured to generate any suitable number of transport audio signals (or channels), for example, in some embodiments, the transport signal generator is configured to generate two transport audio signals. In some embodiments, the transport signal generator 105 is further configured to encode the audio signals. For example, in some embodiments, the audio signals can be encoded using Advanced Audio Coding (AAC) or Enhanced Voice Services (EVS) compression coding. In some embodiments, the transport signal generator 105 can be configured to equalize the audio signals, apply automatic noise control, dynamic processing, or any other suitable processing. In some embodiments, the transport signal generator 105 can take the output of the spatial analyzer 103 as input to facilitate the generation of the transport signals 110. In some embodiments, instead of the FOA signal 102, the transport signal generator 105 can employ the spatial audio signal 100 to generate the transport signals.

[0109] The ambient energy distribution analyzer 107 can be further configured to receive the output of the spatial analyzer 103 and the FOA signal 102 and generate the ambient energy distribution parameters 104.

[0110] The ambient energy distribution parameters 104, the spatial metadata (direction 106 and ratio 108), and the transport audio signals 110 can be transmitted or stored, for example, in some storage device 107 (e.g. memory), or alternatively, directly processed in the same device. In some embodiments, the ambient energy distribution parameters 104, the spatial metadata 106, 108, and the transport audio signals 110 can be encoded or quantized or combined or multiplexed into a single data stream by suitable encoding and / or multiplexing operations. In some embodiments, the encoded audio signals are bundled together with a video stream (e.g. 360-degree video) in a media container, such as an mp4 container, to be transmitted to a suitable receiver.

[0111] The synthesizer 111 is configured to receive the ambient energy distribution parameters 104, the transport audio signals 110, the spatial parameters such as the direction 106 and the ratio 108, and generate the loudspeaker audio signals 112.

[0112] The synthesizer 111 can be configured to generate loudspeaker audio signals by employing spatial sound reproduction, e.g. by positioning sounds in a 3D space in arbitrary directions. The synthesizer 111 can be a computer or a mobile phone (running suitable software), or alternatively a specific device, e.g. utilizing an FPGA or an ASIC. Based on the data stream (transport audio signals and metadata). The synthesizer 111 can be configured to produce output audio signals. For headphone listening, the output signals can be binaural signals. In some other scenarios, the output signals can be ambisonic signals, or signals in some other desired output format.

[0113] In some embodiments, the spatial analyzer and synthesizer (as well as other components described herein) can be implemented within the same device, and can also be part of the same software.

[0114] With regard to Figure 2 , an example summary of the operation of the apparatus shown in Figure 1 is shown.

[0115] The initial operation is receiving spatial audio signals (e.g. loudspeaker-5.0 format, microphone format), as shown in Figure 2 by step 201.

[0116] The received loudspeaker format audio signals can be converted to FOA signals or streams, as shown in Figure 2 by step 203.

[0117] The converted FOA signals can be analyzed to generate spatial metadata (e.g. direction and / or energy ratio), as shown in Figure 2 by step 205.

[0118] Ambient energy distribution parameters can be determined from the converted FOA signals, and the output from the spatial analyzer is shown in Figure 2 by step 207.

[0119] The converted FOA signals can also be processed to generate transport audio signals, as shown in Figure 2 by step 209.

[0120] The ambient energy distribution parameters, transport audio signals and metadata can then optionally be combined to form a data stream, as shown in Figure 2 by step 211.

[0121] The ambient energy distribution parameters, transport audio signals and metadata (or combined data stream) can then be transmitted and received (or stored and retrieved), as shown in Figure 2 by step 213.

[0122] After the environmental energy distribution parameters, the transport audio signal and the metadata (or data stream) have been received or retrieved, the output audio signal can be synthesized based on at least the environmental energy distribution parameters, the transport audio signal and the metadata, as Figure 3 indicated by step 215.

[0123] The synthesized audio signal output signal can then be output to a suitable output.

[0124] With regard to Figure 3 the operation of the environmental energy distribution analyzer 107 is shown in further detail.

[0125] The analysis of the environmental energy distribution is based on the following: analyzing the environmental energy at a spatial sector as a function of time (in a frequency band), finding the direction of at least the maximum environmental energy, and parameterizing the environmental energy distribution based on at least the direction of the maximum environmental energy.

[0126] In some embodiments, the spatial sector for analyzing the environmental energy can be obtained by forming a virtual cardioid signal of the FOA signal to a desired spatial direction. The spatial direction is defined by an azimuth angle Θ and an elevation angle .

[0127] Thus, the environmental energy distribution analyzer can obtain a plurality of such spatial directions using this approach. The spatial directions can be obtained, for example, with azimuth angles in a uniform distribution of 45 degrees intervals.

[0128] In some embodiments, the environmental energy distribution analyzer can also convert the virtual cardioid signal c(k, n) to a spatial direction defined by an azimuth angle Θ and an elevation angle by first obtaining a dipole signal d(k, n). This can be generated, for example, by:

[0129]

[0130] where w(k, n), x(k, n), y(k, n), z(k, n) are the FOA time-frequency signals, where k is the bin index and n is the time index. w(k, n) is the omnidirectional signal, x(k, n), y(k, n), z(k, n) are the dipoles corresponding to the Cartesian coordinate axes. Then, the cardioid signal is obtained as

[0131]

[0132] Although this example describes the use of a cardioid pattern, any suitable pattern can be employed. The approach then computes the environmental energy in the spatial direction corresponding to the cardioid signal as:

[0133]

[0134] where N is the length of the discrete Fourier transform used to convert the signal to the frequency domain, and r(k, n) is the direct-to-total energy ratio.

[0135] In Figure 3 the generation of a virtual cardioid signal based on the FOA signal is shown by step 301.

[0136] The ambient energy distribution analyzer can then be configured to compute a weighted time average of the ambient energy per spatial sector. This can be obtained, for example, by:

[0137]

[0138] where a = 0.1.

[0139] In Figure 3 the generation of a weighted time average of the ambient energy per spatial sector is shown by step 303.

[0140] The ambient energy distribution analyzer can then be configured to determine the spatial sector with the largest average ambient energy. This can be determined as:

[0141]

[0142] where, maximizing the azimuth and elevation angle, of the value of at time n and bin k.

[0143] In Figure 3 the determination of the sector with the largest average ambient energy is shown by step 305.

[0144] The ambient energy distribution analyzer can then use the determined value as the "center" of the ambient energy distribution. In some embodiments, the ambient energy distribution analyzer can also store the maximum ambient energy value.

[0145]

[0146] In Figure 3 the storing of the azimuth and elevation angle of the sector with the largest average ambient energy is shown by step 307.

[0147] The ambient energy distribution analyzer can then determine the extent (or width or spread) of the ambient energy distribution. This can be done by checking the average ambient energy values on other spatial directions so that and is a neighboring spatial sector of p, s. If the context of this neighboring spatial sector is greater than the threshold times the maximum, the context spatial range is extended over the spatial sector p, s. That is, if the condition

[0148]

[0149] A suitable value for the threshold thr = 0.9. If the above condition is true for a spatial sector p, s, the context distribution range is extended over the spatial sector p, s.

[0150] In general, a suitable threshold parameter thr value can be obtained by inputting synthetic context signals with different known energy distributions into the analysis method and monitoring the estimated context energy distribution parameters with different thresholds. Furthermore, the audio signals synthesized with different context energy distribution parameter values obtained with different thresholds can be listened to and the threshold can be selected based on the parameter value giving the closest auditory perception to the original spatial audio field.

[0151] The above check of the average context energy value and the conditional inclusion into the context distribution range is then repeated for all neighboring spatial sectors. After the neighboring spatial sectors have been processed, the context energy distribution analyzer can then repeat the above process for those spatial sectors that satisfy the above condition. Thus, the context energy distribution analyzer can again check neighboring spatial sectors and extend the context energy distribution to span such spatial sectors that satisfy the above condition.

[0152] The range determination terminates when no spatial sectors remain or no more spatial sectors satisfy the condition. As a result, the process returns a list of spatial sectors with a context energy above the threshold. The range of the context energy distribution is defined such that it covers the found spatial sectors.

[0153] In Figure 3 the determination of the range of the context energy distribution is shown by step 309.

[0154] The range of the context energy distribution can then be stored as shown in Figure 4 by step 311.

[0155] The above process is able to find a contiguous spatial sector that dominates the context energy in a certain spatial sector (unimodal context energy distribution).

[0156] Such an example can be shown in Figure 4 by Figure 5The center of the ambient energy distribution defined by the ambianceAzi 401 vector within sector 411 is shown. In addition, the extent of the ambient energy distribution defined by the ambianceExtent 403 angle is also shown, which in this example extends to neighboring sectors labeled 412 and 413. In this example, ambianceAzi is equal to 45 / 2 degrees, and ambianceExtent is equal to 135 degrees.

[0157] In some embodiments, the ambient energy distribution analyzer can optionally determine a second ambient energy sector. This can be implemented in example embodiments in cases where the spatial sector corresponding to the second maximum ambient energy is sufficiently far away from the spatial sector corresponding to the maximum energy. For example, if it is approximately on the opposite side of the spatial audio field. In this case, a second center for the ambient energy distribution can be defined as the direction corresponding to the second maximum ambient energy value. A second portion of the ambient energy distribution can also take a range parameter in a similar manner to the first portion. This enables the ambient energy distribution analyzer to describe a bimodal ambient energy distribution, for example of audio sources at opposite sides of the spatial audio field.

[0158] In some embodiments, the ambient energy distribution analyzer can be configured to output the following parameters (which are signaled to the decoder / synthesizer:

[0159] ambianceAzi: degrees (azimuth of the center of the analyzed ambient energy distribution)

[0160] ambianceEle: degrees (elevation of the center of the analyzed ambient energy distribution)

[0161] ambianceExtent: degrees (width of the analyzed ambient energy distribution)

[0162] In some embodiments, there can be several of the above parameters, each describing a sector of significant ambient energy.

[0163] In some embodiments, there is a ratio parameter for each sector of the ambient energy distribution parameter. This ratio parameter describes the ratio of the ambient energy in a sector to the total ambient energy (ambianceSectorEnergyRatio).

[0164] These parameters can be updated at the encoder for each frame. In some embodiments, these parameters can be signaled at a lower rate (less frequently to the decoder / synthesizer). In some embodiments, a very low update rate (e.g. once per second) is sufficient. A slow update rate can ensure that the rendered spatial energy distribution does not change too quickly.

[0165] In some embodiments where the input is in a loudspeaker input format, some embodiments can perform the analysis directly on the loudspeaker channels. In these embodiments, instead of forming a virtual cardioid signal, the method can directly replace the virtual cardioid signal c(k, n) with the input loudspeaker channels in the time-frequency domain.

[0166] Further, in some embodiments, input / processing can be implemented for higher order ambisonic (HOA) input. In these embodiments, instead of forming a virtual cardioid signal, the method can replace the virtual cardioid signal c(k, n) with a signal having a side- directed pattern (or primarily a side-directed pattern) formed from zeroth order to second order or higher order HOA components or any suitable means, to generate a signal with a side-directed pattern from the HOA signal.

[0167] With respect to Figure 6 , an example synthesizer 111 according to some embodiments is shown.

[0168] In some embodiments, the input to the synthesizer 111 can be the direction 106, the ratio 108 spatial metadata, the transport audio signal stream 110 (which can have been decoded to a FOA signal), and the input environment energy distribution parameter 104. Other inputs to the system can be the enable / disable 550 input.

[0169] The prototype output signal generator 501 can be configured to receive the transport audio signal 110 and generate a prototype output signal therefrom. The transport audio signal stream 110 can be in the time domain and converted to the time-frequency domain before generating the prototype output signal. An example generation of a prototype signal from two transport signals can be by setting the left prototype output channel to a copy of the left transport channel, setting the right prototype output channel to a copy of the right transport channel, and the center (or mid) prototype channel to a mix of the left and right transport channels. An example of a prototype output signal is a virtual microphone signal which attempts to regenerate a virtual microphone signal when the transport signal is actually a FOA signal.

[0170] The square root (ratio) processor 503 can receive the ratio 108 and generate the square root of this value.

[0171] The first gain stage 509 (direct signal generator) can receive the square root of the ratio and apply it to the prototype output signal to generate a direct audio signal portion.

[0172] The VBAP 507 is configured to receive the direction 106 and generate appropriate VBAP gains.

[0173] An example method of generating the VBAP gains can be based on

[0174] 1) triangulating to the loudspeaker setup automatically,

[0175] 2) Select the appropriate triangle based on the direction (so that for a given direction, three loudspeakers are selected that form the triangle to which the given direction belongs), and

[0176] 3) Calculate the gains for the three loudspeakers that form the particular triangle.

[0177] In some embodiments, the VBAP gains (for each azimuth and elevation) and the loudspeaker triplet or other suitable number of loudspeakers or loudspeaker nodes (for each azimuth and elevation) can be pre-formulated as a lookup table stored in memory. In some embodiments, a real-time method then performs the amplitude panning by finding from memory the appropriate loudspeaker triplet (or number) for the desired panning direction and the gains for those loudspeakers that correspond to the desired panning direction.

[0178] The first stage of VBAP is to divide the 3D loudspeaker setup into triangles. There is no single solution for the generation of the triangles, and the loudspeaker setup can be triangulated in many ways. In some embodiments, one tries to find the smallest size triangles or polygons (with no loudspeakers inside the triangles and edges with as equal length as possible). In the general case, this is an effective method because it treats equally any direction of the auditory object and tries to minimize the distance to the loudspeakers used to create the auditory object in that direction.

[0179] Another computationally fast method for triangulation or virtual surface arrangement generation is to generate a convex hull from the data points determined by the loudspeaker angles. This is also a general method that treats equally all directions and data points.

[0180] The next stage or second stage is to select the appropriate triangle or polygon or virtual surface that corresponds to the panning direction.

[0181] The next stage is to formulate the panning gains that correspond to the panning direction.

[0182] The direct part gain stage 515 is configured to apply the VBAP gains to the direct part audio signal to generate the spatially processed direct part.

[0183] The square root (1 - ratio) processor 505 can receive the ratio 108 and generate the square root of the 1 - ratio value.

[0184] The second gain stage 511 (diffuse signal generator) can receive the square root of the 1 - ratio and apply it to the prototype output signal to generate the diffuse audio signal part.

[0185] The decorrelator 513 is configured to receive the diffuse audio signal portion from the second gain stage 511 and generate a decorrelated diffuse audio signal portion.

[0186] The diffuse portion gain determiner 517 can be configured to receive the enable / disable input and the input ambient energy distribution parameter 104. The enable / disable input can be configured to selectively enable or disable the following operations.

[0187] The diffuse portion gain determiner 517 can be configured to selectively (based on the input) distribute energy unevenly to different directions if the original spatial audio field has an uneven distribution of ambient energy. Thus, the energy distribution in the diffuse reproduction can be closer to the original sound field.

[0188] The diffuse gain stage 519 can be configured to receive the diffuse portion gain and apply it to the decorrelated diffuse audio signal portion.

[0189] The combiner 521 can then be configured to combine the processed diffuse audio signal portion and the processed direct signal portion and generate a suitable output audio signal. In some embodiments, these combined audio signals can be further converted into time domain form before being output to a suitable output device.

[0190] With regard to Figure 5 , a flowchart of the operation of the synthesizer 111 shown in Figure 6 is shown.

[0191] The method can comprise receiving the transport audio signal, the metadata, the (enable / disable parameter) and the input ambient energy distribution parameter 104, as shown in Figure 6 by step 601.

[0192] The method can further comprise generating a prototype output signal based on the transport audio signal, as shown in Figure 6 by step 603.

[0193] The method can further comprise determining a direct portion from the prototype output signal and the ratio metadata, as shown in Figure 6 by step 611.

[0194] The method can further comprise determining a diffuse portion from the prototype output signal and the ratio metadata, as shown in Figure 6 by step 607.

[0195] Applying VBAP to the direct portion, as shown in Figure 6 by step 613.

[0196] The method can further comprise determining a diffuse portion gain based on the input ambient energy distribution parameter 104 (and the enable / disable parameter), as shown in Figure 6The direct portion is then determined, as shown in step 603.

[0197] The method can further comprise applying a diffuse portion gain to the determined diffuse portion, as shown in step 605. Figure 6 The direct portion is then determined, as shown in step 603.

[0198] The processed direct portion and diffuse portion can then be combined to generate an output audio signal, as shown in step 607. Figure 7 The direct portion is then determined, as shown in step 603.

[0199] The combined output audio signal can then be output, as shown in step 617. Figure 7 The direct portion is then determined, as shown in step 603.

[0200] With regard to Figure 7 a flowchart showing the operation of an example diffuse portion gain determiner 605 in accordance with some embodiments is shown.

[0201] The example diffuse portion gain determiner 605 can be configured to receive / acquire input environment energy distribution parameters 104, e.g. the previously described ambianceAzi, ambianceEle and ambianceExtent parameters, as shown in step 701. Figure 7 The direct portion is then determined, as shown in step 603.

[0202] In some embodiments, the example diffuse portion gain determiner 605 can then be configured to determine a direction associated with a prototype output signal. In the case of loudspeaker synthesis, a prototype output signal is associated with the direction of each output loudspeaker. In the case of binaural synthesis, prototype output signals can be created with associated directions to fill the spatial audio field uniformly and / or at constant spacing.

[0203] The determination of where the direction associated with the prototype output is directed is shown in step 703. Figure 7 The diffuse portion gain determiner 605 can then determine, for each prototype output signal, whether the direction of the prototype signal (or virtual microphone) is within the receive sector of the environment energy distribution.

[0204] For example, for an environment energy distribution of (azimuth 0, elevation 0) and extent 90 degrees, spatial locations from (azimuth 45, elevation 0) to (azimuth -45, elevation 0) are within the environment energy distribution.

[0205] The determination of whether the direction of the prototype output signal is within the sector of the environment energy distribution is shown in step 705.

[0206] Figure 7 The direct portion is then determined, as shown in step 603.

[0207] ​Therefore, the diffuse gain determiner 605 can be configured to set a gain value of 1 for any prototype output signal within the distribution and a gain value of 0 for any prototype output signal outside the distribution. More generally, the diffuse gain determiner can be configured to set the gain associated with prototype output signals within the sector to be on average greater than the gain associated with virtual microphone signals outside the sector.

[0208] The gain value setting Figure 7 This is illustrated in step 707.

[0209] The sum of squared gains can then be normalized to a unit value, such as Figure 8 The process is shown in step 709.

[0210] These gains can then be passed to diffuse gain stage 519, which is configured to perform compositing using the acquired gains, such as... Figure 4 The process is shown in step 711.

[0211] Therefore, the effect of the above synthesis is to reduce environmental energy or to prevent environmental energy from being synthesized in a direction other than the distribution of received environmental energy.

[0212] If the environmental energy distribution parameters include the environmental energy ratio parameter, then the environmental energy is synthesized into different sectors with a suitable energy ratio.

[0213] In some embodiments, instead of converting to a universal format such as FOA for different spatial audio input formats, the spatial audio input is fed into spatial analysis, environmental energy distribution analysis, and transmission signal generation. This is in Figure 4 The spatial audio 800 is described in detail. The input spatial audio 800 can be a speaker input format, immersive audio (FOA or HOA), a multi-microphone format (i.e., the output signal of a microphone array), or a parameterized format already employing direction and ratio metadata analyzed by the spatial audio capture module. If the input is already in a parameterized format, the spatial analyzer 803 may perform no operation, or may simply perform a conversion from one parameterized representation to another. If the input is not in a parameterized format, the spatial analyzer 803 can be configured to perform spatial analysis to derive direction and ratio metadata. The ambient energy distribution analyzer 807 determines parameters representing the distribution of ambient energy. The determination of the parameters for the ambient energy distribution can differ for different input formats. In some cases, this determination can be based on analyzing the ambient energy at different input channels. It can be based on a signal with a one-sided directional pattern formed from the components of the input spatial audio. A signal with a one-sided directional pattern can be acquired through beamforming or any suitable means.

[0214] The synthesis described herein can also be integrated with a covariance matrix based synthesis. Covariance matrix based synthesis refers to a least-squares optimized signal mixing technique that operates on the covariance matrix of the signals while maintaining good audio quality. This synthesis utilizes the covariance matrix of the input signals and a target covariance matrix (determined by the desired output signal characteristics) and provides a mixing matrix to perform such processing.

[0215] Thus, the key information that needs to be determined is the mixing matrix in the frequency band, which is formulated based on the input and target covariance matrices in the frequency band. The input covariance matrix is measured from the input signals in the frequency band, and the target covariance matrix is formulated as the sum of the ambient part covariance matrix and the direct part covariance matrix. The diagonal entries of the ambient part covariance matrix are created such that the entries corresponding to the spatial directions inside the ambient distribution are set to unity values, and the other entries are set to zero. The diagonal entries are then normalized such that they sum to unity. In some embodiments, the energy within the sector is increased and the energy outside the sector is decreased, and then normalized such that they sum to unity.

[0216] Alternatively, a similar direction index for a spherical surface grid as defined for the directional information can be used to signal the direction of the center of the analyzed ambient energy distribution. For example, the index of the source direction can be obtained by forming a fixed grid of small spheres on a larger sphere and considering the centers of these small spheres as points of a grid that defines almost equidistant directions. The width or extent of the ambient energy distribution can be expressed in radians instead of degrees and quantized to a suitable resolution. Alternatively, the width or extent can be expressed as a number that indicates how many fixed-width spatial sectors it covers. For example, in the example of Figure 9 the value of ambianceExtent can be 3, indicating that it spans three 45-degree sectors. In some embodiments, the ambienceExtent information can include an additional parameter, ambianceExtentSector, that indicates the size of the analysis sector used for the ambient energy distribution analysis. Thus, in the example of ​ the value of ambianceAnalysisSectorWidth can be 45 degrees. Signaling the span of the ambient analysis sector enables the encoder to use different sizes of sectors for the ambient energy analysis. Adapting the size of the ambient energy analysis sector can be advantageous to adjust the system operation for soundfields with different ambient characteristics and to adjust the bandwidth and computational complexity requirements of the encoder and / or decoder.

[0217] With respect to ​FIG. 14 shows an example electronic device that can be used as an analysis or synthesis device. The device can be any suitable electronic device or apparatus. For example, in some embodiments, the device 1400 is a mobile device, user equipment, tablet computer, computer, audio playback apparatus, etc.

[0218] In some embodiments, the device 1400 includes at least one processor or central processing unit 1407. The processor 1407 can be configured to execute various program codes, such as the methods described herein.

[0219] In some embodiments, the device 1400 includes a memory 1411. In some embodiments, the at least one processor 1407 is coupled to the memory 1411. The memory 1411 can be any suitable storage module. In some embodiments, the memory 1411 includes a program code portion for storing program codes that can be implemented on the processor 1407. Further, in some embodiments, the memory 1411 can also include a storage data portion for storing data, such as data that has been processed or is to be processed according to the embodiments described herein. The implemented program codes stored within the program code portion and the data stored within the storage data portion can be retrieved by the processor 1407 through the memory-processor coupling as needed.

[0220] In some embodiments, the device 1400 includes a user interface 1405. In some embodiments, the user interface 1405 can be coupled to the processor 1407. In some embodiments, the processor 1407 can control operation of the user interface 1405 and receive input from the user interface 1405. In some embodiments, the user interface 1405 can enable a user to input commands to the device 1400, for example, via a keypad. In some embodiments, the user interface 1405 can enable a user to obtain information from the device 1400. For example, the user interface 1405 can include a display configured to display information from the device 1400 to a user. In some embodiments, the user interface 1405 can include a touchscreen or touch interface that can both enable information to be input to the device 1400 and display information to a user of the device 1400.

[0221] In some embodiments, the device 1400 includes an input / output port 1409. In some embodiments, the input / output port 1409 includes a transceiver. In such embodiments, the transceiver can be coupled to the processor 1407 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver module can be configured to communicate with other electronic devices or apparatuses via a wired or wired coupling.

[0222] The transceiver can communicate with the further apparatus by any appropriate known communication protocol. For example, in some embodiments the transceiver can use suitable Universal Mobile Telecommunications System (UMTS) protocols, wireless local area network (WLAN) protocols such as IEEE 802.X, suitable short range radio frequency communication protocols such as Bluetooth, or infrared data communication paths (IRDA).

[0223] The transceiver input / output port 1409 can be configured to receive signals and, in some embodiments, determine parameters as described herein by use of the processor 1407 executing suitable code. Furthermore, the apparatus can generate suitable transmission signals and parameter outputs for transmission to a synthesis apparatus.

[0224] In some embodiments, the apparatus 1400 can be used as at least part of a synthesis apparatus. As such, the input / output port 1409 can be configured to receive transmission signals and, in some embodiments, parameters determined at a capture apparatus or processing apparatus as described herein, and generate suitable audio signal format outputs by use of the processor 1407 executing suitable code. The input / output port 1409 can be coupled to any suitable audio output, for example to a multi-channel loudspeaker system and / or headphones or the like.

[0225] In general, the various embodiments of the application can be implemented in hardware or special-purpose circuits, software, logic or any combination thereof. For example, some aspects can be implemented in hardware, while other aspects can be implemented in

[0226] As used in this application, the term "circuitry" can refer to one or more or all of the following:

[0227] (a) an analog and / or digital hardware circuit implementation (for example, an implementation in only analog and / or digital circuitry), and

[0228] (b) a combination of hardware circuits and software, such as (as applicable):

[0229] (i) analog and / or digital hardware circuit(s) with software / firmware and

[0230] (ii) any portions of hardware processor-based systems (including digital signal processors) that require software (including digital signal processors) to achieve certain steps, such as mobile phones or servers, and

[0231] (c) hardware circuitry that is specially adapted for certain functionality, such as a microprocessor or a portion thereof, but requires software (e.g., firmware) to be loaded during an initialization process to program logic into the hardware circuitry, such as a so-called "cryptography processor", or a microprocessor that requires software to program logic into it during an initialization process; and / or

[0232] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term "circuitry" also covers an implementation that is a hardware circuit or processor (or multiple processors) or hardware circuit or processor and accompanying software and / or firmware that work together to cause a device to perform various functionality described herein, but the term circuitry does not cover a mere paper or other conveyance of software or firmware.

[0233] Embodiments of the application can be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Further in this regard, it should be noted that any blocks of the logic flow of the Figures can represent program steps, or interconnected logic circuit, block, and functions, or combinations of program steps and logic circuit, block, and functions. The software can be stored on such physical media as memory chips, or memory blocks implemented in the processor, or hard disk or floppy disks, and can come from suppliers, storage media or a communication source such as the Internet.

[0234] The memory can be of any type appropriate for the local technical environment and can be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processor can be of any type appropriate for the local technical environment, and can include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), gate level circuits and processors based on multi core processor architectures, as non-limiting examples.

[0235] Embodiments of the application can be practiced in a variety of system environments, such as an integrated circuit module. Design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available for conversion of a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0236] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using the geometries of the conductors and the characteristics of the components provided in predetermined libraries. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) can be transmitted to a semiconductor fabrication facility or "fab" for fabrication.

[0237] The foregoing description, for the purpose of clarity, provides a detailed description of the exemplary embodiments of the application. However, it will be apparent to those skilled in the art that many modifications and adaptations can be made in light of the foregoing description and appended claims. All of these and similar modifications and adaptations are within the scope of the application, which is defined solely by the appended claims.

Claims

1. An apparatus for spatial audio signal processing, the apparatus comprising: at least one processor and at least one non-transitory memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to: receive at least two audio signals; determine at least one parameter associated with the at least two audio signals, wherein the at least one parameter is configured to be indicative of an ambient energy distribution of the at least two audio signals, wherein the at least one parameter is associated with at least respective ambient sound energies of the at least two audio signals in a plurality of directions; determine at least one directional parameter representing directional information of the at least two audio signals; and provide at least one output audio signal using at least one of the at least two audio signals based at least on the at least one directional parameter and the at least one parameter, wherein an ambient energy distribution of the at least one output signal is controlled based on the at least one parameter from the ambient energy distribution of the at least two audio signals in the plurality of directions. the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to:

2. The apparatus of claim 1, wherein, determine spatial metadata associated with the at least two audio signals, wherein the determined spatial metadata comprises the at least one parameter and the at least one directional parameter. the at least two audio signals comprise at least one of:

3. The apparatus of claim 1, wherein, transport audio signals; microphone array signals; multi-channel microphone signals; higher order ambisonic signals; first order ambisonic signals, spatial audio signals, or signals generated based on a plurality of spatial audio signals. the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to:

4. The apparatus of claim 1, wherein, divide at least one of the at least two audio signals into a direct portion and a diffuse portion based at least in part on at least one of the at least one directional parameter and the at least one parameter; synthesize a direct audio signal based on the direct portion and the at least one directional parameter; determine a diffuse portion gain based on the at least one parameter; synthesize a diffuse audio signal based on the diffuse portion and the diffuse portion gain; and combine the direct audio signal and the diffuse audio signal to generate the at least one output audio signal. the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to: determine a direction to which a set of prototype output signals points; 5. The apparatus of claim 4, wherein, for each prototype output signal of the set of prototype output signals, determine whether a direction of the each prototype output signal is within a sector defined by the at least one parameter configured to be indicative of an ambient energy distribution of the at least two audio signals; and set a gain associated with prototype output signals within the sector to be on average greater than a gain associated with prototype output signals outside the sector. the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to: ​ 6. The apparatus of claim 1, wherein, ​ analyzing the at least two audio signals to determine the at least one parameter configured to indicate an ambient energy distribution of the at least two audio signals; and receiving the at least one parameter configured to indicate an ambient energy distribution of the at least two audio signals.

7. The apparatus of claim 1, wherein, The at least one directional parameter comprises at least one of: at least one direction parameter representing a direction of arrival; a diffuse parameter associated with the at least one direction parameter; or an energy ratio parameter associated with the at least one direction parameter.

8. The apparatus of claim 1, wherein, The at least one parameter comprises at least one of: a first parameter comprising at least one azimuth and / or at least one elevation associated with at least one spatial sector having a locally maximum average ambient energy; or at least one other parameter based on an angular extent of the at least one spatial sector having the locally maximum average ambient energy.

9. A method comprising: receiving at least two audio signals; determining at least one parameter associated with the at least two audio signals, wherein the at least one parameter is configured to indicate an ambient energy distribution of the at least two audio signals, wherein the at least one parameter is associated with at least respective ambient sound energies of the at least two audio signals in a plurality of directions; determining at least one directional parameter representing directional information of the at least two audio signals; and using at least one of the at least two audio signals to provide at least one output audio signal based at least on the at least one directional parameter and the at least one parameter, wherein an ambient energy distribution of the at least one output signal is controlled based on the at least one parameter from ambient energy distributions of the at least two audio signals in the plurality of directions.

10. The method according to claim 9, further comprising: determining spatial metadata associated with the at least two audio signals, wherein the determined spatial metadata comprises the at least one parameter and the at least one directional parameter.

11. The method of claim 9, wherein, The at least two audio signals comprise at least one of: transport audio signals; microphone array signals; multi-channel microphone signals; higher order ambisonic signals; first order ambisonic signals, spatial audio signals, or signals generated based on a plurality of spatial audio signals.

12. The method according to claim 9, further comprising: partitioning at least one of the at least two audio signals into a direct portion and a diffuse portion based at least in part on at least one of the at least one directional parameter and the at least one parameter; synthesizing a direct audio signal based on the direct portion and the at least one directional parameter; determining a diffuse portion gain based on the at least one parameter; synthesizing a diffuse audio signal based on the diffuse portion and the diffuse portion gain; and combining the direct audio signal and the diffuse audio signal to generate the at least one output audio signal.

13. The method according to claim 12, further comprising: determining a direction to which a prototype output signal set is directed; for each prototype output signal of the set of prototype output signals, determining whether a direction of the respective prototype output signal is within a sector defined by at least one parameter configured to indicate an ambient energy distribution of the at least two audio signals; and setting a gain associated with the prototype output signals within the sector to be on average greater than a gain associated with the prototype output signals outside the sector.

14. The method of claim 9, further comprising: analyzing the at least two audio signals to determine the at least one parameter configured to indicate an ambient energy distribution of the at least two audio signals; and receiving the at least one parameter configured to indicate an ambient energy distribution of the at least two audio signals.

15. The method of claim 9, wherein, The at least one directional parameter comprises at least one of: at least one directional parameter representing a direction of arrival; a diffuse parameter associated with the at least one directional parameter; or an energy ratio parameter associated with the at least one directional parameter.

16. The method of claim 9, wherein, The at least one parameter comprises at least one of: a first parameter comprising at least one azimuth and / or at least one elevation associated with at least one spatial sector having a locally maximum average ambient energy; or at least one other parameter based on an angular extent of the at least one spatial sector having the locally maximum average ambient energy.

17. A non-transitory computer readable medium comprising program instructions stored thereon that, when executed with at least one processor, cause the at least one processor to: receive at least two audio signals; determine at least one parameter associated with the at least two audio signals, wherein the at least one parameter is configured to indicate an ambient energy distribution of the at least two audio signals, wherein the at least one parameter is associated with at least respective ambient sound energies of the at least two audio signals in a plurality of directions; determine at least one directional parameter representing directional information of the at least two audio signals; and cause provision of at least one output audio signal using at least one of the at least two audio signals based on at least the at least one directional parameter and the at least one parameter, wherein an ambient energy distribution of the at least one output signal is controlled based on the at least one parameter from ambient energy distributions of the at least two audio signals in the plurality of directions.

18. The non-transitory computer readable medium of claim 17, further comprising program instructions stored thereon that, when executed with at least one processor, cause the at least one processor to: determine spatial metadata associated with the at least two audio signals, wherein the determined spatial metadata comprises the at least one parameter and the at least one directional parameter.

19. The non-transitory computer-readable medium of claim 17, wherein, The at least two audio signals comprise at least one of: transport audio signals; microphone array signals; multi-channel microphone signals; higher order ambisonic signals; first order ambisonic signals, spatial audio signals, or signals generated based on a plurality of spatial audio signals.

20. The non-transitory computer-readable medium of claim 17, wherein, The at least one parameter comprises at least one of: a first parameter comprising at least one azimuth and / or at least one elevation angle associated with at least one spatial sector having a local maximum average ambient energy; or at least one other parameter based on an angular extent of the at least one spatial sector having the local maximum average ambient energy.

Citation Information

Patent Citations

  • Apparatus and method for generating a plurality of parametric audio streams and apparatus and method for generating a plurality of loudspeaker signals

    EP2733965A1