Spatial audio parameters and associated spatial audio playback

By estimating and reproducing sound field-related parameters in the frequency band using processor and memory in the device, the problem of low quality of sound field parameter estimation and reproduction in the prior art is solved, and a higher quality spatial audio capture and reproduction is achieved.

CN119943061APending Publication Date: 2025-05-06NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510099416.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-04-06
Filing Date
2019-03-28
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively estimate and reproduce the sound field-related parameters in the frequency band, affecting the capture and reproduction quality of spatial audio.

Method used

By means, the device comprises a processor and a memory configured to determine at least one spatial audio parameter for providing spatial audio reproduction for two or more microphone audio signals, and determine at least one coherent parameter associated with the sound field based on these signals so that the sound field can be reproduced.

Benefits of technology

More accurate sound field parameter estimation and reproduction are achieved, the quality of spatial audio capture and reproduction is improved, and the multi-band sound field can be effectively processed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943061A_ABST
    Figure CN119943061A_ABST
Patent Text Reader

Abstract

The invention relates to spatial audio parameters and associated spatial audio playback. An apparatus (105) includes at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus (105) to at least: for two or more microphone audio signals (102), transmit the two or more microphone audio signals (102); determining at least one spatial audio parameter for providing spatial audio reproduction (108, 304); based on the two or more microphone audio signals (102), at least one coherence parameter (112, 114) associated with a sound field is determined such that another sound field is configured to be reproduced based on the at least one spatial audio parameter (108, 304) and the at least one coherence parameter (112, 114).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese invention patent application with the invention name “Spatial Audio Parameters and Associated Spatial Audio Playback” (application number 201980037198.1, application date March 28, 2019). Technical Field

[0002] The present application relates to an apparatus and method for sound field related parameter estimation in a frequency band, but not exclusively for time-frequency domain sound field related parameter estimation of an audio encoder and decoder. Background Art

[0003] Parametric spatial audio processing is a field of audio signal processing in which the spatial aspects of sound are described using a set of parameters. For example, in parametric spatial audio capture from a microphone array, estimating a set of parameters from the microphone array signal (e.g., the direction of the sound in a frequency band and the ratio between the directional and non-directional parts of the captured sound in the frequency band) is a typical and effective choice. It is well known that these parameters describe the perceived spatial characteristics of the captured sound at the location of the microphone array very well. These parameters can be used accordingly for the synthesis of spatial sound for binaural headphones, for loudspeakers, or in other formats such as Ambisonics.

[0004] Therefore, the directional and direct-to-total energy ratios in frequency bands are particularly effective parameterizations for spatial audio capture. Summary of the invention

[0005] According to a first aspect, a device is provided, which includes at least one processor and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to, together with the at least one processor, enable the device to at least: determine at least one spatial audio parameter for providing spatial audio reproduction for two or more microphone audio signals; determine at least one coherence parameter associated with a sound field based on the two or more microphone audio signals, so that the sound field is configured to be reproduced based on the at least one spatial audio parameter and the at least one coherence parameter.

[0006] According to another aspect, a device is provided, comprising at least one processor and at least one memory comprising computer program code, wherein the at least one memory and the computer program code are configured to, together with the at least one processor, enable the device to at least: determine at least one spatial audio parameter for providing spatial audio reproduction for two or more microphone audio signals; determine at least one coherence parameter based on the determination of coherence within a sound field based on the two or more microphone audio signals, so that the sound field is configured to be reproduced based on the at least one spatial audio parameter and the at least one coherence parameter.

[0007] The device caused to determine at least one coherence parameter associated with the sound field based on the two or more microphone audio signals may be further caused to determine at least one of: at least one extended coherence parameter, the at least one extended coherence parameter being associated with the coherence of the directional portion of the sound field; and at least one surround coherence parameter, the at least one surround coherence parameter being associated with the coherence of the non-directional portion of the sound field.

[0008] The device caused to determine at least one spatial audio parameter for providing spatial audio reproduction for two or more microphone audio signals may be further caused to: determine at least one of the following for the two or more microphone audio signals: a directional parameter; an energy ratio parameter; a direct-to-overall energy parameter; a directional stability parameter; an energy parameter.

[0009] The apparatus may be further caused to determine an associated audio signal based on the two or more microphone audio signals, wherein the sound field may be reproduced based on the at least one spatial audio parameter, the at least one coherence parameter and the associated audio signal.

[0010] The device caused to determine at least one coherence parameter associated with the sound field based on the two or more microphone audio signals may be further caused to: determine zero-order and first-order spherical harmonics based on the two or more microphone audio signals; generate at least one universal coherence parameter based on the zero-order and first-order spherical harmonics; and generate the at least one coherence parameter based on the at least one universal coherence parameter.

[0011] The device caused to determine zero-order and first-order spherical harmonics based on the two or more microphone audio signals may be further caused to perform one of the following: determining time-domain zero-order and first-order spherical harmonics based on the two or more microphone audio signals and converting the time-domain zero-order and first-order spherical harmonic functions into time-frequency domain zero-order and first-order spherical harmonics; and converting the two or more microphone audio signals into respective two or more time-frequency domain microphone audio signals, and generating time-frequency domain zero-order and first-order spherical harmonics based on the time-frequency domain microphone audio signals.

[0012] The device caused to generate the at least one coherence parameter based on the at least one general coherence parameter may be caused to: generate at least one extended coherence parameter based on the at least one general coherence parameter and an energy ratio configured to define the relationship between the direct part and the ambient part of the sound field; generate at least one surround coherence parameter based on the at least one general coherence parameter and an energy ratio configured to define the relationship between the direct part and the ambient part of the sound field.

[0013] The device caused to determine at least one coherence parameter associated with the sound field based on the two or more microphone audio signals may be further caused to: convert the two or more microphone audio signals into corresponding two or more time-frequency domain microphone audio signals; determine at least one estimate of non-reverberant sound based on the two or more time-frequency domain microphone audio signals; determine at least one surround coherence parameter based on the at least one estimate of the non-reverberant sound and an energy ratio configured to define a relationship between a direct part and an ambient part of the generated sound field.

[0014] The device caused to determine at least one coherence parameter associated with the sound field based on the two or more microphone audio signals may be further caused to select at least one of: at least one surround coherence parameter based on at least one estimate of the non-reverberant sound and an energy ratio, and at least one surround coherence parameter based on the at least one such general coherence parameter, the surround coherence parameter being maximized based on the at least one such general coherence parameter.

[0015] The device caused to determine at least one coherence parameter associated with the sound field based on the two or more microphone audio signals may be further caused to determine at least one coherence parameter associated with the sound field based on the two or more microphone audio signals and for two or more frequency bands.

[0016] According to a second aspect, a device is provided, comprising at least one processor and at least one memory comprising computer program code, wherein the at least one memory and the computer program code are configured to, together with the at least one processor, enable the device to at least: receive at least one audio signal, the at least one audio signal being based on two or more microphone audio signals; receive at least one coherence parameter associated with a sound field based on the two or more microphone audio signals; receive at least one spatial audio parameter for providing spatial audio reproduction; and reproduce the sound field based on the at least one audio signal, the at least one spatial audio parameter and the at least one coherence parameter.

[0017] The device caused to receive at least one coherence parameter may be further caused to receive at least one of: at least one extended coherence parameter for the at least two frequency bands, the at least one extended coherence parameter being associated with the coherence of the directional portion of the sound field; and at least one surround coherence parameter being associated with the coherence of the non-directional portion of the sound field.

[0018] The at least one spatial audio parameter may include at least one of: a direction parameter; an energy ratio parameter; a direct-to-overall energy parameter; a directional stability parameter; and an energy parameter, and the device caused to reproduce the sound field based on the at least one audio signal, the at least one spatial audio parameter and the at least one coherence parameter may be further caused to: determine a target covariance matrix from the at least one spatial audio parameter, the at least one coherence parameter and the estimated energy of the at least one audio signal; generate a mixing matrix based on the target covariance matrix and the estimated energy of the at least one audio signal; and apply the mixing matrix to the at least one audio signal to generate at least two output spatial audio signals to reproduce the sound field.

[0019] The device caused to determine a target covariance matrix from the at least one spatial audio parameter, the at least one coherence parameter and the energy of the at least one audio signal may be further caused to: determine a total energy parameter based on the energy of the at least one audio signal; determine direct energy and ambient energy based on the energy ratio parameter, the direct to total energy parameter, the directional stability parameter and at least one of the energy parameters; estimate an ambient covariance matrix based on the determined ambient energy and one of the at least one correlation parameter; estimate at least one of the following based on the output channel configuration and / or the at least one directional parameter: a vector of amplitude panning gains, an ambient panning vector or at least one head-related transfer function; estimate a direct covariance matrix based on: the vector of amplitude panning gains, the ambient panning vector or the at least one head-related transfer function; the determined direct partial energy; another of the at least one coherence parameter; and generate the target covariance matrix by combining the ambient covariance matrix and the direct covariance matrix.

[0020] According to a third aspect, a method is provided, comprising: determining, for two or more microphone audio signals, at least one spatial audio parameter for providing spatial audio reproduction; and determining, based on the two or more microphone audio signals, at least one coherence parameter associated with a sound field, so that the sound field is configured to be reproduced based on the at least one spatial audio parameter and the at least one coherence parameter.

[0021] Determining at least one coherence parameter associated with the sound field based on the two or more microphone audio signals may further include determining at least one of: at least one extended coherence parameter, the at least one extended coherence parameter being associated with the coherence of a directional portion of the sound field; and at least one surround coherence parameter, the at least one surround coherence parameter being associated with the coherence of a non-directional portion of the sound field.

[0022] Determining at least one spatial audio parameter for providing spatial audio reproduction for two or more microphone audio signals may also include: determining at least one of the following for the two or more microphone audio signals: a direction parameter; an energy ratio parameter; a direct-to-overall energy parameter; a directional stability parameter; an energy parameter.

[0023] The method may further comprise determining an associated audio signal based on the two or more microphone audio signals, wherein the sound field may be reproduced based on the at least one spatial audio parameter, the at least one coherence parameter and the associated audio signal.

[0024] Determining at least one coherence parameter associated with the sound field based on the two or more microphone audio signals may further include: determining zero-order and first-order spherical harmonics based on the two or more microphone audio signals; generating at least one universal coherence parameter based on the zero-order and first-order spherical harmonics; and generating the at least one coherence parameter based on the at least one universal coherence parameter.

[0025] Determining zero-order and first-order spherical harmonics based on the two or more microphone audio signals may also include one of the following: determining time-domain zero-order and first-order spherical harmonics based on the two or more microphone audio signals and converting the time-domain zero-order and first-order spherical harmonic functions into time-frequency domain zero-order and first-order spherical harmonics; and converting the two or more microphone audio signals into respective two or more time-frequency domain microphone audio signals, and generating time-frequency domain zero-order and first-order spherical harmonics based on the time-frequency domain microphone audio signals.

[0026] Generating the at least one coherence parameter based on the at least one general coherence parameter may further include: generating at least one extended coherence parameter based on the at least one general coherence parameter and an energy ratio configured to define the relationship between the direct part and the ambient part of the sound field; generating at least one surround coherence parameter based on the at least one general coherence parameter and an energy ratio configured to define the relationship between the direct part and the ambient part of the sound field.

[0027] Determining at least one coherence parameter associated with the sound field based on the two or more microphone audio signals may further include: converting the two or more microphone audio signals into corresponding two or more time-frequency domain microphone audio signals; determining at least one estimate of non-reverberant sound based on the two or more time-frequency domain microphone audio signals; determining at least one surround coherence parameter based on the at least one estimate of the non-reverberant sound and an energy ratio configured to define a relationship between a direct portion and an ambient portion of the generated sound field.

[0028] Determining at least one coherence parameter associated with the sound field based on the two or more microphone audio signals may further include selecting at least one of: at least one surround coherence parameter based on at least one estimate of the non-reverberant sound and an energy ratio, and at least one surround coherence parameter based on at least one such general coherence parameter, the surround coherence parameter being maximal based on at least one such general coherence parameter.

[0029] Determining at least one coherence parameter associated with the sound field based on the two or more microphone audio signals may further comprise determining at least one coherence parameter associated with the sound field based on the two or more microphone audio signals and for two or more frequency bands.

[0030] According to a fourth aspect, a method is provided, comprising: receiving at least one audio signal, the at least one audio signal being based on two or more microphone audio signals; receiving at least one coherence parameter associated with a sound field based on the two or more microphone audio signals; receiving at least one spatial audio parameter for providing spatial audio reproduction; and reproducing the sound field based on the at least one audio signal, the at least one spatial audio parameter and the at least one coherence parameter.

[0031] Receiving at least one coherence parameter may further include receiving at least one of: at least one extended coherence parameter for the at least two frequency bands, the at least one extended coherence parameter being associated with the coherence of the directional portion of the sound field; and at least one surround coherence parameter, the at least one surround coherence parameter being associated with the coherence of the non-directional portion of the sound field.

[0032] The at least one spatial audio parameter may include at least one of: a direction parameter; an energy ratio parameter; a direct-to-overall energy parameter; a directional stability parameter; and an energy parameter, and reproducing the sound field based on the at least one audio signal, the at least one spatial audio parameter and the at least one coherence parameter may also include: determining a target covariance matrix from the at least one spatial audio parameter, the at least one coherence parameter and the estimated energy of the at least one audio signal; generating a mixing matrix based on the target covariance matrix and the estimated energy of the at least one audio signal; and applying the mixing matrix to the at least one audio signal to generate at least two output spatial audio signals for reproducing the sound field.

[0033] Determining a target covariance matrix from the at least one spatial audio parameter, the at least one coherence parameter and the estimated energy of the at least one audio signal may further include: determining a total energy parameter based on the energy of the at least one audio signal; determining direct energy and ambient energy based on the energy ratio parameter, the direct to total energy parameter, the directional stability parameter and at least one of the energy parameters; estimating an ambient covariance matrix based on the determined ambient energy and one of the at least one correlation parameter; estimating at least one of the following: a vector of amplitude panning gains, an ambient sound panning vector or at least one head-related transfer function based on the output channel configuration and / or the at least one directional parameter; estimating a direct covariance matrix based on: the vector of amplitude panning gains, the ambient sound panning vector or the at least one head-related transfer function; the determined direct partial energy; and another of the at least one coherence parameter; and generating the target covariance matrix by combining the ambient covariance matrix and the direct covariance matrix.

[0034] According to a fifth aspect, a device is provided, comprising a module for performing the following operations: determining, for two or more microphone audio signals, at least one spatial audio parameter for providing spatial audio reproduction; determining, based on the two or more microphone audio signals, at least one coherence parameter associated with a sound field, so that the sound field is configured to be reproduced based on the at least one spatial audio parameter and the at least one coherence parameter.

[0035] The module for determining at least one coherence parameter associated with a sound field based on the two or more microphone audio signals may be further configured to determine at least one of: at least one extended coherence parameter associated with the coherence of a directional portion of the sound field; and at least one surround coherence parameter associated with the coherence of a non-directional portion of the sound field.

[0036] The module for determining at least one spatial audio parameter for providing spatial audio reproduction for two or more microphone audio signals may be further configured to determine at least one of the following for the two or more microphone audio signals: a direction parameter; an energy ratio parameter; a direct-to-overall energy parameter; a directional stability parameter; an energy parameter.

[0037] The module may be further configured for determining an associated audio signal based on the two or more microphone audio signals, wherein the sound field may be reproduced based on the at least one spatial audio parameter, the at least one coherence parameter and the associated audio signal.

[0038] The module for determining at least one coherence parameter associated with a sound field based on the two or more microphone audio signals may be further configured to: determine zero-order and first-order spherical harmonics based on the two or more microphone audio signals; generate at least one universal coherence parameter based on the zero-order and first-order spherical harmonics; and generate the at least one coherence parameter based on the at least one universal coherence parameter.

[0039] The module for determining zero-order and first-order spherical harmonics based on the two or more microphone audio signals can be further configured to perform one of the following: determining time-domain zero-order and first-order spherical harmonics based on the two or more microphone audio signals and converting the time-domain zero-order and first-order spherical harmonic functions into time-frequency domain zero-order and first-order spherical harmonics; and converting the two or more microphone audio signals into two or more respective time-frequency domain microphone audio signals, and generating time-frequency domain zero-order and first-order spherical harmonics based on the time-frequency domain microphone audio signals.

[0040] The module for generating the at least one coherence parameter based on the at least one general coherence parameter may be configured to: generate at least one extended coherence parameter based on the at least one general coherence parameter and an energy ratio configured to define a relationship between a direct part and an ambient part of the sound field; generate at least one surround coherence parameter based on the at least one general coherence parameter and an energy ratio configured to define a relationship between a direct part and an ambient part of the sound field.

[0041] The module for determining at least one coherence parameter associated with a sound field based on the two or more microphone audio signals may be further configured to: convert the two or more microphone audio signals into corresponding two or more time-frequency domain microphone audio signals; determine at least one estimate of non-reverberant sound based on the two or more time-frequency domain microphone audio signals; determine at least one surround coherence parameter based on the at least one estimate of the non-reverberant sound and an energy ratio configured to define a relationship between a direct portion and an ambient portion of the generated sound field.

[0042] The means for determining at least one coherence parameter associated with the sound field based on the two or more microphone audio signals may be further configured to select at least one of: at least one surround coherence parameter based on at least one estimate of the non-reverberant sound and an energy ratio, and the at least one surround coherence parameter based on at least one such general coherence parameter, and a surround coherence parameter that is maximum based on the at least one such general coherence parameter.

[0043] The module for determining at least one coherence parameter associated with the sound field based on the two or more microphone audio signals may be further configured for determining at least one coherence parameter associated with the sound field based on the two or more microphone audio signals and for two or more frequency bands.

[0044] According to a sixth aspect, a device is provided, comprising a module for performing the following operations: receiving at least one audio signal, the at least one audio signal being based on two or more microphone audio signals; receiving at least one coherence parameter associated with a sound field based on the two or more microphone audio signals; receiving at least one spatial audio parameter for providing spatial audio reproduction; and reproducing the sound field based on the at least one audio signal, the at least one spatial audio parameter and the at least one coherence parameter.

[0045] The module for receiving at least one coherence parameter may be further configured to receive at least one of: at least one extended coherence parameter for the at least two frequency bands, the at least one extended coherence parameter being associated with the coherence of the directional portion of the sound field; and at least one surround coherence parameter being associated with the coherence of the non-directional portion of the sound field.

[0046] The at least one spatial audio parameter may include at least one of: a direction parameter; an energy ratio parameter; a direct-to-overall energy parameter; a directional stability parameter; and an energy parameter, and the module for reproducing the sound field based on the at least one audio signal, the at least one spatial audio parameter and the at least one coherence parameter may be further configured to: determine a target covariance matrix from the at least one spatial audio parameter, the at least one coherence parameter and the estimated energy of the at least one audio signal; generate a mixing matrix based on the target covariance matrix and the estimated energy of the at least one audio signal; and apply the mixing matrix to the at least one audio signal to generate at least two output spatial audio signals to reproduce the sound field.

[0047] The module for determining a target covariance matrix from the at least one spatial audio parameter, the at least one coherence parameter and the estimated energy of the at least one audio signal may be further configured to: determine a total energy parameter based on the energy of the at least one audio signal; determine direct energy and ambient energy based on the energy ratio parameter, the direct to total energy parameter, the directional stability parameter and at least one of the energy parameters; estimate an ambient covariance matrix based on the determined ambient energy and one of the at least one correlation parameter; estimate at least one of the following based on the output channel configuration and / or the at least one directional parameter: a vector of amplitude panning gains, an ambient panning vector or at least one head-related transfer function; estimate a direct covariance matrix based on: the vector of amplitude panning gains, the ambient panning vector or the at least one head-related transfer function; the determined direct partial energy; and another of the at least one coherence parameter; and generate the target covariance matrix by combining the ambient covariance matrix and the direct covariance matrix.

[0048] According to the seventh aspect, a computer program comprising instructions [or a computer-readable medium comprising program instructions] is provided, for causing an apparatus to perform at least the following operations: for two or more microphone audio signals, determining at least one spatial audio parameter for providing spatial audio reproduction; determining at least one coherence parameter associated with a sound field based on the two or more microphone audio signals, so that the sound field is configured to be reproduced based on the at least one spatial audio parameter and the at least one coherence parameter.

[0049] According to an eighth aspect, there is provided a computer program comprising instructions [or a computer-readable medium comprising program instructions] for causing a device to perform at least the following operations: receiving at least one audio signal, the at least one audio signal being based on two or more microphone audio signals; receiving at least one coherence parameter associated with a sound field based on the two or more microphone audio signals; receiving at least one spatial audio parameter for providing spatial audio reproduction; and reproducing the sound field based on the at least one audio signal, the at least one spatial audio parameter and the at least one coherence parameter.

[0050] According to a ninth aspect, a non-temporary computer-readable medium is provided, comprising program instructions for causing a device to perform at least the following operations: determining, for two or more microphone audio signals, at least one spatial audio parameter for providing spatial audio reproduction; determining, based on the two or more microphone audio signals, at least one coherence parameter associated with a sound field, so that the sound field is configured to be reproduced based on the at least one spatial audio parameter and the at least one coherence parameter.

[0051] According to the tenth aspect, a non-temporary computer-readable medium is provided, comprising program instructions for causing a device to perform at least the following operations: receiving at least one audio signal, the at least one audio signal being based on two or more microphone audio signals; receiving at least one coherence parameter associated with a sound field based on the two or more microphone audio signals; receiving at least one spatial audio parameter for providing spatial audio reproduction; and reproducing the sound field based on the at least one audio signal, the at least one spatial audio parameter and the at least one coherence parameter.

[0052] According to the eleventh aspect, a device is provided, comprising: a determination circuit configured to determine at least one spatial audio parameter for providing spatial audio reproduction for two or more microphone audio signals; and the determination circuit is also configured to determine at least one coherence parameter associated with a sound field based on the two or more microphone audio signals, so that the sound field is configured to be reproduced based on the at least one spatial audio parameter and the at least one coherence parameter.

[0053] According to the twelfth aspect, a device is provided, comprising: a receiving circuit configured to receive at least one audio signal, wherein the at least one audio signal is based on two or more microphone audio signals; the receiving circuit is also configured to receive at least one coherence parameter associated with a sound field based on the two or more microphone audio signals; the receiving circuit is also configured to receive at least one spatial audio parameter for providing spatial audio reproduction; and a reproducing circuit configured to reproduce the sound field based on the at least one audio signal, the at least one spatial audio parameter and the at least one coherence parameter.

[0054] According to the thirteenth aspect, a computer-readable medium is provided, comprising program instructions for causing a device to perform at least the following operations: determining, for two or more microphone audio signals, at least one spatial audio parameter for providing spatial audio reproduction; determining, based on the two or more microphone audio signals, at least one coherence parameter associated with a sound field, so that the sound field is configured to be reproduced based on the at least one spatial audio parameter and the at least one coherence parameter.

[0055] According to the fourteenth aspect, a computer-readable medium is provided, comprising program instructions for causing a device to perform at least the following operations: receiving at least one audio signal, the at least one audio signal being based on two or more microphone audio signals; receiving at least one coherence parameter associated with a sound field based on the two or more microphone audio signals; receiving at least one spatial audio parameter for providing spatial audio reproduction; and reproducing the sound field based on the at least one audio signal, the at least one spatial audio parameter and the at least one coherence parameter.

[0056] An apparatus comprises modules for performing the actions of the method as described above.

[0057] An apparatus is configured to perform the actions of the method as described above.

[0058] A computer program comprises program instructions for causing a computer to execute the method as described above.

[0059] A computer program product stored on a medium may cause an apparatus to perform the method described herein.

[0060] An electronic device may include an apparatus as described herein.

[0061] A chipset may include the apparatus as described herein.

[0062] The embodiments of the present application are intended to solve the problems associated with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] For a better understanding of the present application, reference will now be made by way of example to the accompanying drawings, in which:

[0064] Figure 1 A system of apparatus suitable for implementing some embodiments is schematically shown;

[0065] Figure 2 According to some embodiments, Figure 1 A flowchart of the operation of the system shown;

[0066] Figure 3 Schematically illustrates how Figure 1 The analysis processor shown;

[0067] Figure 4 According to some embodiments, Figure 3 A flowchart of the operation of the analysis processor is shown;

[0068] Figure 5 An exemplary coherence analyzer according to some embodiments is shown;

[0069] Figure 6 According to some embodiments, Figure 5 A flow chart illustrating the operation of an exemplary coherence analyzer is shown;

[0070] Figure 7 Another exemplary coherence analyzer according to some embodiments is shown;

[0071] Figure 8 According to some embodiments, Figure 7 A flowchart of the operation of another exemplary coherence analyzer is shown;

[0072] Fig. 9 According to some embodiments, Figure 1 An example synthesis processor is shown;

[0073] Fig.10 According to some embodiments, Fig. 9 A flowchart of the operation of an exemplary synthesis processor is shown;

[0074] Fig.11 According to some embodiments, Fig.10 A flowchart of the operation of generating a target covariance matrix is ​​shown; and

[0075] Fig.12 An example device suitable for implementing the means described herein is schematically illustrated. DETAILED DESCRIPTION

[0076] Suitable apparatus and possible mechanisms for providing efficient spatial analysis-derived metadata parameters for microphone array input format audio signals are described in further detail below.

[0077] The concept expressed in the embodiments below is a system in which the reproduced sound scene is very similar to the original input sound scene, and avoids the reproduction of surround coherent (close, pressurized) sounds as distant ambient and avoids the reproduction of amplitude-shifted sounds as point sources.

[0078] Additionally, some embodiments enable microphone arrays to be virtual collections of microphone beam patterns. For example, a first-order surround sound (FOA) “capture” of a set of speaker and / or audio object signals. Virtual microphones may enable them to be virtual collections of microphone beam patterns.

[0079] Such a system including a real or virtual microphone array as in the embodiments described herein is able to produce an effective representation of the sound scene and provide high-quality spatial audio capture performance so that the perception of the reproduced audio matches the perception of the original sound field (e.g., surround coherent sounds are reproduced as surround coherent sounds, and extended coherent sounds are reproduced as extended coherent sounds).

[0080] Furthermore, some embodiments described herein may be able to recognize when audio is being captured in an anechoic (or at least dry) space and produce an effective representation of such a sound scene. Furthermore, the synthesis stage of some embodiments may include a suitable receiver or decoder that can attempt to recreate the perception of the sound field based on the analyzed parameters and the obtained transmitted audio signal (e.g., reproduce an anechoic sound field in a way that is perceived to be anechoic). This may include processing certain parts of the audio without decorrelation to avoid artifacts.

[0081] Reproducing sound coherently and simultaneously from multiple directions generates a perception that is different from that generated by a single speaker. For example, if the sound is reproduced coherently using the left front and right speakers, the sound may be perceived as more "airy" than if it is reproduced using only the center speaker. Correspondingly, if the sound is reproduced coherently from the left front, right, and center speakers, the sound may be described as close or pressurized. Therefore, spatially coherent sound reproduction is used for artistic purposes, such as adding the presence of certain sounds (e.g., the voice of a lead singer). Coherent reproduction from several speakers is also sometimes used to emphasize low-frequency content.

[0082] The concept discussed in further detail below is to provide methods and modules for performing the following operations: determining spatial coherence by adding a specific analysis method to the microphone array audio input and providing the added related (at least one coherence) parameter in the metadata stream, which can be provided together with other spatial metadata. In the present disclosure, the microphone audio signal can be a real microphone audio signal captured by a physical microphone, such as from a microphone array. Also in some embodiments, the microphone audio signal can be a virtual microphone audio signal generated, for example, synthetically. In some embodiments, the virtual microphone can be determined to have a directional capture pattern corresponding to an panoramic sound beam pattern (e.g., a FOA beam pattern).

[0083] Thus, the concepts discussed in further detail through example implementations relate to audio encoding and decoding using spatial audio or sound field related parameterizations (e.g., other spatial metadata parameters may include direction, energy ratio, direct to overall energy ratio, directional stability, or other suitable parameters). The concept also discloses a method and apparatus provided to improve the reproduction quality of an audio signal encoded with the above parameterizations. The concept embodiments improve the reproduction quality of a microphone audio signal by analyzing an input audio signal and determining at least one coherence parameter. The term coherence or cross-correlation is not strictly interpreted here as a specific similarity value between signals, such as a normalized square value, but generally reflects a similarity value between playback audio signals and may be complex (with phase), absolute, normalized, or squared values. A coherence parameter may be more generally represented as an audio signal relationship parameter that indicates the similarity of audio signals in any manner.

[0084] The coherence of the output signal may refer to the coherence of the reproduced loudspeaker signal or the reproduced binaural signal or the reproduced panoramic sound signal.

[0085] The coherence parameter may also be referred to as a non-reverberant sound parameter in some embodiments, because in some embodiments the coherence parameter is determined based on a non-reverberant estimator, which is caused to estimate a portion of the non-reverberant sound from the (real or virtual) microphone array audio signal and estimates the partial non-reverberant sound.

[0086] Thus, the discussed conceptual implementations can provide two related solutions to two related problems:

[0087] The spatial coherence across an area in a certain direction, which is related to the directional part of the acoustic energy;

[0088] Surround spatial coherence, which is related to the ambient / non-directional portion of the sound energy.

[0089] In some embodiments, the method may include estimating (actually or virtually) whether the sound field already contains spatially separated coherent sound sources (e.g., speakers of a PA system). This may be estimated, for example, by obtaining zero-order and first-order spherical harmonics and comparing the energies of the zero-order and first-order harmonics. This produces a general coherence estimate, which is converted into extended and surround coherence parameters based on the energy ratio parameter.

[0090] In some embodiments, the method may include estimating whether the non-directional portion of the audio should be reproduced incoherently or coherently. This information can be obtained in a variety of ways. For example, it can be obtained by analyzing the input microphone signal. For example, if the microphone signal is analyzed as anechoic, the surround coherence parameter can be set to a large value. As another example, the information can be obtained visually. For example, if the visual depth map shows that the sound source is very close and all reflection sources are very far away, it can be estimated that the input audio signal is mainly anechoic, and therefore the surround coherence parameter should be set to a large value. In this method, the extended coherence parameter can remain unchanged (e.g., zero).

[0091] Furthermore, as will be discussed in further detail below, the ratio parameters may be modified based on the determined spatial coherence or audio signal relationship parameters to further improve the audio quality.

[0092] about Figure 1 , shows an example device and system for implementing an embodiment of the present application. The system 100 is shown to have an "analysis" part 121 and a "synthesis" part 131. The "analysis" part 121 is the part from receiving the microphone array audio signal until the metadata and transmission signal are encoded, and the "synthesis" part 131 is the part from decoding the encoded metadata and transmission signal to presenting the regenerated signal (for example, in the form of a multi-channel speaker).

[0093] The input to the system 100 and the "analysis" part 121 is the microphone array audio signal 102. The microphone array audio signal can be obtained from any suitable capture device, or a virtual microphone recording obtained from, for example, a speaker signal, and the capture device can be local or remote to the example device. For example, in some embodiments, the analysis component 121 is integrated on a suitable capture device.

[0094] The microphone array audio signal is passed to the transmission signal generator 103 and the analysis processor 105 .

[0095] In some embodiments, the transmission signal generator 103 is configured to receive the microphone array audio signal and generate a suitable transmission signal 104. The transmission audio signal may also be referred to as an associated audio signal, and is based on a spatial audio signal that contains directional information of the sound field and is input into the system. For example, in some embodiments, the transmission signal generator 103 is configured to downmix or otherwise select or combine the microphone array audio signal to a determined number of channels, such as by beamforming technology, and output them as a transmission signal 104. The transmission signal generator 103 may be configured to generate a 2-audio channel output of the microphone array audio signal. The determined number of channels may be any suitable number of channels. In some embodiments, the transmission signal generator 103 is optional, and the microphone array audio signal is passed to the encoder unprocessed in the same manner as the transmission signal. In some embodiments, the transmission signal generator 103 is configured to select one or more of the microphone audio signals and output the selection as a transmission signal 104. In some embodiments, the transmission signal generator 103 is configured to apply any suitable encoding or quantization to the microphone array audio signal or the processed or selected form of the microphone array audio signal.

[0096] In some embodiments, the analysis processor 105 is also configured to receive the microphone array audio signal and analyze the signal to generate metadata 106 associated with the microphone array audio signal and therefore associated with the transmission signal 104. The analysis processor 105 can be, for example, a computer (running suitable software stored on a memory and at least one processor), or alternatively a specific device using, for example, an FPGA or ASIC. As shown in more detail herein, the metadata can include, for each time-frequency analysis interval, a directional parameter 108, an energy ratio parameter 110, a surround coherence parameter 112, and an extended coherence parameter 114. The directional parameters and the energy ratio parameters can be considered to be spatial audio parameters in some embodiments. In other words, the spatial audio parameters include parameters intended to characterize the sound field captured by the microphone array audio signal.

[0097] In some embodiments, the parameters generated may differ between frequency bands. Thus, for example, in frequency band X, all parameters are generated and transmitted, while in frequency band Y, only one parameter is generated and transmitted, and further, in frequency band Z, no parameters are generated or transmitted. A practical example may be that for certain frequency bands, such as the highest frequency band, certain parameters are not required for perceptual reasons. The transmission signal 104 and the metadata 106 may be transmitted or stored, which may be Figure 1 107. Before the transmission signal and metadata 106 are sent or stored, they are usually encoded to reduce the bit rate and multiplexed into a stream. Any suitable scheme can be used to achieve encoding and multiplexing.

[0098] On the decoder side, the received or retrieved data (stream) can be demultiplexed and the encoded stream can be decoded to obtain the transmission signal and metadata. Figure 1 Such reception or retrieval of transmission signals and metadata is also shown relative to the right side of dashed line 107 in FIG.

[0099] The "synthesis" portion 131 of the system 100 shows a synthesis processor 109 configured to receive the transmission signal 104 and the metadata 106, and create a suitable multi-channel audio signal output 116 based on the transmission signal 104 and the metadata 106 (which can be any suitable output format, such as a binaural multi-channel speaker or an Atmos signal, depending on the use case). In certain embodiments with speaker reproduction, an actual physical sound field with desired perceptual characteristics is reproduced (using the speakers). In other embodiments, the reproduction of the sound field may be understood to refer to the reproduction of the perceptual characteristics of the sound field by other means than the reproduction of the actual physical sound field in the space. For example, the desired perceptual characteristics of the sound field can be reproduced on headphones using the binaural reproduction method described herein. In another example, the perceptual characteristics of the sound field may be reproduced as an Atmos output signal, and these Atmos signals may be reproduced by an Atmos decoding method to provide, for example, a binaural output with desired perceptual characteristics.

[0100] In some embodiments, the synthesis processor 109 may be a computer (running suitable software stored on a memory and at least one processor), or alternatively a specific device utilizing, for example, an FPGA or ASIC.

[0101] about Figure 2 , showing Figure 1 An example flow chart of the overview shown.

[0102] First, the system (analysis part) is configured to receive microphone array audio signals, such as Figure 2 As shown in step 201.

[0103] The system (analysis part) is then configured to generate a transmission signal (e.g., based on downmixing / selection / beamforming of the microphone array audio signal), such as Figure 2 As shown in step 203.

[0104] The system (analysis part) is also configured to analyze the microphone array audio signal to generate metadata: direction; energy ratio; surround coherence; extended coherence, such as Figure 2 As shown in step 205.

[0105] The system is then configured to (optionally) encode the transmission signal with the coherence parameters and metadata for storage / transmission, such as Figure 2As shown in step 207.

[0106] Thereafter, the system can store / send the transmission signal and metadata with coherence parameters, e.g. Figure 2 As shown in step 209.

[0107] The system can retrieve / receive transmission signals and metadata with coherence parameters, such as Figure 2 As shown in step 211.

[0108] Then, the system is configured to extract from the transmission signal and metadata the coherence parameters, such as Figure 2 As shown in step 213.

[0109] The system (synthesis part) is configured to synthesize an output multi-channel audio signal (which, as previously discussed, may be any suitable output format, such as binaural, multi-channel speaker or panoramic sound signal, depending on the use case) based on the extracted audio signal with coherence parameters and metadata, such as Figure 2 As shown in step 215.

[0110] about Figure 3 , according to some embodiments of the example analysis processor 105 (such as Figure 1 In some embodiments, the analysis processor 105 includes a time-frequency domain converter 301.

[0111] In some embodiments, the time-frequency domain converter 301 is configured to receive the microphone array audio signal 102 and apply an appropriate time-to-frequency domain transform, such as a short-time Fourier transform (STFT), to convert the input time domain signal into a suitable time-frequency signal. These time-frequency signals can be passed to the direction analyzer 303 and to the coherence analyzer 305.

[0112] Thus, for example, the time-frequency signal 302 may be represented in the time-frequency domain as

[0113] s i (b,n),

[0114] Where b is the frequency bin index, n is the frame index, and i is the microphone index. In another expression, n can be viewed as a time index with a lower sampling rate than the original time domain signal. These frequency bins can be grouped into subbands, which group one or more bins into band indexes k=0, ..., K-1. Each subband k has a minimum bin b k,low and the highest warehouse b k,high , and this subband contains k,low to b k,highThe width of the sub-band can be approximated by any suitable distribution, such as the Equirectangular Bandwidth (ERB) scale or the Bark scale.

[0115] In some embodiments, the analysis processor 105 comprises a direction analyzer 303. The direction analyzer 303 may be configured to receive the time-frequency signals 302 and, based on these signals, estimate the direction parameter 108. The direction parameter may be determined based on any audio-based determination of "direction".

[0116] For example, in some embodiments, the direction analyzer 303 is configured to estimate the direction using two or more microphone signal inputs. This represents the simplest configuration to estimate the "direction", and more complex processing can be performed using more microphone signals.

[0117] Direction analyzer 303 may therefore be configured to provide an azimuth angle, denoted as θ(k,n), for each frequency band and time frame. In the case where the directional parameter is a 3D parameter, an example directional parameter may be the azimuth angle θ(k,n), the elevation angle As indicated by the dashed line, the direction parameter 108 may also be passed to the coherence analyzer 305 .

[0118] In some embodiments, in addition to the direction parameter, the direction analyzer 303 is configured to determine other suitable parameters associated with the determined direction parameter. For example, in some embodiments, the direction analyzer is caused to determine an energy ratio parameter 304. The energy ratio can be considered as a determination of the energy of the audio signal, which can be considered to arrive from one direction. For example, the energy ratio r(k,n) can be estimated (directly to the whole) using a stability measure of the directional estimate or using any correlation measure or any other suitable method to obtain the energy ratio parameter. In other embodiments, the direction analyzer is caused to determine and output a stability measure of the direction estimate, a correlation measure, or other parameter associated with the direction.

[0119] The estimated direction 108 parameters may be output (and used in a synthesis processor). The estimated energy ratio parameters 304 may be passed to a coherence analyzer 305. In some embodiments, the parameters may be received in a parameter combiner (not shown), where the estimated direction and energy ratio parameters are combined with the coherence parameters generated by the coherence analyzer 305 described below.

[0120] In some embodiments, the analysis processor 105 includes a coherence analyzer 305. The coherence analyzer 305 is configured to receive parameters (e.g., azimuth angle (θ(k,n)) 108 and direct-to-total energy ratio (r(k,n)) 304) from the direction analyzer 303. The coherence analyzer 305 may be further configured to receive a time-frequency signal (s i(b,n)) 302. All of this is in the time-frequency domain; b is the frequency bin index, k is the frequency band index (each band may consist of several bins b), n is the time index, and i is the microphone index.

[0121] Although the direction and ratio are represented here for each time index n, in some embodiments, the parameters can be combined over multiple time indices. As already expressed, the same applies to the frequency axis, the direction of multiple frequency bins b can be represented by one directional parameter in a frequency band k composed of multiple frequency bins b. The same applies to all spatial parameters discussed herein.

[0122] The coherence analyzer 305 is configured to generate a plurality of coherence parameters. In the following disclosure, there are two parameters: surround coherence (γ(k,n)) and extended coherence The analysis is performed in the time-frequency domain. In addition, in some embodiments, the coherence analyzer 305 is configured to modify the estimated energy ratio (r(k,n)). The modified energy ratio r' can be used to replace the original energy ratio r.

[0123] Next, each of the above spatial coherence issues related to the directional ratio parameterization is discussed, and it is shown how the above new parameters are formed in each case. All processing is performed in the time-frequency domain, so the time-frequency indices k and n are discarded for simplicity. As mentioned earlier, in some cases, the spatial metadata can be represented with another frequency resolution than the frequency resolution of the time-frequency signal.

[0124] These (modified) energy ratio 110, surround coherence 112, and extended coherence 114 parameters may then be output. As discussed, these parameters may be passed to a metadata combiner or processed in any suitable manner, such as encoding and / or multiplexing with a transmission signal, and stored and / or transmitted (and passed to the synthesis portion of the system).

[0125] about Figure 4 , a flowchart outlining the operation of analysis processor 105 is shown.

[0126] The first operation is as follows Figure 4 An operation of receiving a time domain microphone array audio signal as shown in step 401.

[0127] Next, a time-to-frequency transform (e.g., STFT) is applied to generate a suitable time-frequency domain signal for analysis, such as Figure 4 As shown in step 403.

[0128] Then, directional / spatial analysis is applied to the microphone array audio signal to determine the direction and energy ratio parameters, such as Figure 4 As shown in step 405.

[0129] Then, coherence analysis is applied to the microphone array audio signal to determine coherence parameters, such as surround and / or extended coherence parameters, such as Figure 4 As shown in step 407.

[0130] In some embodiments, in this step, the energy ratio may also be modified based on the determined coherence parameter.

[0131] exist Figure 4 A final operation of outputting the determined parameters is shown by step 409.

[0132] about Figure 5 , showing a first example of a coherence analyzer according to some embodiments.

[0133] The first example implements a method for determining spatial coherence using a first-order Atmos (FOA) signal, which can be generated using certain microphone arrays (at least within a defined frequency range). Alternatively, the FOA signal can be virtually generated from other audio signal formats (such as speaker input signals). The following method estimates the extended and surround coherence occurring in the sound field. An example microphone array providing the FOA signal is a B-format microphone that provides an omnidirectional signal and three dipole signals.

[0134] Note that if the FOA signal is generated virtually (in other words, eg converted from a loudspeaker format), the input signal to the coherence analyzer is the FOA signal which is then transformed into the time-frequency domain for directionality and coherence analysis.

[0135] The zeroth-order and first-order spherical harmonics determiner 501 may be configured to receive the time-frequency microphone audio signal 302 and generate a suitable time-frequency spherical harmonics signal 502 .

[0136] The general coherence estimator 503 may be configured to receive the time-frequency spherical harmonics signal 502 (which may be captured at a sound field with spatially separated coherent sound sources or generated by the zeroth and first order spherical harmonics determiner 501). The general coherence parameter μ(k,n) may be generated by monitoring the energy of the FOA components.

[0137] If any microphone capable of producing FOA signals is placed in a diffuse field, the energy of the three dipole signals X, Y, Z has the same sum of energy as the omnidirectional component W (gain balance between W and X, Y, Z according to Schmidt Semi-Normalized (SN3D)). However, if the sound is coherently reproduced at spatially separated loudspeakers, the energy of the X, Y, Z signals becomes small (or even zero) because the X, Y, Z modes have positive amplitudes in one direction and negative amplitudes in another direction, and thus signal cancellation occurs for spatially separated coherent sound sources.

[0138] By generating and monitoring coherent or incoherent surround signals, it is possible to determine a formula that provides an estimate for the general coherence parameter μ based on the energy information of the FOA signal.

[0139] c a,b Denoted as the (a, b) entries of the estimated covariance matrix (W, X, Y, Z) of the FOA signal, and the general coherence parameter μ can be estimated by

[0140]

[0141] Therein, the time-frequency index is omitted. The coefficient p may have the value 1, for example.

[0142] The general coherence to extended coherence and surround coherence splitter 505 is configured to receive the generated general coherence 504 and the energy ratio 304 and to generate estimates of extended and surround coherence parameters based on the general coherence parameter.

[0143] In some embodiments, the energy ratio may be used to split the general coherence into extended coherence and surround coherence. Thus, for example, the diffuse and surround coherences may be estimated as:

[0144]

[0145] γ(k,n)=(1-r(k,n))μ(k,n)

[0146] in, is the extended coherence parameter 114, γ is the surround coherence parameter 112, and r is the energy ratio. In practice, if the direct to total energy ratio is large, the universal coherence is transformed into the extended coherence; if the direct to total energy ratio is small, the universal coherence is transformed into the surround coherence.

[0147] In some embodiments, the common coherence to the extended and surround coherence divider 505 is configured to simply set both the extended and surround coherence parameters to the common coherence parameter.

[0148] about Figure 6 , showing a summary of Figure 5 A flow chart of the operation of a first exemplary coherence analyzer is shown.

[0149] The first operation is an operation of receiving the time-frequency domain microphone array audio signal and energy ratio, such as Figure 6 As shown in step 601.

[0150] Next, apply the appropriate transformation to generate the zeroth and first order spherical harmonics, as Figure 6 As shown in step 603.

[0151] Then, by determining the ratio of the spherical harmonics, the universal coherence can be estimated as Figure 6 As shown in step 605.

[0152] The estimated universal coherence value is then split into extended and surround coherence estimates, as Figure 6 As shown in step 607.

[0153] The final operation is one that outputs a certain coherence parameter, such as Figure 6 As shown in step 609.

[0154] about Figure 7 , another exemplary coherence analyzer is shown.

[0155] These examples estimate whether the non-directional parts of the audio will be reproduced as coherent or incoherent sounds for optimal audio quality. The analyzer provides surround coherence parameters and works with any microphone array, including those that cannot provide FOA signals.

[0156] The undeviated sound estimator 701 is configured to receive the time-frequency microphone array audio signal and estimate a portion of undeviated sound.

[0157] The estimation of the amount of direct and reverberant sound in the captured microphone signal can be implemented according to any known method, or even extracting the direct and reverberant components from the mix. In some embodiments, the estimation can be generated from another source besides the captured audio signal. For example, in some embodiments, visual information can be used to estimate the amount of direct and reverberant sound. For example, if the visual depth map shows that the sound source is very close, while all reflection sources are far away, it can be estimated that the input audio signal is mainly anechoic (and therefore the surround coherence parameter should be set to a large value). In some embodiments, the user can even manually select the estimation.

[0158] An example method for analyzing a microphone audio signal to determine an estimate of the direct sound component may be obtained using spectral subtraction:

[0159] D(k,n)=S(k,n)-R(k,n)

[0160] where D is the estimated direct acoustic energy component and S is the estimated total signal energy (which can be estimated from any microphone signal, for example, S = E[s 2 ], or a mixture thereof), R is the estimated reverberant sound energy component. An estimate of R can be obtained by filtering the estimated direct sound energy component D with an estimated decaying coefficient. The decaying coefficient itself can be estimated, for example, using a blind reverberation time estimation method.

[0161] Using the estimated direct sound component D, the portion of direct sound in the captured microphone signal can be estimated:

[0162]

[0163] The estimated energy values ​​S(k,n) etc. may have been averaged over several time and / or frequency indices (k,n).

[0164] If the non-directional audio is mostly reverberant, then reproducing it as incoherent is optimal, since incoherence is desired in order to reproduce the sense of envelope and spaciousness that is natural to reverberation, and decorrelation is typically required not to deteriorate audio quality in the presence of reverberation. If the non-directional audio is mostly ireverberant, then reproducing it as coherent is desirable, since incoherence is not required for such sounds, and decorrelation can deteriorate audio quality (especially in the case of speech signals). Thus, the choice of coherent / incoherent reproduction of the non-directional audio can be guided based on the analyzed reverberation.

[0165] Surround coherence estimator 703 may receive the estimate of the non-reverberant sound portion 702 and the energy ratio 304 and estimate the surround coherence 112. The directional portion of the captured microphone signal defined by the energy ratio r may be approximated as only direct sound. The ambient portion of the signal (defined by 1-r) may be approximated as a mixture of reverberation, ambient sound, and direct sound during double talk.

[0166] If the ambient portion contains only reverberation and ambient sounds, the surround coherence γ should be set to 0 (these should be reproduced as incoherent). However, if during double talk, the ambient portion contains only direct sound, the surround coherence γ should be set to 1 (it should be reproduced as coherent to avoid decorrelation). For example, using these principles, the equation for surround coherence γ can be formed as:

[0167]

[0168] In this method, the extended coherence can be Set to zero.

[0169] about Figure 8 , showing a summary of Figure 7 A flow chart of the operation of a second exemplary coherence analyzer is shown.

[0170] The first operation is to receive the time-frequency domain microphone array audio signal and the energy ratio operation, such as Figure 8 As shown in step 801.

[0171] Next, the non-reverberant sound portion is estimated, such as Figure 8As shown in step 803.

[0172] Then, the surround coherence is estimated based on the partial and energy ratio of the NR sound, as Figure 8 As shown in step 805.

[0173] The final operation is the one that outputs the determined coherence parameters, such as Figure 8 As shown in step 807.

[0174] In some embodiments, two coherence analyzers may be implemented and the outputs combined. For example, the combination may be achieved by taking the maximum of the two estimates:

[0175]

[0176] γ(k,n)=max(γ1(k,n),γ2(k,n)).

[0177] about Fig. 9 , showing an example synthesis processor 109 in more detail. The example synthesis processor 109 may be configured to utilize a modified method according to any known method, such as a method that is particularly suitable for the case where inter-channel signal coherence needs to be synthesized or manipulated.

[0178] The synthesis method can be a modified least squares optimized signal mixing technique to manipulate the covariance matrix of the signal while attempting to maintain audio quality. The method utilizes the covariance matrix measurement and the target covariance matrix of the input signal (as discussed below) and provides a mixing matrix to perform this processing. The method also provides a means to best utilize decorrelated sounds when there is not a sufficient amount of independent signal energy at the input.

[0179] The synthesis processor 109 may include a time-frequency domain converter 901 configured to receive an audio input in the form of a transmission signal 104 and apply an appropriate time-to-frequency domain transform (e.g., short-time Fourier transform (STFT)) to convert the input time domain signal into a suitable time-frequency signal. These time-frequency signals may be passed to a mixing matrix processor 909 and a covariance matrix estimator 903.

[0180] The time-frequency signal may then be adaptively processed in the frequency band using a hybrid matrix processor (and possibly a decorrelation processor) 909. The output of the hybrid matrix processor 909 may be passed to an inverse time-frequency domain transformer 911 in the form of a time-frequency output signal 912. The inverse time-frequency domain transformer 911 (e.g., an inverse short-time Fourier transformer or I-STFT) is configured to transform the time-frequency output signal 912 into the time domain to provide a processed output in the form of a multi-channel audio signal 116. The hybrid matrix processing method is well documented and will not be described in detail below.

[0181] The mixing matrix determiner 907 may generate a mixing matrix and pass it to a mixing matrix processor 909. The mixing matrix determiner 907 may be caused to generate a mixing matrix for a frequency band. The mixing matrix determiner 907 is configured to receive an input covariance matrix 906 and a target covariance matrix 908 organized in frequency bands.

[0182] By measuring the time-frequency signal (transmission signal in frequency band) from the time-frequency domain transformer 901, the covariance matrix estimator 903 can generate covariance matrices organized in frequency bands 906. These estimated covariance matrices can then be passed to the mixing matrix determiner 907.

[0183] Furthermore, the covariance matrix determiner 903 may be configured to estimate a total energy E 904 and pass it to the target covariance matrix determiner 905. In some embodiments, the total energy E may be determined from the sum of the diagonal elements of the estimated covariance matrix.

[0184] The target covariance matrix determiner 905 is caused to generate a target covariance matrix. In some embodiments, the target covariance matrix determiner 905 may determine the target covariance matrix to reproduce to a surround speaker setup. In the following expressions, the time and frequency indices n and k are deleted for simplicity (when not needed).

[0185] First, the target covariance matrix determiner 905 may be configured to receive the total energy E 904 based on the input covariance matrix from the covariance matrix estimator 903 and the spatial metadata 106 .

[0186] Then, the target covariance matrix determiner 905 may be configured to determine the covariance matrix between the mutually incoherent part, the directional part C D and the ambient or non-directional part C A The target covariance matrix C in T .

[0187] Therefore, the target covariance matrix is ​​determined by the target covariance matrix determiner 905 as C T =C D +C A .

[0188] Environmental Section C A Expressing spatial surround sound energy, which before was only incoherent, but thanks to the invention it can be incoherent or coherent or partially coherent.

[0189] The target covariance matrix determiner 905 may therefore be configured to determine the ambient energy as (1-r)E, where r is a direct to total energy ratio parameter from the input metadata. The ambient covariance matrix may then be determined as follows:

[0190]

[0191] Where I is the identity matrix, U is a matrix of 1s, and M is the number of output channels. In other words, when γ is zero, the environment covariance matrix C A is diagonal, and when γ is 1, the ambient covariance matrix makes sure that all channel pairs are coherent.

[0192] The target covariance matrix determiner 905 may then be configured to determine the direct part covariance matrix C D .

[0193] The target covariance matrix determiner 905 may therefore be configured to determine the direct part energy as rE.

[0194] The target covariance matrix determiner 905 is then configured to determine a gain vector for the speaker signal based on the metadata. First, the target covariance matrix determiner 905 is configured to determine a vector of amplitude panning gains based on the speaker settings and directional information of the spatial metadata, for example using vector basis amplitude panning (VBAP). These gains can be indicated in a column vector vVBAP, which can be implemented in three-dimensional space using any suitable virtual space polygon arrangement (typically triangular in nature, and therefore defined in terms of channel or node triplets in the following examples). In some embodiments, for two speakers active in amplitude panning, the horizontal setting has at most only two non-zero values. In some embodiments, the target covariance matrix determiner 905 can be configured to determine the VBAP covariance matrix as:

[0195]

[0196] The target covariance matrix determiner 905 may be configured to determine the channel triplet i l ,i r ,i c , which is the speaker closest to the estimated direction, and the closest left and right speakers.

[0197] The target covariance matrix determiner 905 may also be configured to determine a translation column vector vLRC, which is at index i l ,i r ,i c Has value The others are zero. The covariance matrix of this vector is

[0198]

[0199] When the extended coherence parameter When θ is less than 0.5, i.e., when the sound is to be reproduced between the “direct point source” scenario and the “three-speaker coherent sound” scenario, the target covariance matrix determiner 305 may be configured to determine the direct part covariance matrix as

[0200]

[0201] When the extended coherence parameter When the target covariance matrix determiner 905 is between 0.5 and 1, that is, when the sound is to be reproduced between the “three-speaker coherent sound” scenario and the “two-extended speaker coherent sound” scenario, the extended distribution vector

[0202]

[0203] Then, the target covariance matrix determiner 905 may be configured to determine the translation vector v DISTR , where the i c The entry is v DISTR,3 The first entry of l and the i-th r The entry is v DISTR,3 The second and third entries of . Then, the target covariance matrix determiner 905 can calculate the direct part covariance matrix as:

[0204]

[0205] Then the target covariance matrix determiner 905 can obtain the target covariance matrix C T =C D +C A As mentioned above, the ambient partial covariance matrix therefore takes into account the spatial coherence contained in the ambient energy and the surround coherence parameter γ, and the direct covariance matrix takes into account the directional energy, the directional parameter and the extended coherence parameter

[0206] The target covariance matrix determiner 905 may be configured to determine the target covariance matrix 908 for binaural output by being configured to synthesize inter-aural characteristics of the surround sound instead of inter-channel characteristics.

[0207] Therefore, the target covariance matrix determiner 905 may be configured to determine the environmental covariance matrix C for binaural sound A The amount of ambient or non-directional energy is (1-r)E, where E is the total energy determined previously. The ambient partial covariance matrix can be determined as:

[0208]

[0209] in

[0210] c(k,n)=γ(k,n)+(1-γ(k,n))c bin (k),

[0211] And among them, c bin (k) is the binaural diffuse field coherence of the frequency indexed by the kth frequency. In other words, when γ(k,n) is 1, then the environmental covariance matrix C A Determine the perfect coherence between the left and right ears. When γ(k,n) is zero, then C A Determine the coherence between the left and right ears in a diffuse field that is natural to the listener (roughly: zero at high frequencies, higher at low frequencies).

[0212] Thus, the target covariance matrix determiner 905 may be configured to determine the direct part covariance matrix C D The amount of directional energy is rE. A similar approach can be used to synthesize the extended coherence parameter ζ as in loudspeaker reproduction, as described in detail below.

[0213] First, the target covariance matrix determiner 905 may be configured to determine a 2×1 HRTF vector v HRTF (k, θ(k,n)), where θ(k,n) is the estimated directional parameter. The target covariance matrix determiner 905 may determine a panning HRTF vector equivalent to reproducing the sound coherently in three directions.

[0214]

[0215] Among them, θ Δ The parameter defines how wide the sound energy is "spread" relative to the azimuth dimension. It could be 30 degrees, for example.

[0216] When the extended coherence parameter When φ is less than 0.5, i.e., when the sound is to be reproduced between the “direct point source” scenario and the “three-speaker coherent sound” scenario, the target covariance matrix determiner 905 may be configured to determine the direct part HRTF covariance matrix as

[0217]

[0218] When the extended coherence parameter ζ is between 0.5 and 1, i.e., when the sound is to be reproduced between the “three-speaker coherent sound” scenario and the “two-extended-speaker coherent sound” scenario, the target covariance matrix determiner 905 can obtain the target covariance matrix by reusing the amplitude distribution vector V DISTR,3 The extension distribution is determined (same as in speaker rendering). Then, the combined head-related transfer function (HRTF) vector can be determined as

[0219] V DISTR_HRTF(k,θ(k,n))

[0220] =[v HRTF (k,θ(k,n))v HRTF (k,θ(k,n)+θ Δ )v HRTF (k,θ(k,n)-θ Δ )]v DISTR,3

[0221] The above formula produces a DISTR,3 The weighted sum of the three HRTFs with the weights in . So the direct part HRTF covariance matrix is

[0222]

[0223] Then, the target covariance matrix determiner 905 is configured to obtain the target covariance matrix C T =C D +C A As mentioned above, the ambient partial covariance matrix therefore takes into account the ambient energy and spatial coherence contained in the surround coherence parameter γ, while the direct covariance matrix takes into account the directional energy, directional parameters and extended coherence parameters

[0224] The target covariance matrix determiner 905 may be configured to determine a target covariance matrix 908 for the panoramic sound output by being configured to synthesize the inter-channel characteristics of the panoramic sound signal instead of the inter-channel characteristics of the speaker surround sound. The first-order panoramic sound (FOA) output is taken as an example below, but it is also simple to extend the same principle to higher-order panoramic sound outputs.

[0225] Therefore, the target covariance matrix determiner 905 may be configured to determine the ambient covariance matrix C for the panoramic sound A The amount of ambient or non-directional energy is (1-r)E, where E is the total energy determined previously. The ambient partial covariance matrix can be determined as

[0226]

[0227] In other words, when γ(k,n) is 1, the environmental covariance matrix C A So that only the 0th order component receives the signal. The meaning of this panoramic sound signal is to reproduce the sound coherently in space. When γ(k,n) is zero, C A corresponds to the Atmos covariance matrix in the diffuse field. The normalization of the 0th and 1st order elements above is according to the known SN3D normalization scheme.

[0228] Thus, the target covariance matrix determiner 905 may be configured to determine the direct part covariance matrix CD The amount of directed energy is rE. A similar approach can be used to synthesize the extended coherence parameter As in loudspeaker reproduction, the detailed description is as follows.

[0229] First, the target covariance matrix determiner 905 may be configured to determine a 4×1 panoramic sound translation vector v Amb (θ(k,n)), where θ(k,n) is the estimated direction parameter. The panoramic sound translation vector v Amb (θ(k,n)) contains the panoramic gain corresponding to the direction θ(k,n). For FOA output with directional parameters in the horizontal plane (using the known ACN channel ordering scheme)

[0230]

[0231] The target covariance matrix determiner 905 may determine a panning panoramic sound vector, which is equivalent to reproducing the sound coherently in three directions.

[0232]

[0233] Among them, θ Δ The parameter defines how wide the sound energy is "spread" relative to the azimuth dimension. It could be 30 degrees, for example.

[0234] When the extended coherence parameter When θ is less than 0.5, i.e., when the sound is to be reproduced between the “direct point source” scenario and the “three-speaker coherent sound” scenario, the target covariance matrix determiner 905 may be configured to determine the direct partial Atmos covariance matrix as

[0235]

[0236] When the extended coherence parameter When the target covariance matrix determiner 305 is between 0.5 and 1, that is, when the sound is to be reproduced between the “three-speaker coherent sound” scenario and the “two-speaker coherent sound” scenario, the target covariance matrix determiner 305 can be configured by reusing the amplitude distribution vector v DISTR,3 The extension distribution is determined (same as in speaker rendering). Then, the combined panoramic sound translation vector can be determined as

[0237] v DISTR_Amb (θ(k,n))=[v Amb (θ(k,n))v Amb (θ(k,n)+θ Δ )v Amb (θ(k,n)-θ Δ )]v DISTR,3 .

[0238] The above formula produces a DISTR,3 The weighted sum of the three panoramic sound translation vectors with the weights in . Therefore, the direct partial panoramic sound covariance matrix is

[0239]

[0240] Then, the target covariance matrix determiner 905 is configured to obtain the target covariance matrix C T =C D +C A As mentioned above, the ambient partial covariance matrix therefore takes into account the ambient energy and spatial coherence contained in the surround coherence parameter γ, while the direct covariance matrix takes into account the directional energy, directional parameters and extended coherence parameters

[0241] In other words, the same general principles apply to constructing binaural or Atmos or speaker target covariance matrices. The main differences are utilizing HRTF data or Atmos panning data instead of speaker amplitude panning data in the rendering of the direct portion, and utilizing binaural coherence (or specific Atmos ambience covariance matrix processing) instead of inter-channel (zero) coherence in the rendering of the ambient portion. It should be understood that a processor may be capable of running software to achieve the above, and therefore capable of rendering each of these output types.

[0242] In the above formula, the energy of the direct and environmental parts of the target covariance matrix is ​​weighted based on the total energy estimate E from the covariance matrix estimated in the covariance matrix estimator 903. Optionally, this weighting can be omitted, that is, the direct part energy is determined as r and the environmental part energy is determined as (1-r). In that case, the estimated input covariance matrix is ​​normalized (i.e., multiplied by 1 / E) with the total energy estimate instead. The mixed matrix based on the results of these determined target covariance matrices and the normalized input covariance matrices may be exactly the same or virtually the same as the formulas previously provided, because the relative energies of these matrices are important rather than their absolute energies.

[0243] about Fig.10 , showing an overview of the synthesis operation.

[0244] Therefore, this method can receive time domain transmission signals, such as Fig.10 As shown in step 1001.

[0245] These transmission signals can then be time-frequency transformed, e.g. Fig.10 As shown in step 1003.

[0246] Then, the covariance matrix can be estimated from the input (transmitted) signal as Fig.10 As shown in step 1005.

[0247] In addition, spatial metadata with direction, energy ratio and coherence parameters can be received, e.g. Fig.10 As shown in step 1002.

[0248] The target covariance matrix can be determined from the estimated covariance matrix, direction, energy ratio and coherence parameter, as Fig.10 As shown in step 1007.

[0249] The mixing matrix can then be determined based on the estimated covariance matrix and the target covariance matrix, as Fig.10 As shown in step 1009.

[0250] The mixing matrix can then be applied to the time-frequency transmission signal, such as Fig.10 As shown in step 1011.

[0251] The result of applying the mixing matrix to the time-frequency transmission signal can then be transformed into the inverse time-frequency domain to generate a spatial audio signal, such as Fig.10 As shown in step 1013.

[0252] about Fig.11 , showing an example method for generating a target covariance matrix according to some embodiments.

[0253] First, the total energy E of the target covariance matrix is ​​estimated based on the input covariance matrix, as Fig.11 As shown in step 1101.

[0254] The method may further include receiving spatial metadata having direction, energy ratio and coherence parameters, such as Fig.11 As shown in step 1102.

[0255] The method may then include determining the ambient energy as (1-r)E, where r is a direct to total energy ratio parameter from the input metadata, such as Fig.11 As shown in step 1103.

[0256] Furthermore, the method may include estimating the environmental covariance matrix, such as Fig.11 As shown in step 1105.

[0257] The method may also include determining the direct partial energy as rE, where r is a direct to total energy ratio parameter from the input metadata, such as Fig.11 As shown in step 1104.

[0258] The method may then include determining a vector of amplitude panning gains based on the loudspeaker setup and the directional information of the spatial metadata, such as Fig.11 As shown in step 1106.

[0259] Thereafter, the method may include determining a channel triplet, the triplet being the speaker closest to the estimated direction, and the closest left and right speakers, such as Fig.11 As shown in step 1108.

[0260] The method can then include estimating the direct covariance matrix, such as Fig.11 As shown in step 1110.

[0261] Finally, the method may include combining the ambient and direct covariance matrix components to generate a target covariance matrix, such as Fig.11 As shown in step 1112.

[0262] The above statements discuss the construction of a target covariance matrix. The method may also use a prototype matrix formed according to any known method. The prototype matrix determines a "reference signal" for rendering, for which a least squares optimized mixing matrix is ​​formulated. If a stereo downmix is ​​provided as an audio signal in the codec, the prototype matrix for speaker rendering may make it possible to determine that the signal for the left-hand speaker is optimal relative to the left channel of the provided stereo track, and the same applies to the right-hand side (the center channel may be optimized for the sum of the left and right audio channels). For binaural output, the prototype matrix may make it possible to determine that the reference signal for the left ear output signal is the left stereo channel, and similarly for the right ear. The determination of the prototype matrix is ​​simple and straightforward for a person skilled in the art who has studied the existing literature. With respect to the existing literature, the novelty of the inventive scheme at the synthesis stage is that the target covariance matrix is ​​also constructed using spatial coherence metadata.

[0263] Although not repeated throughout the document, it should be understood that, typically and in this context, spatial audio processing is performed in frequency bands. These frequency bands can be, for example, frequency bins of a time-frequency transform, or frequency bands that combine multiple frequency bins. The combination can approximate the characteristics of human hearing, such as Bark frequency resolution. In other words, in some cases, audio can be measured and processed in a time-frequency region that combines multiple frequency bins b and / or time indexes n. For simplicity, these aspects are not expressed by all of the above formulas. In the case of combining many time-frequency samples, a set of parameters (e.g., one direction) is typically estimated for the time-frequency region, and all time-frequency samples in the region are synthesized based on the set of parameters (e.g., the one direction parameters).

[0264] Using a frequency resolution for parametric analysis that is different from the frequency resolution of the applied filter bank is a typical approach in spatial audio processing systems.

[0265] Although the examples presented herein have used microphone array audio signals as input, it should be understood that in some embodiments, the examples may be used to process virtual microphone signals as input. For example, a virtual FOA signal may be created from a multi-channel speaker or object signal, for example, by:

[0266]

[0267] For each loudspeaker (or object) signal si with its own azimuth and elevation direction w, y, z, x signals are generated. The output signal combining all such signals is

[0268] After generating the FOA signals, they may be transformed into the time-frequency domain.Directional metadata may be estimated, for example, using techniques such as DirAC, and coherence metadata estimated using the methods described herein.

[0269] Thus, embodiments may improve perceived audio quality in three different aspects:

[0270] 1) In the case of spatially separated coherent sources captured by a real or virtual microphone array, embodiments can detect such a scene and coherently reproduce the audio from spatially separated speakers, thereby maintaining a perception similar to the original audio scene.

[0271] 2) Determining spatial coherence parameters from virtual microphone array inputs provides a straightforward method to estimate these parameters from any loudspeaker / audio object configuration via an intermediate FOA transform.

[0272] 3) In the case of multiple sources existing simultaneously in dry acoustics, embodiments can detect such scenarios and reproduce the audio with less decorrelation, thus avoiding possible artifacts.

[0273] about Fig.12 , shows an example electronic device that can be used as an analysis or synthesis device. The device can be any suitable electronic device or apparatus. For example, in some embodiments, the device 1400 is a mobile device, a user device, a tablet computer, a computer, an audio playback device, etc.

[0274] In some embodiments, device 1400 includes at least one processor or central processing unit 1407. Processor 1407 may be configured to execute various program codes, such as the methods described herein.

[0275] In some embodiments, the device 1400 includes a memory 1411. In some embodiments, at least one processor 1407 is coupled to the memory 1411. The memory 1411 can be any suitable storage module. In some embodiments, the memory 1411 includes a program code portion for storing program code that can be implemented on the processor 1407. In addition, in some embodiments, the memory 1411 can also include a stored data portion for storing data (e.g., data that has been processed or will be processed according to the embodiments described herein). Whenever necessary, the implemented program code stored in the program code portion and the data stored in the stored data portion can be retrieved by the processor 1407 through the memory-processor coupling.

[0276] In some embodiments, device 1400 includes a user interface 1405. In some embodiments, user interface 1405 can be coupled to processor 1407. In some embodiments, processor 1407 can control the operation of user interface 1405 and receive input from user interface 1405. In some embodiments, user interface 1405 can enable a user to enter commands to device 1400, for example, via a keypad. In some embodiments, user interface 1405 can enable a user to obtain information from device 1400. For example, user interface 1405 can include a display configured to display information from device 1400 to a user. In some embodiments, user interface 1405 can include a touch screen or touch interface that enables information to be input to device 1400 and also displays information to a user of device 1400.

[0277] In some embodiments, the device 1400 includes an input / output port 1409. In some embodiments, the input / output port 1409 includes a transceiver. In such embodiments, the transceiver can be coupled to the processor 1407 and configured to enable communication with other devices or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver module can be configured to communicate with other electronic devices or devices via a wire or wired coupling.

[0278] The transceiver can communicate with another device via any suitable known communication protocol. For example, in some embodiments, the transceiver or transceiver module can use a suitable universal mobile telecommunications system (UMTS) protocol, a wireless local area network (WLAN) protocol such as IEEE802.X, a suitable short-range radio frequency communication protocol such as Bluetooth or infrared data communication path (IRDA).

[0279] The transceiver input / output port 1409 may be configured to receive the speaker signal and, in some embodiments, determine parameters as described herein by executing appropriate code using the processor 1407. In addition, the device may generate appropriate transmission signals and parameter outputs for transmission to a synthesis device.

[0280] In some embodiments, device 1400 may be used as at least part of a synthesis device. Thus, input / output port 1409 may be configured to receive the transmission signal and, in some embodiments, parameters determined at a capture device or processing device as described herein, and generate an appropriate audio signal format output using processor 1407 executing appropriate code. Input / output port 1409 may be coupled to any suitable audio output, such as to a multi-channel speaker system and / or headphones or the like.

[0281] In general, various embodiments of the present invention may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Although various aspects of the present invention may be illustrated and described as block diagrams, flow charts, or using some other graphical representation, it is understood that the boxes, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof, as non-limiting examples.

[0282] Embodiments of the present invention can be implemented by computer software that can be executed by a data processor of a mobile device, for example, in a processor entity, or by hardware, or by a combination of software and hardware. Further at this point, it should be noted that any block of the logic flow in the figure can represent a program step, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software can be stored on a physical medium such as a memory chip or a memory block implemented in a processor, on a magnetic medium such as a hard disk or a floppy disk, and on an optical medium such as a DVD and its data variant CD.

[0283] The memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. As non-limiting examples, the data processor may be of any type suitable for the local technical environment and may include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit, and a processor based on a multi-core processor architecture.

[0284] Embodiments of the present invention may be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert logic-level designs into semiconductor circuit designs that are easily etched and formed on semiconductor substrates.

[0285] Programs, such as those offered by Synopsys, Inc. of Mountain View, Calif., and Cadence Design of San Jose, Calif., automatically route conductors and locate components on a semiconductor chip using well-established design rules along with a library of pre-stored design modules. Once the design of a semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be transmitted to a semiconductor manufacturing facility or "fab" for fabrication.

[0286] The foregoing description provides a complete and useful description of exemplary embodiments of the present invention by way of exemplary and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the relevant art in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of the present invention will still fall within the scope of the present invention as defined by the appended claims.

Claims

1. A device comprising: at least one processor; as well as at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: determining, for two or more microphone audio signals, a plurality of spatial audio parameters for providing spatial audio reproduction, wherein the plurality of spatial audio parameters are associated with respective frequency bands of at least two frequency bands of the two or more microphone audio signals; determining at least one coherence parameter associated with a sound field, wherein the sound field is associated with the two or more microphone audio signals; determining at least one audio signal based on the two or more microphone audio signals; and Spatial audio reproduction is enabled based on the plurality of spatial audio parameters, the at least one coherence parameter and the at least one determined audio signal.

2. The device according to claim 1, wherein: The at least one coherence parameter is determined based at least in part on the two or more microphone audio signals.

3. The device according to claim 1, wherein: The at least one coherence parameter is determined based at least in part on visual information associated with the sound field.

4. The device according to claim 1, wherein: The at least one coherence parameter comprises at least one of the following: at least one extended coherence parameter based on a determination of the coherence of the directional part of the sound field; or At least one surround coherence parameter is determined based on the coherence of the non-directional portion of the sound field.

5. The device according to claim 1, wherein: in, The plurality of spatial audio parameters include at least one of: a directional parameter; Energy ratio parameter; Directly to the total energy ratio parameter; Directional stability parameters; or Energy parameters.

6. The device according to claim 1, wherein: The at least one memory stores the instructions which, when executed by the at least one processor, cause the apparatus to: determining zeroth-order spherical harmonics and first-order spherical harmonics based on the two or more microphone audio signals; generating at least one universal coherence parameter based on the zero-order spherical harmonics and the first-order spherical harmonics; as well as Based on the at least one general coherence parameter, the at least one coherence parameter is generated.

7. The device according to claim 6, wherein: The at least one memory stores the instructions which, when executed by the at least one processor, cause the apparatus to: generating at least one extended coherence parameter based on the at least one general coherence parameter and an energy ratio configured to define a relationship between a direct portion and an ambient portion of the sound field; as well as generating at least one surround coherence parameter based on the at least one general coherence parameter and the energy ratio configured to define the relationship between the direct part and the ambient part of the sound field, The at least one coherence parameter at least includes the at least one extended coherence parameter and the at least one surround coherence parameter.

8. The device according to claim 1, wherein: The at least one memory stores the instructions which, when executed by the at least one processor, cause the apparatus to: determining zeroth-order spherical harmonics and first-order spherical harmonics based on the two or more microphone audio signals; as well as At least one of the following: Based on the two or more microphone audio signals, determine time-domain zero-order spherical harmonics and time-domain first-order spherical harmonics, and convert the time-domain zero-order spherical harmonics and the time-domain first-order spherical harmonics into time-frequency domain zero-order spherical harmonics and time-frequency domain first-order spherical harmonics; or The two or more microphone audio signals are converted into corresponding two or more time-frequency domain microphone audio signals, and the time-frequency domain zero-order spherical harmonics and the time-frequency domain first-order spherical harmonics are generated based on the two or more time-frequency domain microphone audio signals.

9. The device according to claim 1, wherein: The at least one memory stores the instructions which, when executed by the at least one processor, cause the apparatus to: converting the two or more microphone audio signals into corresponding two or more time-frequency domain microphone audio signals; determining at least one estimate of undeviated sound based on the two or more time-frequency domain microphone audio signals; and At least one surround coherence parameter is determined based on the at least one estimate of the unaural sound and an energy ratio configured to define a relationship between a direct portion and an ambient portion of the sound field, wherein the at least one coherence parameter is the at least one surround coherence parameter.

10. The device according to claim 9, wherein: The at least one memory stores the instructions, which when executed by the at least one processor cause the apparatus to perform one of the following: selecting the at least one ambient coherence parameter as at least one coherence parameter based on the at least one estimate of the unaural sound and the energy ratio; or The at least one ambient coherence parameter is selected as at least one coherence parameter based on at least one general coherence parameter, wherein the at least one ambient coherence parameter is maximal based on the at least one general coherence parameter.

11. The apparatus of claim 1 , wherein the at least one memory stores the instructions, which when executed by the at least one processor cause the apparatus to: The at least one coherence parameter is determined for the respective frequency band of the at least two frequency bands.

12. A method comprising: determining, for two or more microphone audio signals, a plurality of spatial audio parameters for providing spatial audio reproduction, wherein the plurality of spatial audio parameters are associated with respective frequency bands of at least two frequency bands of the two or more microphone audio signals; determining at least one coherence parameter associated with a sound field, wherein the sound field is associated with the two or more microphone audio signals; determining at least one audio signal based on the two or more microphone audio signals; and Spatial audio reproduction is enabled based on the plurality of spatial audio parameters, the at least one coherence parameter and the at least one determined audio signal.

13. The method according to claim 12, wherein: The at least one coherence parameter is determined based at least in part on the two or more microphone audio signals.

14. The method according to claim 12, wherein: The at least one coherence parameter is determined based at least in part on visual information associated with the sound field.

15. The method according to claim 12, wherein: The at least one coherence parameter comprises at least one of the following: at least one extended coherence parameter based on a determination of the coherence of the directional part of the sound field; or At least one surround coherence parameter is determined based on the coherence of the non-directional portion of the sound field.

16. The method according to claim 12, wherein: in, The plurality of spatial audio parameters include at least one of: a directional parameter; Energy ratio parameter; Directly to the total energy ratio parameter; Directional stability parameters; or Energy parameters.

17. The method according to claim 12, further comprising: The at least one coherence parameter is determined for the respective frequency band of the at least two frequency bands.

18. The method according to claim 12, further comprising: converting the two or more microphone audio signals into corresponding two or more time-frequency domain microphone audio signals; determining at least one estimate of undeviated sound based on the two or more time-frequency domain microphone audio signals; and At least one surround coherence parameter is determined based on the at least one estimate of the unaural sound and an energy ratio configured to define a relationship between a direct portion and an ambient portion of the sound field, wherein the at least one coherence parameter is the at least one surround coherence parameter.

19. The method according to claim 18, further comprising: selecting the at least one ambient coherence parameter as at least one coherence parameter based on the at least one estimate of the unaural sound and the energy ratio; or The at least one ambient coherence parameter is selected as at least one coherence parameter based on at least one general coherence parameter, wherein the at least one ambient coherence parameter is maximal based on the at least one general coherence parameter.

20. A non-transitory computer readable medium comprising program instructions stored thereon, the program instructions being configured to at least: For two or more microphone audio signals, a plurality of spatial audio parameters for providing spatial audio reproduction are determined, wherein: The plurality of spatial audio parameters are associated with respective frequency bands of at least two frequency bands of the two or more microphone audio signals; determining at least one coherence parameter associated with a sound field, wherein the sound field is associated with the two or more microphone audio signals; determining at least one audio signal based on the two or more microphone audio signals; and Spatial audio reproduction is enabled based on the plurality of spatial audio parameters, the at least one coherence parameter and the at least one determined audio signal.