Audio rendering with spatial metadata interpolation

By acquiring parameter values ​​and location information from multiple audio signal sets, and combining them with the listener's location, modified parameter values ​​are generated. The audio signals are then processed to produce spatial audio output, solving the difficulties of linear spatial audio capture with compact microphone devices and achieving high-quality 6-DOF audio rendering.

CN120897148APending Publication Date: 2025-11-04NOKIA TECHNOLOGIES OY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511124357.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-02-26
Filing Date
2021-02-03
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

In the prior art, compact microphone devices have difficulty achieving linear spatial audio capture, especially since the microphone spacing requirements for high-frequency and low-frequency audio signals are difficult to meet simultaneously, leading to inaccurate spatial audio rendering.

Method used

By obtaining parameter values ​​associated with multiple audio signal sets and location, and combining them with the listener's location, modified parameter values ​​are generated. The audio signals are processed to generate spatial audio output. Microphone array and signal interpolation techniques are used to achieve 6-DOF audio rendering.

Benefits of technology

It improves the accuracy and quality of spatial audio rendering, especially in multi-source and complex audio scenes, reduces orientation errors, and supports free movement of listeners within the scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120897148A_ABST
    Figure CN120897148A_ABST
Patent Text Reader

Abstract

The invention relates to audio rendering with spatial metadata interpolation. There is provided an apparatus configured to obtain spatial audio streams, where the spatial audio streams each comprise at least one audio signal and at least one spatial metadata parameter, where the spatial audio streams are each associated with a location; obtaining positions associated with at least two of the spatial audio streams; obtaining a listener position configured to be tracked; generating at least one audio signal based at least in part on at least one of the spatial audio streams, the location, and the listener location; generating at least one modified spatial metadata parameter based at least in part on at least one spatial metadata parameter included by the at least two spatial audio streams, the location, and the listener location; and processing the at least one audio signal to generate a spatial audio output based at least in part on the at least one modified spatial metadata parameter.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese Invention Patent Application with the title of “Audio Rendering with Spatial Metadata Interpolation” (application number 202180016735.1, filing date 2021-02-03). TECHNICAL FIELD

[0002] The present application relates to an apparatus and method for audio rendering with spatial metadata interpolation, but not exclusively to an apparatus and method for audio rendering with spatial metadata interpolation for 6 degrees of freedom systems. BACKGROUND

[0003] Spatial audio capture methods attempt to capture an audio environment so that it can be perceptually recreated to a listener in an effective manner, and in addition can permit the listener to move and / or rotate within the recreated audio environment. For example, in certain systems (3 degrees of freedom - 3DoF), a listener can rotate their head and the rendered audio signals reflect this rotational movement. In certain systems (3 degrees of freedom plus - 3DoF+), a listener can move slightly within the environment as well as rotate their head, while in other systems (6 degrees of freedom - 6DoF), a listener can move arbitrarily within the environment and rotate their head.

[0004] Linear spatial audio capture refers to audio capture methods in which the characteristics of the captured audio are not processed. Instead, the output is a predetermined linear combination of the captured audio signals.

[0005] To linearly record spatial sound at one location in a space, a high-end microphone array is required. One such microphone is the spherical 32-microphone Eigenmike. From a high-end microphone array, a higher-order Ambisonics (HOA) signal can be obtained and used for linear rendering. With the HOA signal, spatial audio can be linearly rendered, separating sounds arriving from different directions satisfactorily within a reasonable auditory bandwidth.

[0006] One problem with linear spatial audio capture techniques is the requirement for a microphone array. Short wavelengths (higher frequency audio signals) require small inter-microphone spacing, while long wavelengths (lower frequencies) require large array size, both of which are difficult to satisfy simultaneously within a single microphone array.

[0007] Most practical capture devices (e.g. virtual reality cameras, single-lens reflex cameras, mobile phones) are not equipped with microphone arrays such as those provided by the Eigenmike, and do not have enough microphone devices to perform linear spatial audio capture. Furthermore, implementing linear spatial audio capture for a capture device results in capturing spatial audio for only a single location.

[0008] Parametric spatial audio capture involves a system that estimates perceptually relevant parameters based on audio signals captured by microphones, and that can synthesize spatial sound based on these parameters and the audio signals. Analysis and synthesis typically occur in a frequency band that is accessible to human spatial hearing resolution.

[0009] It is known that for most compact microphone setups (e.g. VR cameras, multi microphone arrays, mobile phones with microphones, SLR cameras with microphones), parametric spatial audio capture can yield perceptually accurate spatial audio rendering, while linear methods typically do not yield feasible results in terms of spatial aspects of sound. For high-end microphone arrays such as Eigenmike, parametric methods can also provide spatial sound perception that is on average of better quality than linear methods. SUMMARY

[0010] According to a first aspect, there is provided an apparatus comprising means configured to: obtain two or more sets of audio signals, wherein each set of audio signals is associated with a position; obtain at least one parameter value for at least two sets of audio signals of the sets of audio signals; obtain positions associated with at least the two sets of audio signals of the sets of audio signals; obtain a listener position; generate at least one audio signal based on at least one audio signal from at least one set of audio signals of the two or more sets of audio signals based on the positions associated with at least the two sets of audio signals of the sets of audio signals and the listener position; generate at least one modified parameter value based on the obtained at least one parameter value for at least the two sets of audio signals of the sets of audio signals, the positions associated with at least the two sets of audio signals of the sets of audio signals and the listener position; and process the at least one audio signal based on the at least one modified parameter value to generate a spatial audio output.

[0011] The means configured to obtain two or more sets of audio signals can be configured to: obtain the two or more sets of audio signals from microphone apparatuses, wherein each microphone apparatus is at a respective position and comprises one or more microphones.

[0012] Each set of audio signals can be associated with a direction, and the means can be further configured to: obtain directions of the two or more sets of audio signals, wherein the generated at least one audio signal can be further based on the directions associated with the two or more sets of audio signals, and wherein the at least one modified parameter value can be further based on the directions associated with the two or more sets of audio signals.

[0013] The component can be further configured to obtain a listener orientation, wherein the at least one modified parameter value can be further based on the listener orientation.

[0014] The component configured to process the at least one audio signal based on the at least one modified parameter value to generate the spatial audio output can be further configured to process the at least one audio signal further based on the listener orientation.

[0015] The component can be further configured to obtain a control parameter based on the positions associated with the at least two of the sets of audio signals and the listener position, wherein the component configured to generate the at least one audio signal based on the at least one audio signal from the at least one of the two or more sets of audio signals based on the positions associated with the at least two of the sets of audio signals and the listener position can be controlled based on the control parameter.

[0016] The component configured to generate the at least one modified parameter value can be controlled based on the control parameter.

[0017] The component configured to obtain the control parameter can be configured to identify at least three of the sets of audio signals within which the listener position of the set of audio signals is located and generate weights associated with the at least three of the sets of audio signals based on the set of audio signals position and the listener position; or else, identify two of the sets of audio signals of the set of audio signals closest to the listener position and generate weights associated with the two of the sets of audio signals of the set of audio signals based on a perpendicular projection of a straight line between the two of the sets of audio signals of the set of audio signals to the set of audio signals position and the listener position.

[0018] The component configured to generate the at least one audio signal can be configured to perform one of: combine two or more audio signals from two or more of the sets of audio signals based on the weights; select one or more audio signals from one of the two or more sets of audio signals based on which of the two or more sets of audio signals is closest to the listener position; and select one or more audio signals from one of the two or more sets of audio signals based on which of the two or more sets of audio signals is closest to the listener position and a further switching threshold.

[0019] The component configured to generate the at least one modified parameter value can be configured to combine the obtained at least one parameter value for the at least two of the sets of audio signals based on the weights.

[0020] The means configured to process the at least one audio signal based on the at least one modified parameter value to generate a spatial audio output can be configured to generate at least one of: a binaural audio output comprising two audio signals for headphones and / or earphones; and a multi-channel audio output comprising at least two audio signals for a multi-channel loudspeaker setup.

[0021] The at least one parameter value can comprise at least one of: at least one direction value; at least one direct-to-total ratio associated with the at least one direction value; at least one spread coherence associated with the at least one direction value; at least one distance associated with the at least one direction value; at least one surround coherence; at least one diffuse-to-total ratio; and at least one remainder-to-total ratio.

[0022] The at least two of the set of audio signals can comprise at least two audio signals, and the means configured to obtain the at least one parameter value can be configured to: spatially analyse the two or more audio signals from the two or more sets of audio signals to determine the at least one parameter value.

[0023] The means configured to obtain the at least one parameter value can be configured to: receive or retrieve the at least one parameter value for at least two of the set of audio signals.

[0024] According to a second aspect, there is provided a method for an apparatus, comprising: obtaining two or more sets of audio signals, wherein each set of audio signals is associated with a position; obtaining at least one parameter value for at least two of the sets of audio signals; obtaining positions associated with at least the two of the sets of audio signals; obtaining a listener position; generating at least one audio signal based on at least one audio signal from at least one of the two or more sets of audio signals based on the positions associated with at least the two of the sets of audio signals and the listener position; generating at least one modified parameter value based on the obtained at least one parameter value for at least two of the sets of audio signals, the positions associated with at least the two of the sets of audio signals and the listener position; and processing the at least one audio signal based on the at least one modified parameter value to generate a spatial audio output.

[0025] Obtaining the two or more sets of audio signals can comprise obtaining the two or more sets of audio signals from microphone devices, wherein each microphone device can be at a respective position and comprise one or more microphones.

[0026] Each set of audio signals can be associated with a direction, and the method can further comprise obtaining the directions of the two or more sets of audio signals, wherein the generated at least one audio signal can be further based on the directions associated with the two or more sets of audio signals, and wherein the at least one modified parameter value can be further based on the directions associated with the two or more sets of audio signals.

[0027] The method can further comprise obtaining a listener direction, wherein the at least one modified parameter value can be further based on the listener direction.

[0028] Processing the at least one audio signal based on the at least one modified parameter value to generate a spatial audio output can further comprise processing the at least one audio signal further based on the listener direction.

[0029] The method can further comprise obtaining a control parameter based on the positions associated with at least two of the sets of audio signals and the listener position, wherein the operation of generating the at least one audio signal based on the at least one audio signal from at least one of the two or more sets of audio signals based on the positions associated with the at least two of the sets of audio signals and the listener position can be controlled based on the control parameter.

[0030] The operation of generating the at least one modified parameter value can be controlled based on the control parameter.

[0031] Obtaining the control parameter can comprise identifying at least three of the sets of audio signals within which the listener position is located and generating weights associated with the at least three of the sets of audio signals based on the sets of audio signals positions and the listener position, or else identifying two of the sets of audio signals closest to the listener position and generating weights associated with the two of the sets of audio signals based on a perpendicular projection of a straight line between the two of the sets of audio signals and the listener position to the sets of audio signals positions.

[0032] Generating the at least one audio signal can comprise one of: combining two or more audio signals from two or more of the sets of audio signals based on the weights; selecting one or more audio signals from one of the two or more sets of audio signals based on which of the two or more sets of audio signals is closest to the listener position; and selecting one or more audio signals from one of the two or more sets of audio signals based on which of the two or more sets of audio signals is closest to the listener position and a further switching threshold.

[0033] Generating the at least one modified parameter value can comprise combining the obtained at least one parameter value for at least two of the sets of audio signals based on the weights.

[0034] Processing the at least one audio signal based on the at least one modified parameter value to generate a spatial audio output can comprise generating at least one of: a binaural audio output comprising two audio signals for headphones and / or earphones; and a multi-channel audio output comprising at least two audio signals for a multi-channel loudspeaker setup.

[0035] The at least one parameter value can comprise at least one of: at least one direction value; at least one direct-to-total ratio associated with the at least one direction value; at least one spread coherence associated with the at least one direction value; at least one distance associated with the at least one direction value; at least one surround coherence; at least one diffuse-to-total ratio; and at least one residual-to-total ratio.

[0036] The at least two of the sets of audio signals can comprise at least two audio signals, and obtaining the at least one parameter value can comprise spatially analysing the two or more audio signals from the two or more sets of audio signals to determine the at least one parameter value.

[0037] Obtaining the at least one parameter value can comprise receiving or retrieving the at least one parameter value for at least two of the sets of audio signals.

[0038] According to a third aspect, there is provided an apparatus comprising at least one processor and at least one memory including a computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to obtain two or more sets of audio signals, wherein each set of audio signals is associated with a position; obtain at least one parameter value for at least two sets of audio signals of the sets of audio signals; obtain positions associated with at least the two sets of audio signals of the sets of audio signals; obtain a listener position; generate, based on at least one audio signal from at least one set of audio signals of the two or more sets of audio signals, at least one audio signal based on at least the positions associated with the at least two sets of audio signals of the sets of audio signals and the listener position; generate at least one modified parameter value based on the obtained at least one parameter value for the at least two sets of audio signals of the sets of audio signals, the positions associated with the at least two sets of audio signals of the sets of audio signals and the listener position; and process the at least one audio signal based on the at least one modified parameter value to generate a spatial audio output.

[0039] The apparatus caused to obtain the two or more sets of audio signals can be further caused to obtain the two or more sets of audio signals from microphone apparatuses, wherein each microphone apparatus is at a respective position and comprises one or more microphones.

[0040] Each set of audio signals can be associated with a direction, and the apparatus can be further caused to obtain directions of the two or more sets of audio signals, wherein the generated at least one audio signal can be further based on the directions associated with the two or more sets of audio signals, and wherein the at least one modified parameter value can be further based on the directions associated with the two or more sets of audio signals.

[0041] The apparatus can be further caused to obtain a listener direction, wherein the at least one modified parameter value can be further based on the listener direction.

[0042] The apparatus caused to process the at least one audio signal based on the at least one modified parameter value to generate the spatial audio output can be further caused to process the at least one audio signal further based on the listener direction.

[0043] The apparatus can be further caused to obtain a control parameter based on the positions associated with the at least two sets of audio signals of the sets of audio signals and the listener position, wherein the apparatus caused to generate the at least one audio signal based on at least one audio signal from at least one set of audio signals of the two or more sets of audio signals based on at least the positions associated with the at least two sets of audio signals of the sets of audio signals and the listener position can be controlled based on the control parameter.

[0044] The apparatus caused to generate at least one modified parameter value can be caused to combine at least one parameter value obtained for at least two of the sets of audio signals based on the weights.

[0045] The apparatus caused to obtain the control parameters can be further caused to: identify at least three of the sets of audio signals within which a listener position in the set of audio signals is located, and generate weights associated with the at least three of the sets of audio signals based on the set of audio signals position and the listener position; or else, identify two of the sets of audio signals closest to the listener position in the set of audio signals, and generate weights associated with the two of the sets of audio signals in the set of audio signals based on a perpendicular projection of a straight line between the two of the sets of audio signals and the listener position in the set of audio signals.

[0046] The apparatus caused to generate at least one audio signal can be caused to perform one of: combine two or more audio signals from two or more of the sets of audio signals based on the weights; select one or more audio signals from one of the two or more sets of audio signals based on which of the two or more sets of audio signals is closest to the listener position; and select one or more audio signals from one of the two or more sets of audio signals based on which of the two or more sets of audio signals is closest to the listener position and a further switching threshold.

[0047] The apparatus caused to generate at least one modified parameter value can be caused to combine at least one parameter value obtained for at least two of the sets of audio signals based on the weights.

[0048] The apparatus caused to process at least one audio signal based on at least one modified parameter value to generate a spatial audio output can be caused to generate at least one of: a binaural audio output comprising two audio signals for headphones and / or earphones; and a multi-channel audio output comprising at least two audio signals for a multi-channel loudspeaker setup.

[0049] The at least one parameter value can comprise at least one of: at least one direction value; at least one direct-to-total ratio associated with the at least one direction value; at least one early-to-reverberant coherence associated with the at least one direction value; at least one distance associated with the at least one direction value; at least one surround coherence; at least one diffuse-to-total ratio; and at least one residual-to-total ratio.

[0050] The at least two of the sets of audio signals can comprise at least two audio signals, and obtaining at least one parameter value can comprise spatially analysing the two or more audio signals from the two or more sets of audio signals to determine the at least one parameter value.

[0051] The apparatus caused to obtain at least one parameter value can be caused to receive or retrieve at least one parameter value for at least two of the sets of audio signals.

[0052] According to a fourth aspect, there is provided an apparatus comprising: means for obtaining two or more sets of audio signals, wherein each set of audio signals is associated with a position; means for obtaining at least one parameter value for at least two of the sets of audio signals; means for obtaining positions associated with at least the two of the sets of audio signals; means for obtaining a listener position; means for generating at least one audio signal based on at least one audio signal from at least one of the two or more sets of audio signals based on the positions associated with at least the two of the sets of audio signals and the listener position; means for generating at least one modified parameter value based on the obtained at least one parameter value for at least the two of the sets of audio signals, the positions associated with at least the two of the sets of audio signals, and the listener position; and means for processing the at least one audio signal to generate a spatial audio output based on the at least one modified parameter value.

[0053] According to a fifth aspect, there is provided a computer program comprising instructions [or a computer readable medium comprising program instructions] for causing an apparatus to perform at least the following: obtaining two or more sets of audio signals, wherein each set of audio signals is associated with a position; obtaining at least one parameter value for at least two of the sets of audio signals; obtaining positions associated with at least the two of the sets of audio signals; obtaining a listener position; generating at least one audio signal based on at least one audio signal from at least one of the two or more sets of audio signals based on the positions associated with at least the two of the sets of audio signals and the listener position; generating at least one modified parameter value based on the obtained at least one parameter value for at least the two of the sets of audio signals, the positions associated with at least the two of the sets of audio signals, and the listener position; and processing the at least one audio signal to generate a spatial audio output based on the at least one modified parameter value.

[0054] According to a sixth aspect, there is provided a non-transitory computer- readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining two or more sets of audio signals, wherein each set of audio signals is associated with a position; obtaining at least one parameter value for at least two sets of audio signals of the sets of audio signals; obtaining positions associated with at least the two sets of audio signals of the sets of audio signals; obtaining a listener position; generating at least one audio signal based on at least one audio signal from at least one set of audio signals of the two or more sets of audio signals based on the positions associated with at least the two sets of audio signals of the sets of audio signals and the listener position; generating at least one modified parameter value based on the obtained at least one parameter value for at least the two sets of audio signals of the sets of audio signals, the positions associated with at least the two sets of audio signals of the sets of audio signals, and the listener position; and processing the at least one audio signal to generate a spatial audio output based on the at least one modified parameter value.

[0055] According to a seventh aspect, there is provided an apparatus comprising: obtaining circuitry configured to obtain two or more sets of audio signals, wherein each set of audio signals is associated with a position; obtaining circuitry configured to obtain at least one parameter value for at least two sets of audio signals of the sets of audio signals; obtaining circuitry configured to obtain positions associated with at least the two sets of audio signals of the sets of audio signals; obtaining circuitry configured to obtain a listener position; generating circuitry configured to generate at least one audio signal based on at least one audio signal from at least one set of audio signals of the two or more sets of audio signals based on the positions associated with at least the two sets of audio signals of the sets of audio signals and the listener position; generating circuitry configured to generate at least one modified parameter value based on the obtained at least one parameter value for at least the two sets of audio signals of the sets of audio signals, the positions associated with at least the two sets of audio signals of the sets of audio signals, and the listener position; and processing circuitry configured to process the at least one audio signal to generate a spatial audio output based on the at least one modified parameter value.

[0056] According to an eighth aspect, there is provided a computer readable medium comprising program instructions for causing an apparatus to perform at least the following: obtaining two or more sets of audio signals, wherein each set of audio signals is associated with a position; obtaining at least one parameter value for at least two sets of audio signals of the sets of audio signals; obtaining positions associated with at least the two sets of audio signals of the sets of audio signals; obtaining a listener position; generating at least one audio signal based on at least one audio signal from at least one set of audio signals of the two or more sets of audio signals based on the positions associated with at least the two sets of audio signals of the sets of audio signals and the listener position; generating at least one modified parameter value based on the obtained at least one parameter value for at least the two sets of audio signals of the sets of audio signals, the positions associated with at least the two sets of audio signals of the sets of audio signals and the listener position; and processing the at least one audio signal based on the at least one modified parameter value to generate a spatial audio output.

[0057] An apparatus comprising means for performing the acts of the method as described above.

[0058] An apparatus configured to perform the acts of the method as described above.

[0059] A computer program comprising program instructions for causing a computer to perform the method as described above.

[0060] A computer program product stored on a medium can cause an apparatus to perform the method described herein.

[0061] An electronic device can comprise an apparatus as described herein.

[0062] A chipset can comprise an apparatus as described herein.

[0063] Embodiments of the present application aim to address problems associated with the prior art. BRIEF DESCRIPTION OF DRAWINGS

[0064] For a better understanding of the present application, reference will now be made, by way of example, to the accompanying drawings in which:

[0065] Figure 1 A system of apparatuses suitable for implementing some embodiments is schematically illustrated;

[0066] Figure 2 And Figure 3 A system of apparatuses illustrating the effect of distance error on rendering is schematically illustrated;

[0067] Figure 4 An overview of some embodiments relating to capture and rendering of spatial metadata is illustrated;

[0068] Figure 5 schematically illustrating a suitable apparatus for implementing interpolation of audio signals and metadata according to some embodiments;

[0069] Figure 6 schematically illustrating a flowchart of the operation of the apparatus shown in Figure 5

[0070] Figure 7 schematically illustrating source positions inside and outside of an array configuration;

[0071] Figure 8 schematically illustrating a synthesis processor as shown in Figure 5

[0072] Figure 9 schematically illustrating a flowchart of the operation of the synthesis processor shown in Figure 5

[0073] Figure 10 schematically illustrating a suitable apparatus for implementing interpolation of audio signals and metadata according to some embodiments;

[0074] Figure 11 schematically illustrating a flowchart of the operation of the apparatus shown in Figure 5

[0075] Figure 12 schematically illustrating a further view of a suitable apparatus for implementing interpolation of audio signals and metadata according to some embodiments; and

[0076] Figure 13 schematically illustrating an example device suitable for implementing the illustrated apparatus. DETAILED DESCRIPTION

[0077] The concepts as discussed in further detail herein with respect to the following embodiments relate to parametric spatial audio capture with two or more microphone arrays corresponding to different positions in a recording space and enabling a user to move to different positions in the captured sound scene, in other words, the invention relates to 6DoF audio capture and rendering.

[0078] 6DoF is currently common in virtual reality, such as VR games, where movement in the audio scene is directly rendered, as all spatial information is readily available (i.e. position of each sound source and audio signal of each sound source, respectively). The invention relates to also providing robust 6DoF capture and rendering of spatial audio captured with microphone arrays.

[0079] ​​​​For example for the upcoming MPEG-I Audio standard, 6DoF capture and rendering from microphone arrays is relevant, where 6DoF rendering of HOA signals is required. These HOA signals can be obtained from a microphone array at the sound scene.

[0080] In the following examples, an audio signal set is generated by a microphone. For example, a microphone device can comprise one or more microphones and generate one or more audio signals for the audio signal set. In some embodiments, the audio signal set comprises audio signals which are virtual or generated audio signals (e.g. virtual loudspeaker audio signals with associated virtual loudspeaker positions).

[0081] Before discussing this concept in more detail, we will first describe some aspects of spatial capture and reproduction in more detail. For example, with regard to Figure 1 , an example of spatial capture and playback is shown. Thus, for example, Figure 1 On the left side, a spatial audio signal capture environment is shown. The environment or audio scene comprises sources, i.e. source 1 202 and source 2 204, which can be actual audio signal sources or can be abstract audio source representations. In addition, a non-directional or non-specific location environment portion 206 is shown. These can be captured by at least two microphone devices / arrays, each of which can comprise two or more microphones.

[0082] The audio signals can be captured as described above, in addition, they can be encoded, transmitted, received and reproduced as shown by arrows 210 in Figure 1 .

[0083] On the right side of Figure 1 , an example reproduction is shown. The reproduction of the spatial audio signals results in a user 250 (which is shown in this example wearing a head tracking headset) being presented with a reproduced audio environment in the form of a 6DoF spatial rendering 218, which comprises a perceived source 1 212, a perceived source 2 214 and a perceived environment 216.

[0084] As discussed above, traditional linear and parametric spatial audio capture methods for microphone arrays can be used for high quality spatial audio processing, depending on the available microphone devices. However, they are all developed for single position capture and rendering. In other words, the listener cannot move between the microphone arrays. Thus, they cannot be directly applied to 6DOF rendering, where the listener can move arbitrarily between the microphone arrays.

[0085] Embodiments as discussed herein aim to provide broadband 6DOF rendering methods. These methods aim to improve known parametric rendering from microphone arrays. For example, they aim to improve methods in which distance parameters (in addition to direction parameters) are estimated in the frequency band, in other words, in which sound positions are estimated for 6DOF rendering. This improvement relates to the characteristic that sound source distance or position cannot be reliably estimated in all acoustic situations, and in which errors in distance / position estimation generate significant errors in 6DOF playback. This effect is noticeable when the listener’s movement relative to the capture position is significant (e.g. more than 1 meter in any direction).

[0086] With regard to Figure 2 and Figure 3 , a situation with multiple sources is shown. For example, Figure 2 An ideal capture situation is shown. The capture position 306 is shown, and the black dots 301, 303, 305, 307 show the estimated direction and distance for each time-frequency tile. As shown in the figure, when multiple sound sources are active at the same time, the direction parameter at parametric capture does not necessarily point to any of the sound sources, but can point somewhere in between the sound sources. This is not a problem for parametric capture systems, as the such perceptual / dominant direction is a good approximation of the sound situation in perceptual sense. However, as a special and ideal aspect of Figure 2 , the distance is also well estimated. Thus, no matter the listening position 310, the (perceptual / dominant) direction is reproduced at an arc 308 (shown by the dashed line) between the source directions (source 1 302 and source 2 304).

[0087] However, Figure 3 Another example of the same arrangement in a multi-source situation in which the distance estimation is noisy is shown, which is a more realistic example in such multi-source situations. This distance estimation noise results in false estimated positions 321, 323, 325, 327. If the sound is rendered at the listening position 306, this distance estimation does not result in significant directional errors. However, when the sound is rendered at a significantly different listening position 310, then the rendering of the sound direction has a large spatial error. The (perceptual / dominant) direction is reproduced at an arc 318 (shown by the dashed line) that significantly straddles the source directions (source 1 302 and source 2 304). Thus, the spatial reproduction in this example is “expanded” more compared to the “ideal” arc 308 (shown by the dashed line) shown in Figure 2

[0088] ​As a result of the incorrect estimated distance of the rendered audio by the listener in the "full" 6DOF rendering, which can arbitrarily move at the time the user is in the capture position 306 (and not just close to the microphone array position), the sound direction is properly rendered, as the fake distance does not affect the rendered direction. At each time-frequency tile, the perceptual / dominant direction is rendered at the arc line determined by the two simultaneous sources. However, when the user moves to the illustrated 6DOF listening position 310, the influence of the fake distance estimation becomes apparent. At this position, the rendered sound direction is not between the two sources. In other words, the result is a broad and diffuse spatial rendering output (as opposed to an accurate and point-like perception of the sources), in which spatial artifacts can even occasionally appear far from the actual source directions.

[0089] Embodiments therefore seek to provide proper 6DOF audio capture and rendering from a microphone array in which there are multiple sound sources and / or the listener can arbitrarily move.

[0090] While the perception-related parameters can be any suitable parameters, the following examples discussed herein obtain the following set of parameters:

[0091] At least one direction parameter in a frequency band, which indicates the predominant (or dominant or perceptual) direction from which sound arrives, and

[0092] A ratio parameter, which indicates how much energy is arriving from these directions and how much sound energy is ambient / ambient.

[0093] As discussed above, there are different ways to obtain these parameters. One known method is Directional Audio Coding (DirAC), in which, based on a first-order Ambisonic signal (or B-format signal), the direction and diffuseness (i.e., ambient-to-total energy ratio) parameters are estimated in a frequency band. In the following examples, DirAC is used as the main example of parameter generation, although it is known that it can be replaced by other methods to obtain spatial parameters or spatial metadata, such as the higher-order DirAC, high-angle plane wave expansion, and Nokia's Spatial Audio Capture (SPAC) discussed in PCT application WO2018 / 091776.

[0094] The described embodiments aim at good quality spatial sound reproduction with clear identifiable sources and requiring position tracking for audio scenes that are more complex. For example, in outdoor environments there are often many sources active at the same time. When there are multiple sources (more sources than directional parameters), the directional parameters are no longer a physical descriptor pointing to the source, but a perceptual descriptor. This means that, for example, if there are two sources, the directional parameters will typically fluctuate in the area between the two sources according to the source energy in the time-frequency bin. This leads to situations where the distance estimation can fail, as shown in Figure 3 For example, the fluctuation of the directional parameters or the ratio parameters can be used to estimate the distance, as these characteristics are influenced by the room reverberation and the source distance. However, when doing so, the distance parameter is artificially made large, as a certain fluctuation or ratio is not due to the source distance (reverberation) but due to a simultaneous source. Furthermore, if a visual depth map is used for the distance estimation, the fluctuation direction typically does not correspond to the actual source direction, so that the distance is estimated wrongly. The distance can also be estimated from two arrays and finding the intersection of the projected rays from these two arrays towards the estimated direction. However, the fluctuation direction due to a complex sound scene provides a very noisy intersection and thus a noisy distance estimation.

[0095] In other words, the embodiments aim at producing low error parameter estimations in complex audio scenes, as these parameter estimation errors tend to lead to spatial errors at the sound of the 6DOF reproduction. Furthermore, in some embodiments, a 6DOF rendering is provided that does not rely on distance estimations and thus also provides higher robustness for complex situations. The embodiments can interpolate spatial metadata to positions between the actual capture positions.

[0096] Thus, embodiments as discussed herein can involve 6 degrees of freedom (i.e., a listener can move within the scene and the listener position is tracked) binaural rendering of audio captured with at least two microphone arrays in known positions. These embodiments can also provide high quality binaural audio rendering under a wide range of (6DOF tracked) listener positions and sound field conditions, especially improving situations where multiple sources are active at the same time and the listener is not near the array. The embodiments can also determine spatial metadata for array positions using corresponding microphone array signals, predict spatial metadata for listener positions using the determined spatial metadata (based on listener and array positions), determine a selection or mix of array signals (based on listener and array positions), and parametrically render spatial audio output based on the predicted spatial metadata and the determined selection or mix of array signals.

[0097] In some embodiments, the apparatus and method can be further configured such that the determined selection or mixing of the array signal refers to the signal from the closest array, and when the user moves to a position closer (by a threshold) to another array position than the previously closest array, the selection or mixing of the array signal is changed such that the binaural audio signal is rendered based on the audio signal and the predicted spatial metadata from the other array.

[0098] In some embodiments, the array signal can refer to a microphone array signal, or a signal based on a microphone array signal, such as an array signal converted to an Ambisonic format.

[0099] Figure 4 An example system in which embodiments can be implemented is shown in Fig. Figure 4 A system with audio components, i.e. source 1 400, source 2 402 and an environment 410 is shown. In addition, within the system there are capturing apparatuses 401, 403 and 405, which are located at capturing positions within the environment and configured to capture audio signals and obtain or determine spatial metadata from these audio signals (at 404).

[0100] The system further comprises a listener (user) apparatus 407 configured to generate a suitable binaural audio signal. Thus, in some embodiments, the apparatus 407 is configured to determine rendering metadata at the user position (at 406) based on the spatial metadata and the user position (relative to the capturing positions). Furthermore, the apparatus 407 is configured to perform binaural rendering (at 408) using the rendering metadata and the audio signals from at least one microphone array (which can be the closest).

[0101] Thus, even in the case of multiple simultaneous sound sources and even for listening positions not close to the capturing apparatus microphone array positions, embodiments can produce good audio quality. These embodiments omit the use of distance metadata, which is indicated to be unreliable in the case of multiple simultaneous sources and to lead to directional errors when rendering spatial audio in positions far from the microphone array positions. Instead, embodiments show a direct prediction of the direction in the frequency band based on the direction determined at the microphone position (and direct-to-total energy ratio) for the listening position. Since the estimated direction (and direct-to-total energy ratio) is more reliable, the directional error produced by some embodiments is significantly reduced and better audio quality is produced.

[0102] With regard to Figure 5 An example system is shown in Fig. In some embodiments, the system can be implemented on a single apparatus. However, in some other embodiments, the functionality described herein can be implemented on more than one apparatus.

[0103] In some embodiments, the system comprises an input configured to receive a plurality of signal sets based on the microphone array signals 500. The plurality of signal sets based on the microphone array signals can comprise J sets of multi-channel signals. The signals can be the microphone array signals themselves, or some transformed form of the array signals, such as Ambisonic signals. These signals are denoted as s j (m, i), where j is the index of the microphone array from which the signal originates (i.e. the signal set index), m is the sample time, and i is the channel index of the signal set.

[0104] The plurality of signal sets can be passed to a signal interpolator 503 and a spatial analyzer 501.

[0105] In some embodiments, the system comprises a spatial analyzer 501. The spatial analyzer 501 is configured to receive the audio signals s j (m, i), and analyze these audio signals to determine spatial metadata in the time- frequency domain for each array.

[0106] The spatial analysis can be based on any suitable technique, and there are known suitable methods for various input types. For example, if the input signals are in Ambisonic or Ambisonic related form (e.g. they originate from B-format microphones), or the array can be transformed into Ambisonic form in a reasonable way (e.g. Eigenmike), then a Directional Audio Coding (DirAC) analysis can be performed. First order DirAC has been described in Pulkki, Ville, “Spatial sound reproduction with directional audio coding,” Journal of the Audio Engineering Society, Vol. 55, No. 6 (2007): 503-516, where a method for estimating a set of spatial metadata comprising direction and ambient-to-total energy ratio parameters in a frequency band from B-format signals (a variant of first order Ambisonics) is specified.

[0107] Sector-based parametric sound field reproduction in the spherical harmonic domain”, IEEE Selected Topics in Signal Processing, vol. 9, no. 5 (2015): 852-866, by Archontis Politis, Juha Vilkamo and Ville Pulkki, provides a method for obtaining multiple simultaneous directional parameters. Further methods that can be implemented in some embodiments include estimation of spatial metadata from planar devices such as mobile phones and tablets as described in PCT published patent application WO2018 / 091776, and similar delay-based analysis methods for non-planar devices (GB published patent application GB2572368).

[0108] In other words, there are various methods to obtain spatial metadata, and the chosen method can depend on the array type and / or the audio signal format. In some embodiments, one method is applied for one frequency range, while another method is applied for another frequency range. In the following examples, the analysis is based on receiving first-order Ambisonic (FOA) audio signals, which is a well-known signal format in the field of spatial audio. Furthermore, in these examples, a modified DirAC method is used. For example, the input is an Ambisonic audio signal in the form of known SN3D normalisation (Schmidt semi-normalisation) and ACN (Ambisonics Channel Number) channel ordering.

[0109] In some embodiments, the spatial analyser is configured to perform the following operations for each microphone array:

[0110] 1) First, the input signal s j (m,i) is converted into a time-frequency domain format signal. For example, this conversion can be implemented using a short-time Fourier transform (STFT) or a complex modulated quadrature mirror filter (QMF) bank. As an example, the STFT is a process that is typically configured such that for a frame length of N samples, the current frame and the previous frame are windowed (e.g. with a sine window) and processed with a fast Fourier transform (FFT). The result is denoted as S j (b,n,i) of the time-frequency domain signal, where b is the frequency bin and n is the time frame index. The time-frequency signal (in this case a 4-channel FOA signal) is grouped in vector form by:

[0111]

[0112] 2) Next, these time-frequency signals are used in frequency bands. A frequency bin denotes a single complex sample in the STFT domain, while a frequency band denotes a group of these bins. Denote the frequency band index by k = 1..K, and K is the number of frequency bands, each frequency band k has a lowest bin b k,low and a highest bin b k,hign In some embodiments, the signal covariance matrix is estimated in the frequency bands by

[0113]

[0114] In some embodiments, a temporal smoothing over time index n can be applied.

[0115] 3) Further, the inverse sound field intensity vector pointing in the opposite direction of the propagating sound is determined:

[0116]

[0117] Note the channel order, which converts the ACN order to the Cartesian x, y, z order.

[0118] 4) Further, the directional parameter for frequency band k and time index n is determined as the direction of i j (k, n). This directional parameter can for example be expressed as the azimuth angle θ j (k, n) and the elevation angle

[0119] 5) Further, the direct-to-total energy ratio is determined as:

[0120]

[0121] For each frequency band k, for each time index n, and for each signal set (each array) j, the azimuth angle θ j (k, n), the elevation angle and the direct-to-total energy ratio r j (k, n) are determined. This information thus constitutes the metadata 506 for each array, which is output from the spatial analyzer to the metadata interpolator 507.

[0122] In some embodiments, the system further comprises a position pre-processor 505. The position pre-processor 505 is configured to receive information about the microphone array position 502 and the listener position 504 within the audio environment.

[0123] As known in the art, a key goal of parametric spatial audio capture and rendering is to obtain a perceptually accurate spatial audio reproduction for the listener. Thus, the position pre-processor 505 is configured to be able to determine interpolation data for any position (as the listener can move to any position) to allow the metadata to be modified based on the microphone array position 502 and the listener position 504.

[0124] In the examples herein, the microphone arrays are located in one plane. In other words, these arrays have no z-axis displacement component. However, embodiments can be extended to the z-axis in some embodiments, as well as to the case where the microphone arrays are located on a line (in other words, only one axis displacement).

[0125] For example, Figure 7 A microphone apparatus is shown in which the microphone arrays (shown as circular array 1 701, array 2 703, array 3 705, array 4 707 and array 5 709) are located in one plane. Spatial metadata has been determined at the array positions. The apparatus has five microphone arrays in one plane. This plane can be divided into interpolation triangles, for example by Delaunay triangulation. When a user moves to a position within a triangle (e.g. position 1 711), the three microphone arrays forming the triangle containing the position (array 1 701, array 3 705 and array 4 707 in this example case) are selected for interpolation. When the user moves outside the area spanned by these microphone arrays (e.g. position 2 713), the user position is projected to the nearest position at the area spanned by these microphone arrays (e.g. projected position 2 714), in turn where the array-triangle in which the projected position resides is selected for interpolation (in this example, these arrays are array 2 703, array 3 705 and array 5 709). When a position is projected, the projected position overrides the original listener position parameter.

[0126] In the above example, the projection of a position thus maps a position outside the area determined by the microphone apparatus to the edge of the area determined by the microphone apparatus. While this can seem like a limitation, in fact, when considering 6DOF media capture and reproduction, the audio is accompanied by video obtained from VR camera sets that enable 6DOF video reproduction. The area spanned by the VR cameras (as video also needs to be made) is expected to also limit the area in which the user can move in the scene, and further expected that each VR camera also includes a microphone apparatus. Thus, the most important interpolation area is within the area spanned by the microphone arrays. Thus, the projection illustrates that the method does not completely fail outside the determined area. The nearest projected position is a reasonable approximation of the soundfield characteristics at a position slightly outside the area spanned by the microphone apparatus.

[0127] Thus, the position pre-processor 505 can determine:

[0128] a listener position vector p L (in this example a 2x1 vector containing x and y coordinates), which can be the original position or a projected position;

[0129] three microphone device indices j1, j2, j3 and corresponding position vectors These three microphone devices are those microphone devices that encapsulate the position p L .

[0130] The position pre-processor 505 can further formulate interpolation weights w1, w2, w3. These weights can for example be expressed using the known conversion between barycentric coordinates and Cartesian coordinates. First, based on the position vector A 3x3 matrix is determined by appending each vector with a unity value and combining the resulting vectors into a matrix:

[0131]

[0132] Further, the weights are expressed using the matrix inverse and a 3x1 vector obtained by appending the listener position vector p L with unity values:

[0133]

[0134] The interpolation weights (w1, w2, and w3), the position vectors (p L , and ), and the microphone device indices (j1, j2, and j3) together constitute interpolation data 508 and 510, which are provided to the signal interpolator 503 and the metadata interpolator 507.

[0135] In some embodiments, the system comprises a metadata interpolator 507, which is configured to receive the interpolation data 508 and the metadata 506 for each array. The metadata interpolator is further configured to interpolate the metadata using the interpolation weights w1, w2, w3. In some embodiments, this can be achieved by first converting the spatial metadata to vector form:

[0136]

[0137] Further, these vectors are averaged by:

[0138]

[0139] Further, the indices are denoted as:

[0140] v(k, n) = [v1(k, n) v2(k, n) v3(k, n)] T

[0141] The interpolated metadata is obtained by:

[0142] Θ'(k, n) = atan2(v2(k, n), v1(k, n))

[0143]

[0144] Further, the interpolated metadata 514 is output to the synthesis processor 509.

[0145] In the above, an example of metadata interpolation is presented. In other embodiments, other interpolation rules can be devised and implemented. For example, the interpolation ratio parameters can also be determined as a weighted average (according to w1, w2, w3) of the input ratios. Further, in some embodiments, the averaging can also involve a weighting according to the energy of the array signals.

[0146] In some embodiments, the system further comprises a signal interpolator 503. The signal interpolator is configured to receive the input audio signals 500 and the interpolated data 510. In some embodiments, the signal interpolator 503 can first convert the input signals into signals in the time-frequency domain in the same way as in the spatial analyzer 501. In some embodiments, the signal interpolator 503 is configured to receive the time-frequency audio signals directly from the spatial analyzer 501.

[0147] Further, the signal interpolator 503 can be configured to determine the total energy for each signal and for each frequency band. In the example shown herein, the signals are in the form of FOA signals, and thus the total energy can be determined as E j (k, n) = c 1,1,j (k, n). This value can be formulated in the same way as in the spatial analyzer 501 (or obtained from the spatial analyzer 501).

[0148] Further, the signal interpolator 503 can be configured to determine distance values for the indices j1, j2, j3 minD .

[0149] Further, the signal interpolator 503 is configured to determine the selected index j sel . For the first frame (or, when the processing starts), the signal interpolator can set j sel = j minD .

[0150] For the next or subsequent frame (or any desired temporal resolution), the signal interpolator is configured to determine the selection j when the user's position may have changed. sel Does it need to be changed? If j sel If a value is not contained within j1, j2, or j3, then a change is needed. This situation means the user has moved to another instance that does not contain j1, j2, or j3. sel The area. If This also needs to be changed, where α is the threshold. For example, α = 1.2. This situation means that with j sel Compared to the array position, the user is now significantly closer to j. minD The array position. This threshold is needed so that the selection does not change irregularly back and forth when the user is in between these two positions (in other words, a hysteresis threshold is provided to prevent rapid switching between arrays).

[0151] If any of the above conditions are met, then j sel =j minD Otherwise, keep j. sel The previous value.

[0152] The intermediate interpolation signal is determined as follows:

[0153]

[0154] Using this processing, when j sel When changes occur, the selection is changed simultaneously across all frequency bands. In some embodiments, the selection is set to change in a frequency-dependent manner. For example, when j sel When changes occur, some frequency bands are updated immediately, while others are changed in subsequent frames until all bands have been changed. It may be necessary to modify the signal in this frequency-dependent manner to reduce variations in signal S′. interp Potential switching artifacts at (b,n,i). In this configuration, when a switch occurs, the signal S′ may exhibit artifacts within a very short transition period. interp Some frequencies of (b,n,i) come from one microphone array, while others come from another.

[0155] Furthermore, for the intermediate interpolated signal S′ interp (b,n,i) performs energy correction. The equalization gain is defined in the frequency band as:

[0156]

[0157] g max Values ​​limit over-amplification; for example, g max =4. Therefore, balance is achieved through multiplication:

[0158] S(b, n, i) = g(k, n) S' interp (b, n, i)

[0159] where k is the frequency band index at which bin b is located. The signal S(b, n, i) is then the interpolated signal 512 that is output to the synthesis processor.

[0160] The system also includes a synthesis processor 509. The synthesis processor can be configured to receive listener orientation information 516 (e.g., head orientation tracking information) as well as the interpolated signal 512 and the interpolated metadata 514.

[0161] In some embodiments, the synthesis processor is configured to determine a vector rotation function to use in the following formula. According to the principles in Laitinen, M. V., “Binaural reproduction for directional audio coding,” Master’s Thesis, Helsinki University of Technology, pp. 54-55, 2008, the rotation function can be defined as:

[0162]

[0163] where yaw, pitch, and roll are the head orientation parameters, and x, y, z are the values of the unit vector being rotated. The result is x', y', z', which is the rotated unit vector. The mapping function performs the following steps:

[0164] 1. Yaw rotation

[0165] x1 = cos(yaw)x + sin(yaw)y

[0166] y1 = -sin(yaw)x + cos(yaw)y

[0167] z1 = z

[0168] 2. Pitch rotation

[0169]

[0170] y2 = y1

[0171]

[0172] 3. Finally, roll rotation

[0173] x' = x2

[0174]

[0175] Having determined these parameters, the synthesis processor 509 can implement any suitable spatial rendering. For example, in some embodiments, the synthesis processor 509 can implement 3DOF rendering, e.g. according to the principles described in PCT publication WO2019086757. In such embodiments, the parametric audio signal (audio and spatial metadata) can be rendered to binaural, Ambisonic, or surround speaker formats 518.

[0176] With reference to Figure 6 , a flowchart illustrating Figure 5 the operation of the method is shown.

[0177] Thus, as shown in step 601 of Figure 6 , in some embodiments, a plurality of signal sets based on the microphone array signal can be obtained.

[0178] As shown in step 603 of Figure 6 , having obtained the plurality of signal sets, a spatial analysis can be performed on each array.

[0179] As shown in step 602 of Figure 6 , the microphone array positions can also be obtained.

[0180] Further, as shown in step 610 of Figure 6 , the listener position / orientation can be obtained.

[0181] As shown in step 604 of Figure 6 , having obtained the microphone array positions and listener orientation / position, the method can obtain interpolation factors by processing the relative positions.

[0182] Having obtained the interpolation factors by processing the relative positions and signals / metadata, the method can interpolate the signals as shown in step 606 of Figure 6 , and interpolate the metadata as shown in step 605 of Figure 6 .

[0183] As shown in step 611 of Figure 6 , having determined the interpolated metadata and signals and listener orientation / position, the method can apply a synthesis process.

[0184] As shown in step 613 of Figure 6 , the spatialized audio is output.

[0185] The synthesis processor 509 is shown in more detail in Figure 8 .

[0186] In some embodiments, the synthesis processor 509 comprises a prototype signal generator 801. In some embodiments, the prototype signal generator 801 is configured to receive the interpolated signal 512 (which is received in the time-frequency domain) as well as the head (user / listener) orientation information 516.

[0187] The prototype signal is at least partially similar to the processed output, and thus it can serve as a good starting point for performing the parametric rendering. In the illustrated example, the output is a binaural signal, and thus the prototype signal is designed such that it has two channels (left and right) and is oriented in the spatial audio scene according to the head orientation of the user. The two-channel (i = 1,2) prototype signal can be formulated, for example, by:

[0188]

[0189] where p i,i is a mixing weight according to the head orientation information. For example, the prototype signal can be two cardioid pattern signals generated from the interpolated FOA signal, one pointing in the left direction (relative to the head orientation of the user), and one pointing in the right direction. When p 1,1 = p 2,1 = 0.5 and as follows (assuming WYZX channel order):

[0190] p 1,2 = 0.5 [cos(yaw)cos(roll) + sin(yaw)sin(pitch)sin(roll)]

[0191] p 1,3 = -0.5cos(pitch)sin(oll)

[0192] p 1,4 = 0.5 [cos(yaw)sin(pitch)sin(roll) - sin(yaw)cos(roll)]

[0193] and

[0194]

[0195] The above example of a cardioid prototype signal is merely one example. In other examples, the prototype signal can be different for different frequencies, e.g., the directivity of the spatial pattern can be less compared to the cardioid at lower frequencies, while the shape can be cardioid-like at higher frequencies. This choice is motivated because it is more similar to the binaural signal than a wideband cardioid pattern. However, it is not very critical which pattern design is applied, as long as the general trend is to obtain some left-right difference for the prototype signal. This is because the parametric processing steps described below will correct for inter-channel features anyway.

[0196] Further, the prototype signal can be represented in vector form as:

[0197]

[0198] Further, the prototype signal can be output to a covariance matrix estimator 803 and a mixer 809.

[0199] In some embodiments, the synthesis processor 509 is configured to estimate a covariance matrix of the time-frequency prototype signal and its total energy estimate in the frequency band. As previously mentioned, the covariance matrix can be estimated as:

[0200]

[0201] The estimation of the covariance matrix can involve a time average, such as an IIR average or an FIR average over several time indices n. The covariance matrix estimator 803 can further be configured to formulate a total energy estimate E(k, n), i.e., the sum of the diagonal values of C x (k, n). In some embodiments, instead of estimating the total energy from the prototype signal, the total energy estimate can be estimated based on the interpolated signal 512. For example, the total energy estimate has been determined in the signal interpolator shown in Figure 5 and can be obtained therefrom.

[0202] The total energy estimate 806 can be provided as an output to a target covariance matrix determiner 805. The estimated covariance matrix can be output to a mixing rule determiner 807.

[0203] The synthesis processor 509 can further comprise a target covariance matrix determiner 805. The target covariance matrix determiner 805 is configured to receive the interpolated spatial metadata 514 and the total energy estimate E(k, n) 806. In this example, the spatial metadata comprises an azimuth angle θ'(k, n), an elevation angle and a direct-to-total energy ratio r'(k, n). In some embodiments, the target covariance matrix determiner 805 further receives head orientation (yaw, pitch, roll) information 516.

[0204] In some embodiments, the target covariance matrix determiner is configured to rotate the spatial metadata according to the head orientation by:

[0205]

[0206] Further, the rotated directions are:

[0207] θ"(k, n) = atan2(v'2(k, n), v'1(k, n))

[0208]

[0209] The target covariance matrix determiner 805 can also use a pre-existing HRTF (Head Related Transfer Function) dataset at the synthesis processor. It is assumed that from this HRTF set, a 2x1 complex valued head related transfer function (HRTF) H(k, θ) can be obtained for any angle θ, and frequency band k For example, the HRTF data can be a dense HRTF set that has been pre-transformed to the frequency domain, so that the HRTF can be obtained at the mid-frequency of the frequency band k. In turn, at rendering time, the pair of HRTFs closest to the desired direction can be selected. In some embodiments, interpolation between the two or more closest data points can be performed. Various means for interpolating HRTFs have been described in the literature.

[0210] At the HRTF dataset, a diffuse field covariance matrix has also been formulated for each frequency band k. For example, by taking a set of uniformly distributed directions θ d , (where d = 1..D) and by estimating the diffuse field covariance matrix as:

[0211]

[0212] The diffuse field covariance matrix can be obtained.

[0213] In turn, the target covariance matrix determiner 805 can formulate the target covariance matrix C

[0214]

[0215] In turn, the target covariance matrix C y (k, n) is output to the mixing rule determiner 807.

[0216] In some embodiments, the synthesis processor 509 further comprises a mixing rule determiner 807. The mixing rule determiner 807 is configured to receive the target covariance matrix C y (k, n) and the measured covariance matrix C x (k, n), and generate a mixing matrix M(k, n). The mixing process can use Vilkamo, J., The method described in T. and Kuntz, A., "Optimized covariance domain framework for time-frequency processing of spatial audio", Journal of the Audio Engineering Society, Vol. 61, No. 6, pp. 403-411, 2013, can be used to generate the mixing matrix.

[0217] The formulas provided in the appendix of the above referenced document can be used to formulate the mixing matrix M(k, n). In the present disclosure, we use the same notation for the matrices for the sake of clarity. In some embodiments, the mixing rule determiner 807 is further configured to determine a prototype matrix Q(k, n) that guides the generation of the mixing matrix 812:

[0218]

[0219] The rationale of these matrices and the formulas used to obtain the mixing matrix M(k, n) based on these matrices are described in detail in the above referenced document and will not be repeated here. In short, the method provides a mixing matrix M(k, n) such that when applied to a signal with a covariance matrix C x (k, n), a signal with a covariance matrix C y (k, n) that is substantially the same or similar to C y (k, n) is produced in a least square optimized manner. In these embodiments, the prototype matrix Q is the identity matrix because the generation of the prototype signal has already been implemented by the prototype signal generator 801. Having an identity prototype matrix means that the processing aims at producing an output that is as similar as possible to the input (i.e. with respect to the prototype signal) while obtaining the target covariance matrix C

[0220] In some embodiments, the synthesis processor 509 comprises a mixer 809. The mixer 809 is configured to receive the time-frequency prototype audio signal 802 and the mixing matrix 812. The mixer 809 processes the input prototype audio signal 802 to generate two processed (binaural) time-frequency signals 814.

[0221]

[0222] where bin b is located in frequency band k.

[0223] The above process assumes that the input signals x(b, n) have suitable inter-channel incoherence between them to render the output signals y(b, n) with the desired target covariance matrix properties. In some cases, the input signals can not have suitable inter-channel incoherence. In these cases, a decorrelation operation needs to be used to generate decorrelated signals based on x(b, n) and mix these decorrelated signals into a specific residual signal that is added to the signal y(b, n) in the above equation. The process of obtaining such a residual signal has been explained in the previously cited references.

[0224] Further, the mixer 809 is configured to output the processed binaural time- frequency signals y(b, n) 814, which are provided to the inverse T / F transformer 811.

[0225] In some embodiments, the synthesis processor 509 comprises an inverse T / F transformer 811 that applies an inverse time-frequency transform corresponding to the applied time- frequency transform (such as an inverse STFT in case the signals are in the STFT domain) to the processed binaural time-frequency signals 814 to generate a spatialized audio output 518, which can be in binaural form reproducible through headphones.

[0226] Figure 8 The operation of the synthesis processor shown in Figure 9 is shown in the flowchart of

[0227] Hence, as shown in step 901 in Figure 9 the method comprises obtaining an interpolated (time-frequency) signal.

[0228] Further, as shown in step 902 in Figure 9 a listener head orientation is obtained.

[0229] Further, as shown in step 903 in Figure 9 a prototype signal is generated based on the interpolated (time-frequency) signal and the head orientation.

[0230] Additionally, as shown in step 905 in Figure 9 a covariance matrix is generated based on the prototype signal.

[0231] Further, as shown in step 906 in Figure 9 interpolated metadata can be obtained.

[0232] As shown in step 907 in Figure 9 a target covariance matrix is determined based on the interpolated metadata and the covariance matrix.

[0233] Further, as shown in step 909 in Figure 9 a mixing rule can be determined.

[0234] As Figure 9 Based on the mixing rules and the prototype signals, a mix can be generated to generate a spatialized audio signal, as indicated in step 911.

[0235] Further, as indicated in step 913, the spatialized audio signal can be output. Figure 9

[0236] Figure 10 Some further embodiments are shown in Fig. 10. In these embodiments, the system is as in Fig. 9, with the difference that the system is implemented in two separate devices, an encoder processor 1040 and a decoder processor 1060, and an additional encoder / MUX 1001 and DEMUX / decoder 1009. Figure 5

[0237] In these embodiments, the encoder processor 1040 is configured to receive the plurality of signal sets 500 and the microphone array positions 502 as input. The encoder processor 1040 further comprises a spatial analyzer 501 configured to receive the plurality of signal sets 500 and output metadata 506 for each array. The encoder processor 1040 further comprises an encoder / MUX 1001 configured to receive the plurality of signal sets 500, the metadata 506 for each array (from the spatial analyzer 501), and the microphone array positions 502. The encoder / MUX 1001 is configured to apply a suitable encoding scheme to the audio signals, e.g. any method for encoding Ambisonic signals that has been described in the context of MPEG-H. The encoder / MUX 1001 block can also downmix or otherwise reduce the number of audio channels to be encoded. Furthermore, the encoder / MUX 1001 can quantize and encode the spatial metadata and the array position information and embed the encoded result in a bitstream 1006 together with the encoded audio signals. The bitstream 1006 can further be provided with the encoded video signals at the same media container. Further, the encoder / MUX 1001 outputs the bitstream 1006. Depending on the bit rate employed, the encoder can have omitted the encoding of some signal sets, in which case it can have omitted the encoding of the corresponding array positions and metadata (however, they can also be retained in order to use them for metadata interpolation).

[0238] ​​The decoder processor 1060 includes a DEMUX / decoder 1009. The DEMUX / decoder 1009 is configured to receive the bitstream 1006, and to decode and de-multiplex the multiple signal sets based on the microphone array 500', (and provide them to the signal interpolator 503), the microphone array positions 502' (and provide them to the position pre-processor 505), and the metadata 506' for each array (and provide them to the metadata interpolator 507).

[0239] The decoder processor 1060 also includes the signal interpolator 503, the position pre-processor 505, the metadata interpolator 507, and the synthesis processor 509, as discussed in relation to Figure 5 and Figure 8 are discussed in further detail.

[0240] In the above example, information relating to the array positions is transmitted from the encoder processor 1040 to the decoder processor 1060 via the bitstream 1006, but in some embodiments this can not be necessary as the system can be configured such that the position pre-processor 505 is implemented within the encoder processor 1040. In such an example, the encoder processor is configured to generate the necessary interpolation data at a suitable grid of predefined expected user positions (e.g. at a spatial resolution of 10cm). This interpolation data can be encoded using a suitable means, and provided to the decoder (for decoding) in the form of a bitstream. In turn, this interpolation data would be used at the decoder processor 1060 as a look-up table based on the user position by selecting the closest existing data set to the user position.

[0241] In relation to Figure 11 , a flowchart showing the operation of the system as shown in Figure 10 is shown.

[0242] As shown in step 1101 of Figure 11 , the method can begin with obtaining multiple signal sets based on microphone array signals.

[0243] In turn, as shown in step 1103 of Figure 11 , the method can include spatially analysing the signal sets to generate spatial metadata.

[0244] In turn, as shown in step 1105 of Figure 11 , the metadata, signals and other information can be encoded and multiplexed.

[0245] In turn, as shown in step 1107 of Figure 11 , the encoded and multiplexed signals and information can be decoded and de-multiplexed.

[0246] As shown in Figure 11As shown in step 1109, once the microphone array position and listener orientation / position have been obtained, the method can obtain the interpolation factors by processing the relative positions.

[0247] Once the interpolation factors have been obtained by processing the relative positions and signals / metadata, the method can interpolate the signals as shown in step 1111 and interpolate the metadata as shown in step 1113. Figure 11 Figure 11 Once the interpolated metadata and signals have been determined, as well as the listener orientation / position, the method can apply the synthesis processing as shown in step 1115.

[0248] As shown in step 1117, the spatialized audio is output. Figure 11

[0249] As shown in step 1117, the spatialized audio is output. Figure 12

[0250] With respect to Figure 10 , example applications of the encoder and decoder processors of Figure 12 are shown.

[0251] In this example, there are three microphone arrays, which can be, for example, spherical arrays with a sufficient number of microphones (e.g., 30 or more), or VR cameras (e.g., OZO, etc.) with microphones mounted on their surface. Thus, microphone array 1 1201, microphone array 2 1211, and microphone array 3 1221 are shown, which are configured to output audio signals to computer 1 1205 (and, in this example, to FOA / HOA converter 1215).

[0252] In addition, each array is also equipped with a localizer that provides position information for the corresponding array. Thus, microphone array 1 localizer 1203, microphone array 2 localizer 1213, and microphone array 3 localizer 1223 are shown, which are configured to output position information to computer 1 1205 (and, in this example, to encoder processor 1040).

[0253] Figure 10 The system infurther includes a computer, computer 1 1205, which includes FOA / HOA converter 1215, which is configured to convert the array signals to first order Ambisonic (FOA) or higher order Ambisonic (HOA) signals. Converting microphone array signals to Ambisonic signals is known and not described in detail herein, but if the array is, for example, Eigenmikes, there are available means for converting the microphone signals to Ambisonic form.

[0254] ​​The FOA / HOA converter 1215 outputs a converted Ambisonic signal in the form of multiple signal sets based on the microphone array signal 1216 to the encoder processor 1040, which can operate as the encoder processor 1040 as described above.

[0255] Microphone array positioners 1203, 1213, and 1223 are configured to provide microphone array position information to an encoder processor in computer 11205 via a suitable interface (e.g., via Bluetooth connection). In some embodiments, the array positioner also provides rotation alignment information, which can be provided to rotate and align the FOA / HOA signals at computer 11205.

[0256] The encoder processor 1040 at computer 1 1205 is configured as follows: Figure 8 In the context described above, it processes multiple signal sets and microphone array positions based on microphone array signals and provides an encoded bitstream of 1006 as output.

[0257] The bitstream 1006 can be stored and / or transmitted, and the decoder processor 1060 of computer 2 1207 is configured to receive or obtain the bitstream 1006 from the storage device. The decoder processor 1060 can also obtain listener position and orientation information from the position / orientation tracker of the HMD (Head-Mounted Display) 1231 worn by the user. Based on the bitstream 1006 and the listener position and orientation information 1230, the decoder processor of computer 2 1207 is configured to generate binaural spatialized audio output signals 1232 and provide them via a suitable audio interface for reproduction through headphones 1233 worn by the user.

[0258] In some embodiments, computer 2 1207 is the same device as computer 1 1205; however, in typical cases, they are different devices or computers. In this context, "computer" can refer to a desktop / laptop computer, cloud processing, game controller, mobile device, or any other device capable of performing the processes described in this disclosure.

[0259] In some embodiments, bitstream 1006 is an MPEG-I bitstream. In some other embodiments, it can be any suitable bitstream.

[0260] In the above embodiments, the spatial parameterization analysis of directional audio coding can be replaced by an adaptive beamforming approach. The adaptive beamforming approach can for example be the COMPASS method outlined in Archontis Politis, Sakari Tervo, and Ville Pulkki, “COMPASS: Coding and Multidirectional Parameterization of Ambisonic Sound Scenes,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018.

[0261] In such embodiments, the spatial covariance matrix C HOA,j (k,n) can be computed from the Ambisonic signals as previously defined, but including higher order Ambisonic (HOA) channels if available. For example, the signals can be represented as:

[0262]

[0263] where N is the Ambisonic order. In some embodiments, the spatial covariance matrix can be decomposed by eigenvalue decomposition:

[0264] C HOA,j (k,n) = E(k,n)V(k,n)E H (k,n)

[0265] where E(k,n) contains the eigenvectors and V(k,n) contains the eigenvalues. Further, the diffuse or non-diffuse condition determination can be performed based on a statistical analysis of the ordered eigenvalues contained in the diagonal of V(k,n).

[0266] If a non-diffuse condition is detected, the number of dominant sources S' is also estimated based on a statistical analysis of the distribution of the ordered eigenvalues. For a robust estimation, the number of sources is bounded by:

[0267] S = min(S', (N + 1) 2 / 2)

[0268] After estimating the number of sources, their approximate directions of arrival (DOA) are determined. For a dense pre-computed grid of m = 1,..., M directions (θ m , φ m ) uniformly arranged on a sphere over a range of M = 1000 ~ 5000 angles, the spatial power spectrum can be computed as:

[0269]

[0270] where y N is a spherical harmonic vector up to order N and has the appropriate ordering and normalization for the applied Ambisonic convention. Further, the estimated DOA corresponds to the grid direction with S highest peaks.

[0271] In some other embodiments, the DOA estimation can employ higher resolution subspace methods (especially at low Ambisonic order) to overcome the limitation of wide low order beams to distinguish close angles sources. For example, MUSIC can be used, where the spatial spectrum is computed as:

[0272]

[0273] where E noise (k,n) is the last (N+1) 2 S ordered eigenvectors are formed. After performing MUSIC for all grid points, similarly, the DOA is found by looking for the S highest peaks.

[0274] After the DOAs (θ s ,φ s ) for s = 1,..,S have been determined, the direct-to-total (DTR) energy ratio per source can be determined as:

[0275]

[0276] Further, the source with the highest DTR can be selected as the dominant source and the corresponding parameters r j,s (k,n), θ s (k,n), φ s (k,n) are passed to the metadata interpolator, similar to the DirAC analysis described above.

[0277] In some further embodiments, instead of selecting a single dominant DOA and DTR, some or all of the detected DOAs and DTRs are passed to the metadata interpolator. In other words, in some embodiments, there are multiple simultaneous directions and ratios for each time-frequency tile.

[0278] Thus, while the foregoing embodiments discuss estimating one simultaneous direction estimate for each time-frequency interval in some embodiments, multiple directions for each time-frequency tile can be estimated or otherwise determined.

[0279] For example, the metadata interpolation principles described herein can also be extended for two or more simultaneous direction estimates (at each time-frequency interval) and corresponding two or more direct-to-total energy ratios. In this case, the interpolated metadata also contains two or more direction estimates.

[0280] In some embodiments, the implemented method can be, for example:

[0281] 1) Form the direction vectors from all involved direction parameters (and corresponding ratios) using the method described in the foregoing.

[0282] 2) Determine the closest array to the listener.

[0283] 3) Select the longest direction vector (i.e. with the largest direct-to-total ratio) from the closest array.

[0284] 4) For the remaining arrays involved in the interpolation, select those with the direction vector (one per array) that has the largest dot product with the selected vector of the closest array.

[0285] 5) Form the combined vector based on the selected vectors (of steps 3 and 4) and the interpolation weights (as described in the foregoing) and obtain the direction and ratios based on it (as described in the foregoing).

[0286] 6) Discard the vector data that was selected for use in the above steps 3 and 4.

[0287] 7) If there are still direction vectors in the closest array, repeat steps 3-6 to determine the next direction and its corresponding ratio until a sufficient number of interpolated directions and ratios are obtained.

[0288] In some embodiments, a minimum distance assignment algorithm such as the Hungarian algorithm is used to pair the closest DOAs between the sets. Since the number of DOAs between the microphones can differ, the assignment can occur between an equal number of DOAs for pairs of microphones, while additional DOAs that are not assigned in a certain microphone can still be interpolated at other microphones with zero DOA vectors. With this approach, as many DOAs as the maximum number of DOAs detected across the three microphone arrays can be passed to the synthesis stage.

[0289] In some embodiments, when there are multiple simultaneous directions of arrival, at the target covariance matrix determiner 805 of the synthesis processor 509 shown in Figure 5 the target covariance matrix is constructed with more than one direct part (for each direction and its corresponding direct-to-total energy ratio). Otherwise, the synthesis process can be the same.

[0290] In some embodiments, the signal interpolator 503 is configured to interpolate the audio signals using any suitable method, as indicated in Figure 10 For example, instead of switching the signals, the signals are linearly interpolated based on the weight factors (w1, w2 and w3). In some cases, this interpolation method can lead to undesired comb filtering, but in some cases it can provide better quality.

[0291] In some embodiments, the interpolated data 508 / 510, the microphone array positions 502, and / or the listener position 504 are also forwarded to the synthesis processor 509. These can be used, for example, to determine prototype signals (e.g. using a wider pattern when the listener is far away from any array so as not to lose any signal energy).

[0292] In some embodiments, the functions or processing blocks described in the foregoing embodiments can be combined and / or divided into other functional blocks or further processing blocks in various ways. For example, in some embodiments, the functions (or processing steps) associated with the signal interpolator 503, the position pre-processor 505 and the metadata interpolator 507 are integrated within the synthesis processor 509. In some embodiments, combining the functions (or processing steps) results in more compact code and an efficient implementation.

[0293] In some embodiments, the prototype signals can already be determined in the signal interpolator 503. In such embodiments, the listener orientation 516 is provided to the signal interpolator 503.

[0294] In some embodiments, the target total energy is determined in the signal interpolator 503 and passed to the synthesis processor 509. In these embodiments, the interpolated signals 512S(b,n,i) can not need to be energy corrected in the signal interpolator 503, as the energy correction can be performed in the synthesis processor 509 (using the received target energy instead of the target energy determined based on the received audio signals). This can be beneficial in some practical systems, as the energy correction can be performed simultaneously with the spatial synthesis, thereby potentially reducing the computational complexity. Furthermore, these embodiments can feature improved audio quality, as all gains can be applied simultaneously (and hence potentially only one time gain smoothing can be applied).

[0295] In some embodiments, the interpolation weights (w1, w2 and w3) can be determined using any suitable scheme. For example, in some embodiments, the foregoing embodiments can be adapted so as to use the closest array more predominantly.

[0296] In the embodiments described herein, the signal interpolator 503 is configured to determine the selected microphone array j selso that it is always one of the microphone arrays j1, j2, j3 in which the listener position is located. In some cases, if the listener is on the edge of two determined triangles, this determination can result in a switch between two microphone arrays. To prevent such fast switching, in some embodiments a threshold can be applied in the selection of the microphone array. For example, only when some of the microphone arrays j1, j2, j3 are closer to the listener position than j sel The selected microphone array j is changed only when the microphone array j sel .

[0297] In some embodiments, a combination of different methods can be used to perform the parameter interpolation. For example, two different methods for interpolating the direct-to-total energy ratio are proposed above. In some embodiments, a combination of these methods can be implemented. For example, if the first method (in other words, the length of the combination vector) provides a value below a threshold, the result of the first method is selected, otherwise the result of the second method (in other words, the direct-to-total ratio is weighted directly) is selected. The threshold can be fixed or adaptive. For example, in some embodiments, the threshold can be determined relative to the original ratio.

[0298] In some embodiments discussed above, an encoder and a decoder are provided as shown in Figure 7 In some other embodiments, the spatial analysis is performed in the decoder (at least at some frequencies). In these embodiments, only the audio signal and the microphone positions need to be passed from the encoder to the decoder. In some embodiments, also the spatial metadata at certain frequencies is transmitted.

[0299] As shown in Figure 13 When the listener is located outside the area related to the microphone array position, the listener position can be projected into that area, as shown in

[0300] In some embodiments, the signal interpolator 503 can compute the sound scene energy at each microphone from all Ambisonic channels (including higher order Ambisonic channels) rather than using only the energy of the first channel as normalized for SN3D Ambisonic channels or normalized for N3D Ambisonic channels where N is the Ambisonic order.

[0301] The above embodiments assume that the microphone arrays are positioned with the same orientation, or alternatively are converted to the same orientation (in other words, the "x-axis" of each microphone array is aligned and pointing in the same direction). In some embodiments, in addition to the position information, microphone array orientation information is also communicated. In turn, this information can be used at any point in the processing in order to account for different orientations and "align" these microphone orientations.

[0302] With respect to ​ , an example electronic device that can be used as a computer, an encoder processor, a decoder processor, or any of the functional blocks described herein is shown. The device can be any suitable electronic device or apparatus. For example, in some embodiments, the device 1400 is a mobile device, a user equipment, a tablet computer, a computer, an audio playback apparatus, etc.

[0303] In some embodiments, the device 1400 includes at least one processor or central processing unit 1407. The processor 1407 can be configured to execute various program codes, such as the methods described herein.

[0304] In some embodiments, the device 1400 includes a memory 1411. In some embodiments, the at least one processor 1407 is coupled to the memory 1411. The memory 1411 can be any suitable storage component. In some embodiments, the memory 1411 includes a program code portion for storing program codes that can be implemented on the processor 1407. Further, in some embodiments, the memory 1411 can also include a storage data portion for storing data (e.g., data that has been processed or is to be processed according to the embodiments described herein). The implemented program codes stored in the program code portion and the data stored in the storage data portion can be retrieved by the processor 1407 via the memory-processor coupling as needed.

[0305] In some embodiments, the device 1400 includes a user interface 1405. In some embodiments, the user interface 1405 can be coupled to the processor 1407. In some embodiments, the processor 1707 can control the operation of the user interface 1405 and receive input from the user interface 1405. In some embodiments, the user interface 1705 can enable a user to enter commands to the device 1400, for example, via a keypad. In some embodiments, the user interface 1405 can enable the user to obtain information from the device 1400. For example, the user interface 1405 can include a display for displaying information from the device 1400 to the user. In some embodiments, the user interface 1405 can include a touch screen or touch interface that is capable of both entering information into the device 1400 and displaying information to the user of the device 1400.

[0306] In some embodiments, the device 1400 includes an input / output port 1409. In some embodiments, the input / output port 1409 includes a transceiver. In such embodiments, the transceiver can be coupled to the processor 1407 and configured to enable communication with other apparatuses or electronic devices, for example, via a wireless communication network. In some embodiments, the transceiver or any suitable transceiver or transmitter and / or receiver components can be configured to communicate with other electronic devices or apparatuses via a wired or wired coupling.

[0307] The transceiver can communicate with other apparatuses by any suitable known communication protocol. For example, in some embodiments, the transceiver can use suitable Universal Mobile Telecommunications System (UMTS) protocols, wireless local area network (WLAN) protocols such as IEEE 802.X, suitable short-range radio frequency communication protocols such as Bluetooth, or infrared data communication paths (IRDA).

[0308] The transceiver input / output port 1409 can be configured to send / receive audio signals, bit streams, and in some embodiments, perform operations and methods as described above by executing suitable code using the processor 1407.

[0309] In general, the various embodiments of the application can be implemented in hardware or special-purpose circuits, software, logic or any combination thereof. For example, some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software which can be executed by a controller, microprocessor or other computing device, although the application is not limited thereto. While various aspects of the application can be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein can be implemented in, as non-limiting examples, hardware, software, firmware, special-purpose circuits or logic, general purpose hardware or controler or other computing devices, or some combination thereof.

[0310] Embodiments of the application can be implemented by computer software executable by a data processor of the mobile device such as in the processor entity, or by hardware, or by a combination of software and hardware. Further, in this regard, it should be noted that any blocks of the logic flow of the accompanying figures can represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software can be stored on such physical media as memory chips, or memory blocks implemented in the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD.

[0311] The memory can be of any type appropriate for the local technical environment and can be implemented using any appropriate data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processor can be of any type appropriate for the local technical environment, and can include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASIC), gate level circuits and processors based on multi core processor architectures, as non-limiting examples.

[0312] Embodiments of the application can be practiced in a variety of components such as integrated circuit modules. The design of integrated circuits is by nature a highly automated process. Complex and powerful software programs are available for conversion of a logic level design into a semiconductor circuit design ready to be etched and formed on semiconductor chips.

[0313] Programs, such as those provided by Synopsys, Inc. of Mountain View, California and Cadence Design, of San Jose, California automatically route conductors and locate components on a semiconductor chip using well-established rules of design as well as libraries of pre-stored design modules. Once the design for a semiconductor circuit has been completed, the resultant design, in a standardized electronic format (e.g., Opus, GDSII, or the like) can be transmitted to a semiconductor fabrication facility or "fab" for fabrication.

[0314] The foregoing description has provided by way of exemplary and non-limiting examples a full and informative description of the exemplary embodiments of the application. However, various modifications and adaptations to the foregoing illustrative embodiments of this application can become apparent to those skilled in the relevant arts in view of the foregoing description, when read, for example, in the light of the appended claims. However, all such and similar modifications of the teachings of this application will still fall within the scope of the application as defined in the appended claims.

Claims

1. An apparatus comprising: At least one processor; as well as At least one memory storing instructions, wherein when executed with the at least one memory, the means causes the device to at least: Obtain two or more spatial audio streams, wherein each of the two or more spatial audio streams includes at least one audio signal and at least one spatial metadata parameter, wherein each of the two or more spatial audio streams is associated with a location; Obtain the position associated with at least two of the two or more spatial audio streams; Obtain the listener's location, wherein the listener's location is configured to be tracked; At least one audio signal is generated based at least in part on at least one of the two or more spatial audio streams, the location, and the listener's location; At least one modified spatial metadata parameter is generated, based at least in part on at least one spatial metadata parameter included by the at least two spatial audio streams, the location, and the listener's location; and The at least one audio signal is processed to generate spatial audio output, based at least in part on the at least one modified spatial metadata parameter.

2. The apparatus according to claim 1, wherein, The at least one memory stores instructions, which, when executed using the at least one memory, cause the device to at least: An updated listener location is obtained, at least in part based on tracking the listener's location, wherein the at least one modified spatial metadata parameter is generated at least in part based on the updated listener location.

3. The apparatus according to claim 1, wherein, The at least one memory stores instructions, which, when executed using the at least one memory, cause the device to at least: The two or more spatial audio streams are obtained from a microphone device, wherein the microphone device is located at a corresponding position and each includes one or more microphones.

4. The apparatus according to claim 1, wherein, The two or more spatial audio streams are each associated with a direction, wherein the at least one memory stores instructions that, when executed with the at least one memory, cause the device to: Obtain the orientation of the two or more spatial audio streams, wherein at least one generated audio signal is further based on the orientation associated with the two or more spatial audio streams, and wherein the at least one modified spatial metadata parameter is further based on the orientation associated with at least two of the two or more spatial audio streams.

5. The apparatus according to claim 1, wherein, The at least one memory stores instructions, which, when executed using the at least one memory, cause the device to: Obtain listener orientation, wherein the at least one modified spatial metadata parameter is further based on the listener orientation.

6. The apparatus according to claim 5, wherein, Processing at least one audio signal to generate the spatial audio output based on the at least one modified spatial metadata parameter includes the at least one memory storage instruction, which, when executed with the at least one memory, causes the device to: The at least one audio signal is further processed based on the listener orientation.

7. The apparatus according to claim 1, wherein, The at least one memory stores instructions, which, when executed using the at least one memory, cause the device to: Control parameters are obtained based on the positions associated with at least two of the two or more spatial audio streams and the listener position, wherein at least one of the at least one audio signal or at least one of the at least one modified spatial metadata parameters is controlled at least in part based on the control parameters.

8. The apparatus according to claim 7, wherein, The at least one memory stores instructions, which, when executed with the at least one memory, cause the device to perform at least one of the following: Identify at least three spatial audio streams in which the listener's position is located, and generate weights associated with the at least three spatial audio streams based on the associated positions and the listener's position; or Identify the two spatial audio streams that are closest to the listener's position from among the two or more spatial audio streams, and generate weights associated with the two spatial audio streams based on the associated positions and the vertical projection of the listener's position onto the line between the two spatial audio streams.

9. The apparatus according to claim 8, wherein, The at least one memory stores instructions, which, when executed with the at least one memory, cause the device to perform one of the following: Two or more audio signals from the two or more spatial audio streams are combined based on at least one of the following: weights associated with the at least three spatial audio streams, or weights associated with the two spatial audio streams; Based on which of the two or more spatial audio streams is closest to the listener's location, one or more audio signals are selected from one of the two or more spatial audio streams; or Based on which of the two or more spatial audio streams is closest to the listener's location and a further switching threshold, one or more audio signals are selected from the one of the two or more spatial audio streams.

10. The apparatus according to claim 8, wherein, The at least one memory stores instructions, which, when executed using the at least one memory, cause the device to: Based on at least one of the following, combine at least one spatial metadata parameter included in at least two of the two or more spatial audio streams: weights associated with the at least three spatial audio streams, or weights associated with the two spatial audio streams.

11. The apparatus according to claim 1, wherein, Processing the at least one audio signal to generate the spatial audio output includes the at least one memory storage instruction, which, when executed with the at least one memory, causes the device to: Generate at least one of the following: Binaural audio output, comprising two audio signals for over-ear headphones and / or in-ear headphones; or Multi-channel audio output, which includes at least two audio signals for a multi-channel speaker group.

12. The apparatus according to claim 1, wherein, At least one spatial metadata parameter includes at least one of the following: At least one direction; At least one direct-to-total ratio associated with at least one directional value; At least one extended coherence associated with at least one direction value; At least one distance associated with at least one direction value; At least one circumferential coherence; At least one diffusion pair ratio; or At least one remaining pair of total ratios.

13. A method comprising: Obtain two or more spatial audio streams, wherein each of the two or more spatial audio streams includes at least one audio signal and at least one spatial metadata parameter, wherein each of the two or more spatial audio streams is associated with a location; Obtain the position associated with at least two of the two or more spatial audio streams; Obtain the listener's location, wherein the listener's location is configured to be tracked; At least one audio signal is generated based at least in part on at least one of the two or more spatial audio streams, the location, and the listener's location; At least one modified spatial metadata parameter is generated, based at least in part on at least one spatial metadata parameter included by the at least two spatial audio streams, the location, and the listener's location; and The at least one audio signal is processed to generate spatial audio output, based at least in part on the at least one modified spatial metadata parameter.

14. The method of claim 13, further comprising: The updated listener location is obtained at least in part based on the tracking of the listener's location, wherein the at least one modified spatial metadata parameter is generated at least in part based on the updated listener location.

15. The method of claim 13, further comprising: The two or more spatial audio streams are obtained from a microphone device, wherein the microphone device is located at a corresponding position and each includes one or more microphones.

16. The method according to claim 13, wherein, The two or more spatial audio streams are each associated with a orientation, wherein the method further includes: Obtain the orientation of the two or more spatial audio streams, wherein at least one generated audio signal is further based on the orientation associated with the two or more spatial audio streams, and wherein the at least one modified spatial metadata parameter is further based on the orientation associated with at least two of the two or more spatial audio streams.

17. The method of claim 13, further comprising: Obtain listener orientation, wherein the at least one modified spatial metadata parameter is further based on the listener orientation.

18. The method according to claim 17, wherein, Processing at least one audio signal to generate the spatial audio output based on the at least one modified spatial metadata parameter includes: The at least one audio signal is further processed based on the listener orientation.

19. The method according to claim 13, wherein, At least one spatial metadata parameter includes at least one of the following: At least one direction; At least one direct-to-total ratio associated with at least one directional value; At least one extended coherence associated with at least one direction value; At least one distance associated with at least one direction value; At least one circumferential coherence; At least one diffusion pair ratio; or At least one remaining pair of total ratios.

20. An apparatus comprising: At least one processor; as well as At least one memory storing instructions, wherein when executed with the at least one memory, the means causes the device to at least: Acquire at least one signal, wherein the at least one signal comprises one or more audio signals from corresponding capture locations of two or more capture locations, wherein the two or more capture locations are within an audio scene; Obtain at least one capture location, wherein the at least one capture location includes one or more of the two or more capture locations; Obtain at least one spatial metadata parameter, wherein the at least one spatial metadata parameter includes one or more spatial metadata parameters associated with a corresponding capture location among the two or more capture locations; Obtain multiple predefined listener locations; At least in part, based on the at least one spatial metadata parameter, the at least one capture location, and the plurality of predefined listener locations, a plurality of modified spatial metadata parameters are generated; and Two or more spatial audio streams are generated, at least in part based on the at least one signal, the at least one capture location, and the plurality of modified spatial metadata parameters.

Citation Information

Patent Citations

  • Analysis of spatial metadata from multi-microphones having asymmetric geometry in devices

    WO2018091776A1

  • Determination of targeted spatial audio parameters and associated spatial audio playback

    WO2019086757A1