Binaural audio rendering of spatial audio

By analyzing the inter-channel characteristics and head orientation of the transmitted audio signal, an adapted binaural audio signal is generated, which solves the problems of sound quality degradation and artifacts in traditional methods, and achieves high-quality audio rendering under any head orientation.

CN120303957APending Publication Date: 2025-07-11NOKIA TECHNOLOGIES OY
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202380082697.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-01
Filing Date
2023-11-06
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In prior art In head tracking binaural audio rendering, traditional methods cannot effectively adapt to changes in user head orientation, resulting in a decrease in sound quality and artifacts.

Method used

By analyzing the inter-channel characteristics of the transmitted audio signal and the user's head orientation, the mixed information is determined, and the adapted binaural audio signal is generated to avoid negative artifacts and achieve head tracking binaural rendering.

Benefits of technology

Maintain sound quality under any head orientation, avoiding artifacts, and delivering a high-quality binaural audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303957A_ABST
    Figure CN120303957A_ABST
Patent Text Reader

Abstract

A method for generating a spatial output audio signal, the method comprising: obtaining a spatial audio signal, the spatial audio signal comprising: at least two channel audio signals; and at least one spatial parameter associated with the at least two channel audio signals; analyzing the at least two channel audio signals to determine at least one inter-channel characteristic; obtaining orientation and / or position parameters; determining mixing information based on the at least one inter-channel characteristic and the orientation and / or location parameter; and generating at least two channel output audio signals based on the at least two channel audio signals, the orientation and / or location parameters, and the mixing information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to apparatuses and methods for binaural audio rendering for spatial audio, but is not limited to apparatuses and methods for leveraging adaptive prototype generation for head-tracked binaural rendering within parametric spatial audio rendering. Background Art

[0002] There are various methods for capturing spatial audio. One option is to use a microphone array (e.g., as part of a mobile device) to capture spatial audio. Using the microphone signals, spatial analysis of the sound scene can be performed to determine spatial metadata in frequency bands. Additionally, the microphone signals can be used to determine the transmitted audio signal. The spatial metadata and the transmitted audio signal can be combined to form a spatial audio stream.

[0003] Metadata-Assisted Spatial Audio (MASA) is an example of a spatial audio stream. It is one of the input formats that the upcoming Immersive Voice and Audio Service (IVAS) codec will support. It uses an audio signal together with corresponding spatial metadata (e.g., containing direction and direct-to-total energy ratio in frequency bands) and descriptive metadata (containing additional information related to, for example, the original capture and (transmitted) audio signal). A MASA stream can be obtained, for example, by capturing spatial audio using a microphone of a mobile device, where the set of spatial metadata is estimated based on the microphone signals. A MASA stream can also be obtained from other sources (such as specific spatial audio microphones (such as Ambisonics), studio mixes (e.g., 5.1 mix) or other content) by means of a suitable format conversion. MASA tools within the codec can also be used to encode a multi-channel signal by converting it into a MASA stream and encoding the stream. Summary of the Invention

[0004] According to a first aspect, there is provided a method for generating a spatial output audio signal, the method comprising: obtaining a spatial audio signal, the spatial audio signal comprising: at least two-channel audio signals; and at least one spatial parameter associated with the at least two-channel audio signals; analyzing the at least two-channel audio signals to determine at least one inter-channel characteristic; obtaining a direction and / or position parameter; determining mixing information based on the at least one inter-channel characteristic and the direction and / or position parameter; and generating at least two-channel output audio signals based on the at least two-channel audio signals, the direction and / or position parameter, and the mixing information.

[0005] Generating the at least two-channel output audio signals may further comprise: generating the at least two-channel output audio signals based on at least one spatial parameter associated with the at least two-channel audio signals.

[0006] Determining the mixed information may further include: determining the mixed information further based on at least one spatial parameter.

[0007] Analyzing at least two channel audio signals to determine at least one inter-channel characteristic may include: generating an inter-channel characteristic based on at least one spatial parameter associated with the at least two channel audio signals.

[0008] At least one spatial parameter associated with the at least two channel audio signals may include: a spatial parameter associated with a respective channel audio signal among the at least two channel audio signals; and a spatial parameter associated with the at least two channel audio signals.

[0009] Generating at least two channel output audio signals based on the at least two channel audio signals, at least one spatial parameter associated with the at least two channel audio signals, a directivity parameter, and the mixed information may include: generating at least one prototype matrix based on the mixed information; rendering at least two channel output audio signals from the at least two channel audio signals based on: at least one spatial parameter associated with the at least two channel audio signals, the directivity parameter, the directivity parameter, and the at least one prototype matrix.

[0010] Generating at least two channel output audio signals based on the at least two channel audio signals, at least one spatial parameter associated with the at least two channel audio signals, a directivity parameter, and the mixed information may include: processing the at least two channel audio signals based on the mixed information to generate at least two channel-adapted audio signals; rendering at least two channel output audio signals from the at least two channel-adapted audio signals based on: at least one spatial parameter associated with the at least two channel audio signals; and the directivity parameter.

[0011] Processing the at least two channel audio signals based on the mixed information to generate at least two channel-adapted audio signals may include: adapting the at least two channel audio signals based on a current directivity and an inter-channel characteristic.

[0012] Adapting the at least two channel audio signals based on a current directivity and an inter-channel characteristic may include: determining a mono factor based on the current directivity and the inter-channel characteristic, the mono factor being configured to indicate how the at least two channel audio signals should be mixed to avoid negative artifacts within the at least two channel output audio signals.

[0013] Analyzing at least two-channel audio signals to determine at least one inter-channel characteristic may include analyzing at least two-channel audio signals to determine at least one of the following: an inter-channel level difference between the at least two-channel audio signals; an inter-channel level difference between the at least two-channel audio signals of the modified at least two-channel audio signals, the modification being based on orientation and / or position parameters; an inter-channel phase difference between the at least two-channel audio signals; an inter-channel time difference between the at least two-channel audio signals; an inter-channel similarity measure between the at least two-channel audio signals; and an inter-channel correlation between the at least two-channel audio signals.

[0014] Processing at least two-channel audio signals to generate at least two-channel adapted audio signals based on mixing information may include: mixing the at least two-channel audio signals based on an inter-channel difference such that an audio component in substantially one of the at least two-channel audio signals is mixed into a corresponding one of the at least two-channel adapted audio signals and further at least partially cross-mixed into another one of the at least two-channel adapted audio signals.

[0015] Processing at least two-channel audio signals to generate at least two-channel adapted audio signals based on mixing information may further include: switching at least two of the at least two-channel adapted audio signals generated based on orientation and / or position parameters indicating an orientation towards a rear direction.

[0016] The at least two-channel output audio signals may be binaural audio signals.

[0017] The method may further include: obtaining a user's head orientation and / or position, and wherein obtaining the orientation and / or position parameters includes: processing the user's head orientation and / or position to generate the orientation and / or position parameters.

[0018] According to a second aspect, there is provided an apparatus for generating a spatial output audio signal, the apparatus including components configured to perform the following operations: obtaining a spatial audio signal, the spatial audio signal including: at least two-channel audio signals; and at least one spatial parameter associated with the at least two-channel audio signals; analyzing the at least two-channel audio signals to determine at least one inter-channel characteristic; obtaining orientation and / or position parameters; determining mixing information based on the at least one inter-channel characteristic and the orientation and / or position parameters; and generating at least two-channel output audio signals based on the at least two-channel audio signals, the orientation and / or position parameters, and the mixing information.

[0019] A component configured to generate at least two channel output audio signals may be further configured to generate at least two channel output audio signals based on at least one spatial parameter associated with the at least two channel audio signals.

[0020] A component configured to determine mixing information may be further configured to determine the mixing information further based on at least one spatial parameter.

[0021] A component configured to analyze at least two channel audio signals to determine at least one inter-channel characteristic may be configured to generate the inter-channel characteristic based on at least one spatial parameter associated with the at least two channel audio signals.

[0022] At least one spatial parameter associated with the at least two channel audio signals may include: a spatial parameter associated with a respective audio channel audio signal among the at least two audio channel audio signals; and a spatial parameter associated with the at least two audio channel audio signals.

[0023] A component configured to generate at least two channel output audio signals based on at least two channel audio signals, at least one spatial parameter associated with the at least two channel audio signals, a directivity parameter, and mixing information may be configured to generate at least one prototype matrix based on the mixing information; and render at least two channel output audio signals from the at least two channel audio signals based on: at least one spatial parameter associated with the at least two channel audio signals, the directivity parameter, the directivity parameter, and at least one prototype matrix.

[0024] A component configured to generate at least two channel output audio signals based on at least two channel audio signals, at least one spatial parameter associated with the at least two channel audio signals, a directivity parameter, and mixing information may be configured to process the at least two channel audio signals based on the mixing information to generate at least two channel-adapted audio signals; and render at least two channel output audio signals from the at least two channel-adapted audio signals based on: at least one spatial parameter associated with the at least two channel audio signals; and the directivity parameter.

[0025] A component configured to process at least two channel audio signals based on mixing information to generate at least two channel-adapted audio signals may be configured to adapt the at least two channel audio signals based on a current directivity and an inter-channel characteristic.

[0026] A component configured to adapt at least two channel audio signals based on a current directivity and an inter-channel characteristic may be configured to determine a mono factor based on the current directivity and the inter-channel characteristic, the mono factor being configured to indicate how the at least two channel audio signals should be mixed to avoid negative artifacts within the at least two channel output audio signals.

[0027] A component configured to analyze at least two-channel audio signals to determine at least one inter-channel characteristic may be configured to analyze at least two-channel audio signals to determine at least one of the following: an inter-channel level difference between the at least two-channel audio signals; a modified inter-channel level difference between the at least two-channel audio signals, the modification being based on orientation and / or position parameters; an inter-channel phase difference between the at least two-channel audio signals; an inter-channel time difference between the at least two-channel audio signals; an inter-channel similarity metric between the at least two-channel audio signals; and an inter-channel correlation between the at least two-channel audio signals.

[0028] A component configured to process at least two-channel audio signals based on mixing information to generate at least two-channel adapted audio signals may be configured to: mix the at least two-channel audio signals based on an inter-channel difference such that an audio component in substantially one of the at least two-channel audio signals is mixed into a corresponding one of the at least two-channel adapted audio signals and is further at least partially cross-mixed into another one of the at least two-channel adapted audio signals.

[0029] A component configured to process at least two-channel audio signals based on mixing information to generate at least two-channel adapted audio signals may further be configured to: switch at least two of the at least two-channel adapted audio signals generated based on orientation and / or position parameters indicating an orientation towards a rear direction.

[0030] The at least two-channel output audio signals may be binaural audio signals.

[0031] The above component may further be configured to: obtain a user's head orientation and / or position, and wherein the component configured to obtain orientation and / or position parameters may be configured to: process the user's head orientation and / or position to generate orientation and / or position parameters.

[0032] According to a third aspect, there is provided an apparatus for generating a spatial output audio signal, the apparatus including at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the system to at least perform: obtaining a spatial audio signal, the spatial audio signal including: at least two-channel audio signals; and at least one spatial parameter associated with the at least two-channel audio signals; analyzing the at least two-channel audio signals to determine at least one inter-channel characteristic; obtaining orientation and / or position parameters; determining mixing information based on the at least one inter-channel characteristic and the orientation and / or position parameters; and generating at least two-channel output audio signals based on the at least two-channel audio signals, the orientation and / or position parameters, and the mixing information.

[0033] The apparatus configured to generate at least two-channel output audio signals may further be configured to: generate at least two-channel output audio signals based on at least one spatial parameter associated with the at least two-channel audio signals.

[0034] The apparatus configured to determine mixing information may further be configured to: further determine the mixing information based on at least one spatial parameter.

[0035] The apparatus configured to analyze at least two-channel audio signals to determine at least one inter-channel characteristic may be configured to: generate the inter-channel characteristic based on at least one spatial parameter associated with the at least two-channel audio signals.

[0036] At least one spatial parameter associated with the at least two-channel audio signals may include: spatial parameters associated with respective audio-channel audio signals among the at least two audio-channel audio signals; and spatial parameters associated with the at least two audio-channel audio signals.

[0037] The apparatus configured to generate at least two-channel output audio signals based on the at least two-channel audio signals, at least one spatial parameter associated with the at least two-channel audio signals, a directivity parameter, and mixing information may be configured to: generate at least one prototype matrix based on the mixing information; render at least two-channel output audio signals from the at least two-channel audio signals based on: at least one spatial parameter associated with the at least two-channel audio signals, the directivity parameter, the directivity parameter, and at least one prototype matrix.

[0038] The apparatus configured to generate at least two-channel output audio signals based on the at least two-channel audio signals, at least one spatial parameter associated with the at least two-channel audio signals, a directivity parameter, and mixing information may be configured to: process the at least two-channel audio signals based on the mixing information to generate at least two-channel adapted audio signals; render at least two-channel output audio signals from the at least two-channel adapted audio signals based on: at least one spatial parameter associated with the at least two-channel audio signals; and the directivity parameter.

[0039] The apparatus configured to process at least two-channel audio signals based on the mixing information to generate at least two-channel adapted audio signals may be configured to: adapt the at least two-channel audio signals based on a current directivity and an inter-channel characteristic.

[0040] The apparatus configured to adapt at least two channel audio signals based on a current orientation and inter-channel characteristics may be configured to: determine a mono factor based on the current orientation and inter-channel characteristics, the mono factor being configured to indicate how the at least two channel audio signals should be mixed to avoid negative artifacts in the at least two channel output audio signals.

[0041] The apparatus configured to analyze at least two channel audio signals to determine at least one inter-channel characteristic may be configured to analyze the at least two channel audio signals to determine at least one of the following: an inter-channel level difference between the at least two channel audio signals; a modified inter-channel level difference between the at least two channel audio signals, the modification being based on orientation and / or position parameters; an inter-channel phase difference between the at least two channel audio signals; an inter-channel time difference between the at least two channel audio signals; an inter-channel similarity measure between the at least two channel audio signals; and an inter-channel correlation between the at least two channel audio signals.

[0042] The apparatus configured to process at least two channel audio signals based on mixing information to generate at least two channel-adapted audio signals may be configured to: mix the at least two channel audio signals based on an inter-channel difference such that an audio component in substantially one of the at least two channel audio signals is mixed into a corresponding one of the at least two channel-adapted audio signals and further at least partially cross-mixed into another one of the at least two channel-adapted audio signals.

[0043] The apparatus configured to process at least two channel audio signals based on mixing information to generate at least two channel-adapted audio signals may further be configured to: switch at least two of the at least two channel-adapted audio signals based on orientation and / or position parameters indicating an orientation towards a rear direction.

[0044] The at least two channel output audio signals may be binaural audio signals.

[0045] The apparatus may further be configured to: obtain a user head orientation and / or position, and wherein the apparatus configured to obtain orientation and / or position parameters may be configured to: process the user head orientation and / or position to generate orientation and / or position parameters.

[0046] According to a fourth aspect, there is provided an apparatus for generating a spatial output audio signal, the apparatus comprising: means for obtaining a spatial audio signal, wherein the spatial audio signal comprises: at least two channel audio signals; and at least one spatial parameter associated with the at least two channel audio signals; means for analyzing the at least two channel audio signals to determine at least one inter-channel characteristic; means for obtaining a direction and / or position parameter; means for determining mixing information based on the at least one inter-channel characteristic and the direction and / or position parameter; and means for generating at least two channel output audio signals based on the at least two channel audio signals, the direction and / or position parameter, and the mixing information.

[0047] According to a fifth aspect, there is provided an apparatus for generating a spatial output audio signal, the apparatus comprising: an obtaining circuit configured to obtain a spatial audio signal, the spatial audio signal comprising: at least two channel audio signals; and at least one spatial parameter associated with the at least two channel audio signals; an analyzing circuit configured to analyze the at least two channel audio signals to determine at least one inter-channel characteristic; an obtaining circuit configured to obtain a direction and / or position parameter; a determining circuit configured to determine mixing information based on the at least one inter-channel characteristic and the direction and / or position parameter; and a generating circuit configured to generate at least two channel output audio signals based on the at least two channel audio signals, the direction and / or position parameter, and the mixing information.

[0048] According to a sixth aspect, there is provided a computer program [or a computer-readable medium comprising program instructions] comprising instructions [or program instructions] for causing an apparatus for generating a spatial output audio signal to at least perform the following operations: obtaining a spatial audio signal, the spatial audio signal comprising: at least two channel audio signals; and at least one spatial parameter associated with the at least two channel audio signals; analyzing the at least two channel audio signals to determine at least one inter-channel characteristic; obtaining a direction and / or position parameter; determining mixing information based on the at least one inter-channel characteristic and the direction and / or position parameter; and generating at least two channel output audio signals based on the at least two channel audio signals, the direction and / or position parameter, and the mixing information.

[0049] According to a seventh aspect, there is provided a non - transitory computer - readable medium including program instructions for causing an apparatus for generating a spatial output audio signal to at least perform the following operations: obtaining a spatial audio signal, the spatial audio signal including: at least two - channel audio signals; and at least one spatial parameter associated with the at least two - channel audio signals; analyzing the at least two - channel audio signals to determine at least one inter - channel characteristic; obtaining a direction and / or position parameter; determining mixing information based on the at least one inter - channel characteristic and the direction and / or position parameter; and generating at least two - channel output audio signals based on the at least two - channel audio signals, the direction and / or position parameter, and the mixing information.

[0050] According to an eighth aspect, there is provided a computer - readable medium including program instructions for causing an apparatus for generating a spatial output audio signal to at least perform the following operations: The method includes: obtaining a spatial audio signal, the spatial audio signal including: at least two - channel audio signals; and at least one spatial parameter associated with the at least two - channel audio signals; analyzing the at least two - channel audio signals to determine at least one inter - channel characteristic; obtaining a direction and / or position parameter; determining mixing information based on the at least one inter - channel characteristic and the direction and / or position parameter; and generating at least two - channel output audio signals based on the at least two - channel audio signals, the direction and / or position parameter, and the mixing information.

[0051] An apparatus includes components for performing the actions of the method as described above.

[0052] An apparatus is configured to perform the actions of the method as described above.

[0053] A computer program includes program instructions for causing a computer to perform the method as described above.

[0054] A computer - program product stored on a medium can cause an apparatus to perform the method as described herein.

[0055] An electronic device can include the apparatus as described herein.

[0056] A chipset can include the apparatus as described herein.

[0057] Embodiments of the present application are intended to solve problems associated with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] To better understand the present application, reference will now be made, by way of example, to the accompanying drawings, in which:

[0059] Figure 1 Schematically shows an example system suitable for capturing and playing spatial audio signals for implementing some embodiments;

[0060] Figure 2 Shows the operation of an example system of a capture and playback spatial audio signal capture device as shown in Figure 1 a flowchart;

[0061] Figure 3 Schematically shows an example device system suitable for implementing some embodiments;

[0062] Figure 4 Schematically shows, as in Figure 1 an example playback device as shown in

[0063] Figure 5 Shows the operation of an example playback device as shown in Figure 4 a flowchart;

[0064] Figure 6 Schematically shows, as in Figure 4 an example spatial processor as shown in

[0065] Figure 7 Shows the operation of an example spatial processor as shown in Figure 6 a flowchart;

[0066] Figure 8 Schematically shows, as in Figure 6 an example transmission signal adapter as shown in

[0067] Figure 9 Shows the operation of an example transmission signal adapter as shown in Figure 8 a flowchart;

[0068] Figure 10 Shows an example processing output;

[0069] Figure 11 Schematically shows another example capture and playback system of a device suitable for implementing some embodiments. DETAILED DESCRIPTION

[0070] Suitable devices and possible mechanisms for rendering a suitable output audio signal from a parameterized spatial audio stream (or signal) from the captured or otherwise obtained audio signal will be described in further detail below.

[0071] As discussed above, Metadata-Assisted Spatial Audio (MASA) is an example of a parameterized spatial audio format and representation suitable as an input format for IVAS.

[0072] It can be considered as an audio representation consisting of "N channels + spatial metadata". It is a scene-based audio format, especially suitable for spatial audio capture on practical devices such as smart phones. The idea is to describe the sound scene in terms of the direction of the sound varying in time-frequency and, for example, the energy ratio. The sound energy not defined (described) by the direction is described as diffuse (coming from all directions).

[0073] As discussed above, the spatial metadata associated with the audio signal can include multiple parameters per time-frequency block (such as multiple directions and the direct-to-total ratio, spread coherence, distance, etc. associated with each direction (or direction value)). The spatial metadata can also include other parameters, or can be associated with other parameters considered non-directional (such as surround coherence, diffuse-to-total energy ratio, remainder-to-total energy ratio), but when combined with the directional parameters, can be used to define the characteristics of the audio scene. For example, a reasonable design choice that can produce good quality output is to determine that the spatial metadata includes one or more directions (and the direct-to-total ratio, spread coherence, distance values, etc. associated with each direction) for each time-frequency portion.

[0074] As described above, the parameterized spatial metadata representation can use multiple concurrent spatial directions. With MASA, the proposed maximum number of concurrent directions is two. For each concurrent direction, there can be associated parameters such as: direction index; direct-to-total ratio; spread coherence; and distance. In some embodiments, other parameters are defined, such as diffuse-to-total energy ratio; surround coherence; and remainder-to-total energy ratio.

[0075] The parameterized spatial metadata values can be used for each time-frequency map block (the MASA format defines 24 frequency bands and 4 time sub-frames per frame). The frame size in IVAS is 20 ms. In addition, currently MASA supports 1 or 2 directions for each time-frequency map block.

[0076] Example metadata parameters can be:

[0077] Format descriptor, which defines the MASA format for IVAS;

[0078] Channel audio format, which defines the combined subsequent fields stored in two bytes;

[0079] Number of directions, which defines the number of directions described by the spatial metadata (each direction is associated with spatial metadata related to a set of directions, as described below);

[0080] Number of channels, which defines the number of transmission channels in the format;

[0081] Source format, which describes the original format in which the MASA is created.

[0082] Examples of MASA format spatial metadata parameters related to the number of directions can be:

[0083] Direction index, which defines the direction of arrival of the sound in the time-frequency parameter interval. (Typically, this is a spherical representation with an accuracy of approximately 1 degree);

[0084] Direct-to-total energy ratio, which defines the energy ratio for the direction index (i.e., the time-frequency subframe);

[0085] Extended coherence, which defines the energy spread for the direction index (i.e., the time-frequency subframe).

[0086] Examples of MASA format spatial metadata parameters independent of the number of directions can be:

[0087] Diffuse-to-total energy ratio, which defines the energy ratio of non-directional sound over the surrounding directions;

[0088] Ambient coherence, which defines the coherence of non-directional sound over the surrounding directions;

[0089] Residual-to-total energy ratio, which defines the energy ratio of the residual (such as microphone noise) sound energy to meet the requirement that the sum of the energy ratios is 1.

[0090] In addition, example spatial metadata bands can be:

[0091]

[0092]

[0093] The MASA stream can be rendered into various outputs, such as a multi-channel loudspeaker signal (e.g., 5.1) or a binaural signal.

[0094] In Vilkamo, J., An example rendering method is described in "Optimized covariance domain framework for time–frequency processing of spatial audio" by T. and Kuntz, A. (Journal of the Audio Engineering Society, Vol. 61, No. 6, pp. 403 - 411, 2013). The rendering method is based on multichannel mixing. The method processes a given audio signal in a frequency band such that a desired covariance matrix is obtained for the output signal in the frequency band. The covariance matrix contains the channel energies of all channels and the inter-channel relationships between all pairs of channels, i.e., the cross-correlations and the inter-channel phase differences. These features are known to convey the perceptually relevant spatial features of multichannel sound in various playback scenarios, such as binaural for headphones, surround speakers, Ambisonics, and crosstalk-cancelled stereo.

[0095] The above rendering method uses a prototype signal (or a prototype matrix that provides a prototype signal based on the input signal). The prototype signal or matrix can be frequency-invariant or frequency-varying, depending on the use case. A prototype signal is a signal that provides an example of what the signal content of a given output channel should be.

[0096] This information is needed because the covariance matrix only represents the spatial image but cannot represent what signals arrive from different directions. For example, if in a frequency band there is a tone in one direction and narrowband noise in another direction, the covariance matrix may be the same or highly similar even though the channels will be opposite, assuming the signals have the same energy. Therefore, the rendering method uses a prototype matrix or prototype signal to guide the rendering of the spatial output. The rendering method discusses providing an output with the desired covariance matrix properties but making the output signal waveform maximally similar to the prototype signal.

[0097] In addition, there are other parametric rendering schemes. Nevertheless, they generally use a prototype signal in some way for rendering (i.e., the rendering is based on the transmitted audio signal) and modify it in the frequency band based on spatial metadata to obtain the desired spatial audio signal (such as a binaural signal).

[0098] The above examples use the terms "prototype signal" and "prototype matrix". Generally speaking, these terms refer to preprocessing the transmitted audio signal to provide an audio signal suitable for spatial audio rendering. In some examples, the prototype signal is the transmitted audio signal, in other words, no processing is done to the transmitted audio signal to generate the prototype signal.

[0099] The embodiments discussed in this document focus on head-tracking binaural reproduction (however, other embodiments may use these methods for other multi-channel reproduction formats without significant creative effort). In the following examples, the transmitted audio signal (the audio signal generated from the capture device) may be a two-channel transmission signal, where the left channel contains sounds that are predominantly on the left side within the acoustic audio environment, and the right channel contains sounds that are predominantly on the right side within the acoustic audio environment. For example, these signals may be obtained from two coincident cardioids pointing left and right. Such signals are generally conducive to generating binaural signals. The left and right binaural audio channels may be synthesized mainly based on the corresponding left and right transmission signals. The binaural cues required for spatial processing are synthesized, and the fine spectral content of the left and right ears tends to follow the spectral content of the transmitted audio signal.

[0100] When the listener or user of the playback device rotates their head by more than 90 degrees, the left transmitted audio channel signal becomes closer to / more similar to the sound intended for the right ear, and vice versa. Using such a signal as a starting point for spatial synthesis, the above-mentioned rendering method can render an appropriate covariance matrix for the binaural signal, but performs poorly in many cases because the fine spectral content of the left and right binaural signals does not match the expected content well. The sound can further acquire vocoder-like characteristics because even though the channel energies are properly synthesized, the fine spectral content still mainly comes from the wrong source.

[0101] Although (as disclosed in, for example, GB2007904.2) when the user looks towards 180 degrees close to the original viewing direction (i.e., they are looking in the "backward" direction), the left and right transmission channels can be flipped to improve performance, this flipping of the transmission channels performs poorly in other directions, such as when the user is looking towards directions close to ±90 degrees.

[0102] For example, stereo transmission sound is obtained using two cardioids pointing left and right respectively. This means that any sound coming directly from the left or right will be in only one of these channels. This is a situation where channel flipping does not help because one of these transmission signals does not contain the previously mentioned signal at all. In the case where the sound source is at 90 degrees and the user's head is oriented at 90 degrees, the sound will be rendered approximately in the center, i.e., at the same level in both ears. The spatial renderer synthesizes such binaural cues, but it may do so by amplifying the wrong signal content because that particular signal content may be missing in one of these channels. In other words, the rendering method as shown above is given a very poor starting point for rendering the binaural output, and in such cases, the perceived sound quality is usually very poor.

[0103] Although it is possible to mix the transmission signals into a dual-mono signal before spatial synthesis and thus the desired signals will always be available for both binaural outputs, this would also mean losing the inherent incoherence between the transmission channels, which is required for rendering the environment (or generally, the width) without using a large amount of audio decorrelation processing, which is detrimental to the sound quality of signals such as applause / cheers or speech.

[0104] IVAS use cases (e.g., MASA format) complicate the situation because the cardioid example is only one of many potential transmission signal format types. The transmission signal can, for example, be a downmix of a 5.1-channel format sound or can be generated from spaced microphones with or without significant directional characteristics.

[0105] As discussed generally in the following embodiments and concepts of this application, an effective method is implemented for adapting a transmission audio signal for spatial audio rendering to be suitable for any head orientation and any transmission signal type. Thus, the sound quality produced in this way will be superior under certain head orientations and / or certain transmission signal types. Therefore, these embodiments create a good user experience because the sound quality is maintained regardless of the user's head position / rotation.

[0106] In summary, the concepts, as will be further discussed in detail in the following embodiments, relate to head-tracking binaural rendering of parametric spatial audio consisting of spatial metadata and a transmission audio signal. In some embodiments, this can be that the transmission audio signal is at least two different types. In such embodiments, a binaural renderer is provided that can render binaural audio from the transmission audio signal and the spatial metadata, thereby achieving high-quality (accurate direction reproduction and no significant additional noise) head-tracking binaural audio rendering at any head orientation from a transmission audio signal (having at least 2 channels) with arbitrary inter-channel characteristics (such as the direction pattern and spacing of microphones). In some embodiments, this can be achieved by the following operations: determining the inter-channel characteristics based on an analysis of the transmission audio signal (such as the level difference in frequency bands), and then determining the mixing information based on the determined inter-channel characteristics and the orientation of the head. Further, this mixing information can enable the mixing of the transmission audio signal to obtain two audio signals (sometimes referred to as "prototype signals") that represent the appropriate audio signal content for the left and right output channels. Further, these embodiments can also be configured to perform binaural audio rendering using the determined mixing information, head orientation, and spatial metadata.

[0107] As described in further detail herein, there are at least two ways in which the mixing information can be used during binaural audio rendering. In some embodiments, the mixing information can be used to preprocess the transmitted audio signal to make it suitable for spatial audio rendering for the current head orientation and the determined inter-channel characteristics. This method will be described in detail in the following exemplary embodiments.

[0108] Alternatively, in some embodiments, the mixing information is used as a prototype matrix during spatial rendering. This is the case, for example, when using the rendering method discussed above and using the mixing information as a prototype matrix, and making the left and right binaural audio signals similar to a preprocessed version of the transmitted audio signal in terms of the desired fine spectral content, but the preprocessed version of the transmitted audio signal is not actually generated as a separate intermediate signal (in the program memory). This method will be further described herein.

[0109] In the description herein, the term "audio signal" can refer to an audio signal having one channel or an audio signal having multiple channels. When it is involved that the specified signal has one or more channels, it will be clearly stated. In addition, the term "audio signal" can mean that the signal is in any form, such as encoded or non-encoded form, for example, a series of values that define the signal waveform or spectral values.

[0110] Embodiments will be described with respect to an exemplary capture (or encoder / analyzer) and playback (or decoder / synthesizer) device or system 150 as shown in Figure 1 .

[0111] In the following example, the audio signal input is an audio signal input from a microphone array, but it will be understood that the audio input can be any suitable audio input format, and the following description will detail where differences in processing occur when different input formats are used.

[0112] System 150 is shown as having a capture part and a playback (decoder / synthesizer) part.

[0113] In some embodiments, the capture part includes a microphone array audio signal input 100. This input audio signal can come from any suitable source, such as: two or more microphones installed on a mobile phone, other microphone arrays, such as, B-format microphones or Eigenmikes. In some embodiments, as described above, the input can be any suitable audio signal input, such as an Ambisonic signal, for example, first-order Ambisonics (FOA), higher-order Ambisonics (HOA) or speaker surround mixing and / or objects.

[0114] The microphone array audio signal input 100 can be provided to the microphone array front end 101. In some embodiments, the microphone array front end 101 is configured to implement an analysis processor function that is configured to generate or determine suitable (spatial) metadata 104 associated with the audio signal, and to implement a suitable transmission signal generator function to generate a transmission audio signal 102.

[0115] Thus, the analysis processor function is configured to perform a spatial analysis on the input audio signal to generate suitable spatial metadata 104 in a frequency band. For all of the above input types, there are known methods to generate suitable spatial metadata in a frequency band, e.g., direction and direct-to-total energy ratio (or similar parameters such as diffuseness, i.e., ambient-to-total ratio). These methods are not detailed herein, however, some examples can include performing a suitable time-frequency transform on the input signal and then estimating, when the input is a mobile phone microphone array, the delay values between microphone pairs that maximize the inter-microphone correlation in a frequency band, formulating corresponding direction values for the delays (as described in GB patent application No. 1619573.7 and PCT patent application No. PCT / FI2017 / 050778), and formulating ratio parameters based on the correlation values.

[0116] The metadata can take various forms and in some embodiments includes spatial metadata and other metadata. A typical parameterization for spatial metadata is a direction parameter in each frequency band, which is characterized by an elevation value φ(k,n) and an azimuth value θ(k,n) and an associated direct-to-total energy ratio r(k,n) in each frequency band, where k is the frequency band index and n is the time frame index.

[0117] In some embodiments, the generated parameters vary by frequency band. Thus, for example, in frequency band X, all parameters are generated and sent, while in frequency band Y, only one parameter is generated and sent, and furthermore, in frequency band Z, no parameters are generated or sent. A practical example can be that for some frequency bands such as the highest frequency band, some of these parameters are not needed for perceptual reasons.

[0118] In some embodiments, the microphone array front end 101 can use a machine learning model to determine the spatial metadata 104 based on the microphone array signal 100, as described in NC322440 and NC322439.

[0119] Thus, the output of the analysis of the processor functionality is the (spatial) metadata 104 determined in the time-frequency map blocks. The (spatial) metadata 104 can relate to direction and energy ratios in a frequency band, but can also have any of the metadata types listed above. The (spatial) metadata 104 can vary over time and frequency.

[0120] In some embodiments, the analysis functionality is implemented external to the system 150. For example, in some embodiments, the spatial metadata associated with the input audio signal can be provided to the encoder 103 as a separate bitstream. In some embodiments, the spatial metadata can be provided including a set of spatial (direction) index values.

[0121] As described above, the microphone array front end 101 is further configured to implement a transmission signal generator functionality to generate a suitable transmitted audio signal 102. The transmission signal generator functionality is configured to receive an input audio signal (which can be, for example, the microphone array audio signal 100) and generate the transmitted audio signal 102. The transmitted audio signal can be a multi-channel, stereo, binaural or mono audio signal. The generation of the transmitted audio signal 102 can be implemented using any suitable method.

[0122] In some embodiments, the transmitted signal 102 is the input audio signal, e.g., the microphone array audio signal. The number of transmission channels can also be any suitable number (other than one or two channels discussed in the examples).

[0123] In some embodiments, the transmitted signal 102 is determined based on the kind or type of the input microphone array signal. For example, if the microphone array signal 100 is from a mobile device, the microphone array front end 101 is configured to select the microphone signal from the left side of the device as the left transmitted signal and select another microphone signal from the right side of the device as the right transmitted signal. As another example, a dedicated microphone array can be used to capture the audio signal, in which case the transmitted audio signal 102 can have been captured by the dedicated microphone.

[0124] In some embodiments, the microphone array front end 101 is configured to apply any suitable preprocessing steps, such as equalization, microphone noise suppression, wind noise suppression, automatic gain control, beamforming and other spatial filtering, ambient noise suppression, and limiting. The transmitted audio signal 102 can have any kind of directional characteristic, e.g., having an omnidirectional or cardioid-like directional pattern.

[0125] In some embodiments, the capture portion may include an encoder 103. The encoder 103 may be configured to receive a transmitted audio signal 102 and spatial metadata 104. The encoder 103 may also be configured to generate a bitstream 106 that includes the metadata information and the transmitted audio signal in an encoded or compressed form.

[0126] For example, the encoder 103 may be implemented as an IVAS encoder or any other suitable encoder. In such an embodiment, the encoder 103 is configured to encode the audio signal and the metadata and form an IVAS bitstream. The bitstream 106 includes the transmitted audio signal 102 and the spatial metadata 104 in an encoded form. The transmitted audio signal 102 may be encoded, for example, using an IVAS core codec, EVS, or an AAC encoder (or any other suitable encoder), and the metadata 104 may be encoded, for example, using the methods proposed in GB1811071.8, GB1913274.5, PCT / FI2019 / 050675, GB2000465.1 (or any other suitable method).

[0127] Furthermore, the bitstream 106 may be sent / stored.

[0128] In addition, the system 100 may include a player or decoder 105 portion. The player or decoder 105 is configured to receive, retrieve, or otherwise obtain the bitstream 106 and generate a suitable spatial audio signal 110 from the bitstream for presentation to a listener / listener playback device.

[0129] Thus, the decoder 105 is configured to receive the bitstream 106 and demultiplex the encoded stream, and then decode the audio signal and the metadata to obtain the transmitted signal and the metadata. In some embodiments, the decoder 105 may be an IVAS decoder (or any other suitable decoder). The decoder 105 may also receive, for example, head orientation 108 information from a head tracker, which the decoder may use during rendering to reproduce an output spatial audio signal 110 from the transmitted audio signal and the spatial metadata. For example, a binaural audio signal may be reproduced on headphones, especially in the case of binaural rendering. The decoder 105 and the encoder 103 may be implemented in different devices or within the same device.

[0130] Regarding Figure 2 , a flowchart of operations implemented by the Figure 1 device system shown in

[0131] Therefore, as shown by 201, the first operation is to obtain a microphone array audio signal.

[0132] Then, as shown by 203, this step is to generate a transmission audio signal and spatial metadata from the microphone array audio signal.

[0133] The next operation is shown by 205, which is to encode the transmission audio signal and spatial metadata to generate a bitstream.

[0134] In addition, the operation of obtaining head orientation information is shown by 206.

[0135] Furthermore, as shown by 207, the bitstream is decoded, and a (binaural) spatial audio signal is rendered based on the decoded transmission audio signal, spatial metadata, and head orientation information.

[0136] Finally, as shown by 209, the rendered spatial audio signal is output.

[0137] Regarding Figure 3 , an example (playback) device for implementing some embodiments is shown. In the example shown in Figure 3 , it is shown that a mobile phone 301 is coupled to headphones 321 worn by a user of the mobile phone 301 via a wired or wireless connection 307. Hereinafter, the example device or apparatus is the mobile phone as shown in Figure 3 . However, the example device or apparatus can also be any other suitable device, such as a tablet computer, a laptop computer, a computer, or any conference phone device. In addition, the device or apparatus itself can be headphones, such that the operations of the example mobile phone 301 are performed by the headphones.

[0138] In this example, the mobile phone 301 includes a processor 315. The processor 315 can be configured to execute various program codes, such as the methods described herein. The processor 315 is configured to communicate with the headphones 321 using a wired or wireless headphone connection 307. In some embodiments, the wired or wireless headphone connection 307 is a Bluetooth 5.3 or Bluetooth LE audio connection. The connection 307 provides a (dual-channel) audio signal 304 from the processor 315 to be reproduced for the user by the headphones 321.

[0139] The headphones 321 can be the over-ear headphones as shown in Figure 1 , or any other suitable type, such as in-ear headphones, bone conduction headphones, or any other type of headphones. In some embodiments, the headphones 321 have a head orientation sensor for providing head orientation information to the processor 315. In some embodiments, the head orientation sensor is separate from the headphones 321, and the data is provided to the processor 315 separately. In other embodiments, head orientation is tracked by other means / components, such as using the device 301 camera and machine learning-based face orientation analysis.

[0140] In some embodiments, the processor 315 is coupled to a memory 303 that has program code 305 that provides processing instructions in accordance with the following embodiments. The program code 305 has instructions for processing a transmitted audio signal received by the transceiver 313 or retrieved from the storage device 311 to obtain a rendering suitable for efficient output to the headphones.

[0141] The transceiver 313 may communicate with other devices via any suitable known communication protocol. For example, in some embodiments, the transceiver may use a suitable radio access architecture based on: Long Term Evolution Advanced (LTE-A), or New Radio (NR) (which may also be referred to as 5G), Universal Mobile Telecommunications System (UMTS) Radio Access Network (UTRAN or E-UTRAN), Long Term Evolution (LTE, same as E-UTRA), 2G network (legacy network technology), Wireless Local Area Network (WLAN or Wi-Fi), Worldwide Interoperability for Microwave Access (WiMAX), Personal Communications Service (PCS), Wideband Code Division Multiple Access (WCDMA), systems using Ultra-Wideband (UWB) technology, sensor networks, Mobile Ad-hoc Networks (MANET), Cellular Internet of Things (IoT) RAN, and Internet Protocol Multimedia Subsystem (IMS), any other suitable option, and / or any combination thereof.

[0142] A remote capture device (or encoder device) configured to generate an encoded audio bitstream may be a system similar or identical to the Figure 3 devices and headphone systems shown. In such capture device or equipment, the spatial audio signal is the transmitted audio signal and metadata that are encoded, passed to the transceiver or stored in the storage device, and then provided to the playback device or device processor for decoding and rendered as binaural spatial sound, which (using a wired or wireless headphone connection) is forwarded to the headphones for reproduction to the listener (user).

[0143] In some embodiments, the device (operable to capture or play or both) includes a user interface (not shown) which, in some embodiments, may be coupled to a processor. In some embodiments, the processor may control the operation of the user interface and receive input from the user interface. In some embodiments, the user interface may enable a user to input commands to the device, for example, via a keyboard. In some embodiments, the user interface may enable a user to obtain information from the device. For example, the user interface may include a display configured to display information from the device to the user. In some embodiments, the user interface may include a touchscreen or touch interface that can enable information to be input into the device and can also display information to a user of the device. In some embodiments, the user interface may be a user interface for communication.

[0144] Regarding Figure 4 , a schematic diagram of the processor 103 with respect to the decoder 105 is shown, where the encoded bitstream is processed to generate spatial audio (e.g., binaural audio signals) suitable for the headphones 321.

[0145] In some embodiments as shown in Figure 1 , the decoder 105 is configured to receive as input a bitstream 402 obtained from a capture / encoder device (which may be the same device or a remote device of the device) (which is reference 106 in Figure 1 and reference 302 in Figure 3 ).

[0146] In addition, in some embodiments, the decoder 105 may be configured to receive or otherwise retrieve head orientation information 400 (which is reference 108 in Figure 1 and reference 306 in Figure 3 ).

[0147] In some embodiments, the decoder includes a demux (demultiplexer) and a decoder 401 which demultiplex and decode the bitstream 402 into two streams: a transmitted audio signal 404 and spatial metadata 406. The decoding corresponds to the encoding applied in the encoder 103 as shown in Figure 1 . It should be noted that the decoded transmitted audio signal 404 and spatial metadata 406 may be different from those before encoding and decoding, but are substantially or in principle the same as the transmitted audio signal 102 and spatial metadata 104 presented in Figure 1 and described above. Any differences are due to errors introduced during encoding or decoding or in the transmission channel. Nevertheless, for simplicity, the same terms are used hereinafter to refer to these signals.

[0148] The transmitted audio signal 404, the spatial metadata 406, and the head orientation signal 400 can be received by a spatial synthesizer 403 configured to synthesize a spatial audio output 408 in a desired format (which is reference 110 in Figure 1 and reference 304 in Figure 3 ). For example, the output can be a binaural audio signal.

[0149] Regarding Figure 5 , an example flowchart showing the operation of the processor as shown in Figure 4 is presented.

[0150] Thus, a first operation can include, as shown by 501, obtaining the head orientation signal and an encoded spatial audio bitstream.

[0151] Then, as shown by 503, the encoded spatial audio bitstream is demultiplexed and decoded to generate the transmitted audio signal and the spatial metadata.

[0152] After that, as shown by 505, a spatial audio signal is synthesized from the transmitted audio signal based on the spatial metadata and the head orientation information.

[0153] Furthermore, as shown by 507, the spatial audio signal is output (e.g., a binaural audio signal is output to headphones).

[0154] Regarding Figure 6 , the spatial synthesizer 403 of Figure 4 is shown in further detail. In some embodiments, the spatial synthesizer 403 is configured to receive the transmitted audio signal 404, the head orientation 400, and the spatial metadata 406.

[0155] In some embodiments, the head orientation 400 takes the form of a rotation matrix that represents the rotation to be performed on a direction vector to compensate for head rotation. In some embodiments, if the head orientation takes another form (such as conventional yaw, pitch, roll), the head orientation information can be converted to a rotation matrix R(n), where n is a time index and the angles are given in radians by:

[0156] c1 = cos(-roll)

[0157] c2 = cos(-pitch)

[0158] c3 = cos(-yaw)

[0159] s1 = sin(-roll)

[0160] s2 = sin(pitch)

[0161] s3 = sin(-yaw)

[0162]

[0163] It should be noted that the signs and order of the angles are merely conventions based on the determined axes of rotation and order of rotation. Other equivalent transformations can be created similarly. Additionally, the rotation matrix can be obtained from quaternions or direction cosine matrices, which are also commonly used in representing the tracked orientation.

[0164] In some embodiments, the spatial synthesizer 403 includes a forward filter bank 601. The transmitted audio signal 404 is provided to the forward filter bank 601, which transforms the transmitted audio signal into a time-frequency representation, the time-frequency transmitted audio signal 600. Any filter bank suitable for audio processing can be used, such as a complex modulated quadrature mirror filter (QMF) bank or its low-latency variant, or a short-time Fourier transform (STFT). Similarly, the forward filter bank 601 can be implemented by any suitable time-frequency transformer. In the example described herein, the forward filter bank 601 is configured to have 60 frequency bins and sufficient stopband attenuation to avoid significant aliasing when processing the frequency bin signals. In this configuration, all frequency bins can be processed independently of each other, but some frequency bins share the same spatial metadata. For example, the spatial metadata 406 can include spatial parameters in a finite number of frequency bands (e.g., 5 bands), and each of these bands corresponds to a set of one or more frequency bins provided by the forward filter bank 601. Although there are 5 bands in this example, there can be any suitable number of bands, e.g., the number of bands can be 8, 12, 18, or 24 bands. The time-frequency transmitted signal S(b,t,i) can be labeled in vector or scalar form as:

[0165]

[0166] where b is the frequency bin index, t is the time-frequency signal time index, and i is the channel index.

[0167] In some embodiments, the spatial synthesizer 403 includes a transmitted signal adapter 607. The transmitted signal adapter 607 is configured to receive the time-frequency transmitted audio signal 600 and the head orientation 400 information or signal or data.

[0168] The transmitted signal adapter 607 is configured to process the time-frequency transmitted audio signal 600 based on the head orientation 400 data to provide an adapted time-frequency transmitted audio signal 606 that is "more favorable" for the current head orientation for subsequent spatial synthesis processing. The adapted time-frequency transmitted audio signal 606 can be labeled, for example, as:

[0169]

[0170] The adapted time-frequency transmitted audio signal 606 can be provided to the decorrelator and mixer 611 block, the processing matrix determiner 609, and the input and target covariance matrix determiner 605.

[0171] In some embodiments, the spatial synthesizer 403 includes a spatial metadata rotator 603. The spatial metadata rotator 603 is configured to receive spatial metadata 406 and head orientation data 400 (which, for this example, takes the form of a derived rotation matrix R(n)).

[0172] In some embodiments, the spatial metadata rotator 603 is configured to convert the direction parameters of the spatial metadata into vector form (if they are not provided in this format).

[0173] For example, if the direction parameters consist of an azimuth angle θ(k,n) and an elevation angle where k is the frequency band index, they are converted by the following equation:

[0174]

[0175] The spatial metadata rotator 603 is configured to rotate the direction vector v DOA (k,n) by the rotation matrix R(n):

[0176]

[0177] In some embodiments, the rotated matrix can then be converted into a rotated spatial metadata direction by the following equation:

[0178] θ R (k,n) = atan2(v y (k,n), v x (k,n))

[0179]

[0180] The rotated spatial metadata 602 is otherwise the same as the original spatial metadata 406, but with the rotated direction parameters θ R (k,n) and replacing the original direction parameters θ(k,n) and In effect, this rotation compensates for head rotation by rotating the direction parameters in the opposite direction.

[0181] In some embodiments, the spatial synthesizer 403 includes an input and target covariance matrix determiner 605. The input and target covariance matrix determiner 605 is configured to receive the rotated spatial metadata 602 and the adapted time-frequency transmission signal 606, and it determines a covariance matrix 604, which includes an input covariance matrix representing the adapted time-frequency transmission audio signal 606 and a target covariance matrix representing the time-frequency spatial audio signal 610 (to be rendered). The input covariance matrix can be measured from the adapted time-frequency transmission signal 606 (labeled as the column vector x(b,t)), where the rows indicate the transmission signal channels. This can be achieved by the following formula:

[0182]

[0183] where the superscript H indicates the conjugate transpose, and t1(n) and t2(n) are the first and last time-frequency signal time indices corresponding to frame n (or sub-frame n in some embodiments). In this example, there are four time indices t at each frame n. However, the number of time indices can be more than four or less than four. In some embodiments, the covariance matrix is determined for each bin as described above. In other embodiments, the covariance matrix is also averaged (or summed) over multiple frequency bins at a resolution close to human auditory resolution, or at the resolution of the determined spatial metadata parameters, or at any suitable resolution.

[0184] In some embodiments, the target covariance matrix is determined based on the spatial metadata and the total signal energy. The total signal energy E o (b,n) can be obtained, for example, as the average or sum of the diagonal values of C x (b,n). Further, in one example, the spatial metadata consists of the rotated direction parameter θ R (k,n) and and the direct total energy ratio parameter r(k,n). In this example, the frequency band index k is the index where the bin b is located.

[0185] In some embodiments where the output is a binaural signal, the target covariance matrix can be determined by the following formula:

[0186]

[0187] where, is for bin b, azimuth angle θ R (k,n) and elevation angle The column vector of the head-related transfer function, and it is a column vector of length two with complex values, where these values correspond to the HRTF amplitudes and phases for the left and right ears. At high frequencies, the HRTF values can also be real-valued because phase differences are not required for perceptual reasons at high frequencies. Obtaining the HRTF for a given direction and frequency is known. C d (b) is the diffuse-field binaural covariance matrix, which can be determined, for example, in an offline phase by taking a set of spatially uniform HRTFs, independently formulating their covariance matrices, and averaging the results.

[0188] Input covariance matrix C x (b,n) and the target covariance matrix C y (b,n) can be output as covariance matrix 604.

[0189] The above examples have considered direction and ratio. However, more generally, generating the target covariance matrix can be performed based on GB2572650, in which, in addition to direction and ratio, spatial coherence parameters and output types other than binaural outputs are also described.

[0190] In some embodiments, the spatial synthesizer 403 includes a processing matrix determiner 609. The processing matrix determiner 609 is configured to receive the covariance matrices C x (b,n) and C y (b,n) 604 and the adapted time-frequency transmitted audio signal 606, and determine the processing matrices M(b,n) and M r (b,n). In some embodiments, determining the processing matrix based on the covariance matrix can be based on Juha Vilkamo, Tom and Achim Kuntz's "Optimized covariance domain framework for time-frequency processing of spatial audio" (Journal of the Audio Engineering Society, Vol. 61, No. 6, pp. 403-411, 2013).

[0191] In this method, the processing matrix 608 is determined as a mixing matrix for processing the input audio signal with the measured covariance matrix C x (b,n) such that the output audio signal (the processed input audio signal) reaches the determined target covariance matrix C y (b,n).

[0192] This method can be used in various use cases, including the generation of binaural or surround speaker signals. When formulating the processing matrix, this method can further implement a prototype matrix, which includes a matrix identifying which signals are typically intended for each output in the optimization process (with the constraint that the output must reach the target covariance matrix). In the examples described herein, the generation of such appropriate signals has been implemented in the transmission signal adapter 607, and thus the prototype matrix can be simply represented, for example, as or Furthermore, the processing matrix determiner 609 can be configured to output the processing matrices 608M(b,n) and M r (b,n).

[0193] In some embodiments, the spatial synthesizer 403 includes a decorrelator and a mixer 611. The decorrelator and the mixer 611 are configured to receive the adapted time-frequency transmitted audio signal x(b,t) 606 and the processing matrices 608M(b,n) and M r (b,n). The decorrelator and the mixer 611 are configured to first process the adapted time-frequency audio signal 606 using the decorrelator to generate a decorrelated signal x D (b,t). Then, the decorrelator and the mixer 611 are configured to apply a mixing process to generate the time-frequency spatial audio signal 610:

[0194] y(b,t) = M(b,n)x(b,t) + M r (b,n)x D (b,t)

[0195] In the above, although not explicitly written in the above equation, the processing matrix can be linearly interpolated between frames n such that at each time index of the time-frequency signal, the matrix advances from M(b,n - 1) towards M(b,n). If a start is detected (fast interpolation) or no start is detected (normal interpolation), the interpolation rate can be adjusted. Furthermore, the time-frequency spatial audio signal 610y(b,t) can be output.

[0196] In some embodiments, the spatial synthesizer 403 includes an inverse filter bank 613, which is configured to apply an inverse transform corresponding to the transform used by the forward filter bank 601 to convert the time-frequency spatial audio signal 610 into a spatial audio output 408 (which is a binaural audio signal in this example).

[0197] Regarding Figure 7 , an example flowchart showing the operation of the spatial synthesizer as shown in Figure 6 is shown according to some embodiments.

[0198] Thus, as shown by 701, the first operation may include obtaining a head orientation signal and transmitting an audio signal and spatial metadata.

[0199] Then, as shown by 703, the transmitted audio signal is subjected to a time-frequency transform to generate a time-frequency transmitted audio signal.

[0200] After that, as shown by 707, the time-frequency transmitted audio signal is adapted based on the head orientation information.

[0201] In addition, as shown by 705, the spatial metadata is rotated based on the head orientation.

[0202] As shown by 709, the input and target covariance matrices are determined from the adapted time-frequency audio signal. In some embodiments, the target covariance matrix is also determined based on the rotated spatial metadata.

[0203] Furthermore, as shown by 711, a processing matrix is determined from the input and target covariance matrices.

[0204] As shown by 713, based on the processing matrix, the adapted transmitted audio signal is decorrelated and mixed.

[0205] Then, as shown by 715, an inverse time-frequency transform is performed on the time-frequency spatial audio signal.

[0206] Furthermore, as shown by 717, a spatial audio signal is output.

[0207] Figure 8 Shown in further detail is the transmission signal adapter 607 as Figure 6 shown. As described above, the general concept of the operation of the transmission signal adapter 607 is that when the listener is looking straight ahead, for example, the transmitted audio signal is directly suitable for rendering because the head is basically in the same pose as the capture device when capturing spatial audio. In other words, sounds that are mainly on the left side are mainly in the left transmission signal, and the same is true for sounds on the right side.

[0208] However, when the listener is looking, for example, at ±90 degrees, the transmitted audio signal can be adapted for subsequent rendering operations based on the inter-channel characteristics of the transmission signal. In these embodiments, when the level difference between the channels is small, both signals may contain all the sources of the sound scene, and likewise, the transmitted audio signal is not modified. For example, the transmitted audio signal may originate from a substantially omnidirectional microphone pair, such as two microphones integrated into the left and right edges of a mobile phone.

[0209] However, if the inter-channel level difference is large, one of the channels may not contain at least some sources of the sound scene, which will result in a quality degradation during rendering if they are used to perform rendering when the head orientation is, for example, ±90 degrees in the yaw direction. For example, the transmitted audio signal may originate from a pair of cardioid microphones facing opposite directions, and the associated sound source (e.g., a speaker) may be located at or near the maximum attenuation direction of one of these cardioid patterns. In this case, the sound of this speaker will be rendered at the center (i.e., in front or behind, since the head is yawed at ±90 degrees). However, the signal of this speaker only appears at one of the transmitted channels. This affects the accuracy of subsequent rendering operations that mainly generate left and right binaural channels from the corresponding left and right transmitted audio signals.

[0210] Therefore, in this example, when the speaker signal is active, the audio should be cross-mixed to ensure that the specific signal content (the speaker signal in this example) appears at both channels so that rendering can be performed without the artifacts mentioned above. Similarly, when it is determined that cross-mixing is not required, it is not performed. For example, when the user looks at ±90 degrees but the sound scene includes applause / cheers, cross-mixing should not be performed. In this case, in this example, the channel content is kept completely separate at the transmitted signal adapter 607 because then the subsequent spatial audio renderer can generate the appropriate incoherence for the applause / cheers without having to use a large amount of decorrelators to recover the loss of inter-channel incoherence, which is a side effect of the cross-mixing process.

[0211] In some embodiments, the transmitted signal adapter 607 is configured to receive the time-frequency transmitted audio signal 600 (labeled as S(b,t,i), where b is the frequency bin index, t is the sample time index, and i is the channel index) and the head orientation data 400. In some embodiments, the transmitted signal adapter 607 includes an inter-channel level difference (ILD) determiner 801. The ILD determiner 801 is configured to receive the time-frequency transmitted audio signal 600 and determine the inter-channel level difference (ILD) between the channels of the time-frequency transmitted audio signal. In some embodiments, this can be determined by the following operations. First, for example, the energy of a channel is calculated by the following formula:

[0212] E(b,t,i) = |S(b,t,i)| 2

[0213] These energy values can be smoothed in time by, for example, the following formula:

[0214] E′(b,t,i) = (1 - α)E(b,t,i) + αE′(b,t - 1,i)

[0215] Where α is a smoothing factor (e.g., α = 0.95, however, any value can be used), and E′(b,0,i) = 0. In some embodiments, this smoothing can be omitted (i.e., e′(b,n,i) = e(b,n,i)).

[0216] The L of ILD dB (b,n) (in decibels):

[0217]

[0218] Furthermore, the ILD value 802L can be output dB (b,n).

[0219] In some embodiments, the value E′(b,t,i) can be clamped to a very small value prior to the above operations to avoid numerical instability.

[0220] In some embodiments, the transmission signal adapter 607 includes a mono factor determiner 803. The mono factor determiner 803 is configured to obtain the ILD value 802L dB (b,n) and the head orientation 400, and determine how the transmission signal should be mixed to avoid negative artifacts due to using the unprocessed transmission signal in head-tracked rendering. This determination is based on the inter-channel characteristics of the transmitted audio signal and the head orientation. In these embodiments, the inter-channel characteristics are represented by the ILD value 802 to guide or configure the mixing. In other embodiments, other inter-channel characteristics can be used.

[0221] In some embodiments, the mono factor determiner 803 is configured to determine the ILD-based mono factor by the following formula:

[0222]

[0223] Where L linit1 and L linit0 are values that control the mixing (e.g., L linit1 = 4 dB, and L limit0 = 1 dB). In some embodiments, other values can also be used. In some embodiments, the absolute value of the ILD is used. In other words, the mono factor can become larger as the negative or positive ILD becomes larger. Basically, if the absolute ILD is less than L limit0 , the ILD-based mono factor takes a value of 0; if the absolute ILD is greater than L limit1 , the ILD-based mono factor takes a value of 1; and if the ILD is between L limit0 and L linit1 , the value of the ILD-based mono factor is between 0 and 1.

[0224] In some embodiments, the mono factor determiner 803 is configured to determine a direction-based mono factor, for example, by the following formula:

[0225]

[0226] y rot (n) = 1 - |R 2,2 (n)|

[0227] where R 2,2 (n) is the entry in the second column and the second row of the rotation matrix R(n). This entry of the rotation matrix indicates the degree to which the y-axis component of the vector affects the y-axis component of the provided output vector when processed using the rotation matrix R(n). In other words, when the user orientation is aligned with the y-axis (i.e., such that the left and right ears are in line with the y-axis), its absolute value is close to 1. Thus, when the user is oriented close to perpendicular to the y-axis (e.g., when facing ±90 degrees in yaw), y rot (n) approaches 1 (and thus, R 2,2 (n) approaches 0). y limit1 and y limit0 are values that control the mixing (e.g., y limit1 = 0.8, and y limit0 = 0.4, however, other values may also be used). Thus, when y rot (n) is less than y limit0 , the direction-based mono factor takes a value of 0; if y rot (n) is greater than y limit1 , the direction-based mono factor takes a value of 1; and if y rot (n) is between y limit0 and y limit1 , the value of the direction-based mono factor is between 0 and 1. In some embodiments, y rot (n) can be calculated using an applied exponent, such as or where can be any number. Since this factor is rotation-related and the change in one coordinate is not linear with rotation, in some cases, an exponential change for this factor can provide better quality. Nevertheless, the value is also valid and provides good quality.

[0228] Furthermore, these two mono factors (the ILD-based and direction-based mono factors) are combined, and an overall mono factor Ξ(b,t,i) 804 is formulated for the left and right channels. For example, the combination can be:

[0229] Ξ(b,t,1) = Ξ ILD (b,t)Ξorient (n)f(-L dB (b,t))

[0230] Ξ(b,t,2) = Ξ ILD (b,t)Ξ orient (n)f(L dB (b,t))

[0231] Where f(x) is an operator that gives the value 1 if x is greater than zero and 0 otherwise. Using this operator enables the determination of non-zero mono factors only for channels with lower energy.

[0232] Therefore, the above Ξ ILD is determined for the sample index t (of the time-frequency audio signal), and Ξ orient is determined using the time index n, which is the time resolution of the parameterized spatial metadata. In other words, for the time index n, there can be multiple sample indices t. In this case, when formulating Ξ(b,t,i), Ξ orient can also be the same for multiple instances of t. In some embodiments, the time resolutions can be the same.

[0233] Therefore, the resulting mono factor 804 takes a large value (1 or close to 1) only when both the ILD-based mono factor and the orientation-based mono factor have large values (1 or close to 1). In other words, when configuring the mix, there is a significant level difference and head rotation (e.g., exceeding a threshold).

[0234] In these embodiments, the mono factor is non-zero only for the softer channels. As required, the mono factor 804 takes the same value regardless of whether the head is pointing forward or backward (for mirror-symmetric directions), because if the user is looking backward, the transmitted signal can be "flipped" (as discussed further below).

[0235] In some embodiments, the transmission signal adapter 607 includes a mixer 805. The mixer 805 is configured to receive the mono factor 804 Ξ(b,t,i) and the time-frequency transmitted audio signal 600S(b,t,i), and mix the time-frequency transmitted audio signal 600 based on the value of the mono factor 804. This mixing can be, for example, based on the following formula:

[0236]

[0237] Where N ch is the number of channels, typically 2.

[0238] Thus, in summary, if the ILD is large and the user's head orientation is towards the side direction, the mono factor Ξ(b,n,i) for the softer channels has a large value (1 or close to 1), and thus, the sum of the left and right transmission signals is mainly used for the softer channels (and the original transmission signal is used for the louder channels). If the user is facing away from the side direction, and / or if the absolute value of the ILD is small, the original transmission signal is mainly used for both channels. In other words, for both channels, the mono factor 804 Ξ(b,n,i) is small or zero. In some embodiments, the transmission signals may be multiplied by a certain factor (e.g., 0.5 or 0.7 or any other value) before summation to control the loudness of the sum signal; while in some other embodiments, the transmission signals are not multiplied by such a factor.

[0239] In some embodiments, since mixing can amplify or attenuate the signal with respect to the original signal (e.g., depending on the phase relationship between the channels), thus, in some embodiments, the resulting signal may be equalized to minimize the impact on the loudness of the transmission signal.

[0240] This equalization can be implemented as follows:

[0241] First, calculate the energy of the mixed signal S′ mixed (b,t,i):

[0242] E mixed (b,t,i) = |S′ mixed (b,t,i)| 2

[0243] These energies can be smoothed over time, for example, by the following formula:

[0244] E′ mixed (b,t,i) = (1 - α)E mixed (b,t,i) + αE′ mixed (b,t - 1,i)

[0245] where α is the smoothing factor (similar to that presented above), and E′ mixed (b,0,i) = 0.

[0246] Using the smoothed energy E′(b,t,i) of the original transmission signal and the smoothed energy E′ mixed (b,t,i) of the mixed transmission signal, the equalization value can be calculated, for example, by the following formula:

[0247]

[0248] where c max is the maximum allowable gain (e.g., c max= 4 or any other value), which can be used to limit the allowed equalization amount to avoid excessive amplification such as noise. Additionally, a lower bound can be placed on the denominator to avoid numerical instability.

[0249] Furthermore, finally, for example, the mixed time-frequency transmitted audio signal 806S is obtained by the following formula mixed (b,t,i):

[0250] S mixed (b,t,i) = g eq (b,t)S′ mixed (b,t,i)

[0251] In some embodiments, the transmission signal adapter 607 includes a transmission channel switch 807. The transmission channel switch 807 is configured to obtain the resulting mixed time-frequency transmission signal 806S mixed (b,t,i) and the head orientation R(n). The adapter 607 processes the case where the user is facing a direction such as ±90 degrees prior to the transmission channel switch 807, and the transmission channel switch 807 is configured to determine and process the case where the user is facing, for example, the rear direction (e.g., approximately 180 degrees of yaw). The transmission channel switch 807 is also configured to monitor the R 2,2 (n) entry in R(n). When this value is below a threshold, e.g., below -0.17 (or any other suitable value), this indicates that, for example, the user has exceeded the head orientation of 90 degrees of yaw by approximately 10 degrees. Furthermore, the transmission channel switch is configured to determine that a switch is needed. The transmission channel switch 807 is further configured to continuously monitor R 2,2 (n) until it exceeds 0.17 (or any other suitable value), which means that, for example, the yaw of the user's head orientation has returned to the front, exceeding the yaw of 90 degrees in the front direction by approximately 10 degrees. Furthermore, the transmission channel switch 807 is configured to determine that no (further) switch is needed.

[0252] When the transmission channel switch 807 has determined that a switch is needed, it switches the channel order. In other words, such a switch can be achieved by the following formula:

[0253] S adapted (b,t,i) = S mixed (b,t,3 - i)

[0254] When the channel index is i = 1, 2. Otherwise, S adapted (b,t,i) = S mixed (b,t,i). When the transmission channel switch 807 changes from the switched (mode) to the non-switched (mode) or from the non-switched (mode) to the switched (mode), it can be achieved by interpolation. For example, when moving from the non-switched mode to the switched mode, it can be formulated as:

[0255] S adapted (b,t,i) = g eq′ (b,t)(g interp (t)S mixed (b,t,3 - i)+(1 - g interp (t))S mixed (b,t,i))

[0256] where g interp (t) is an interpolation coefficient that starts from 0 and ends at 1 during the interpolation interval, where the interval can be, for example, 400 samples t. The interpolation can also have an equalizer g eq′ (b,t) that ensures the energy of S adapted (b,n,i) is the same as the sum of the energy of the signal g interp (t)S mixed (b,t,3 - i) and (1 - g interp )S mixed (b,t,i). The equalizer g eq′ (b,t) can be capped to a certain value, such as 4 (or any other suitable value). When the mode changes from "switching" to "non - switching", the interpolation can be the same, except that g interp (t) starts from 1 and decreases to 0 over 400 sample intervals.

[0257] The outputs of the transmission channel switch 807 and the transmission channel adapter 607 are the adapted time - frequency transmission signal 606S adapted (b,n,i), which can be labeled as a column vector for both channels:

[0258]

[0259] Regarding Figure 9 , according to some embodiments, an example flowchart of the operation of the transmission signal adapter 607 shown in Figure 8 is shown.

[0260] Thus, as shown by 901, the first operation can include obtaining a head - orientation signal and a time - frequency transmission audio signal.

[0261] Then, as shown by 903, the inter - channel level difference is determined from the time - frequency transmission audio signal.

[0262] After that, as shown by 905, the mono - factor is determined based on the inter - channel level difference and the head - orientation.

[0263] In addition, as shown by 907, the time - frequency transmission audio signal is mixed based on the mono - factor.

[0264] Then, as shown by 909, the method determines whether to switch channels (and switches them when determined) based on the head orientation.

[0265] Furthermore, as shown by 911, an adapted time-frequency transmitted audio signal can be output.

[0266] Regarding Figure 10 , an example of the effect of the application of the above embodiment is shown. The first row shows the spectrograms of the time-frequency transmitted signals S(b,n,i) of the left 1001 and right 1003. These signals are from an analog capture scenario where there is pink noise arriving from 36 uniformly spaced directions in the horizontal plane and a voice sound arriving directly from the left. In this example, the sound is captured using two coincident cardioid signals pointing left and right. Thus, the voice sound only appears in the left capture mode, and both signals contain noise / environment that is partially incoherent between the transmitted audio signals.

[0267] The second row shows the absolute value of the customized interaural level difference 1004|L dB (b,t)| as described previously.

[0268] The third row shows the customized monaural factors Ξ(b,t,i) for the left 1005 and right 1007 channels (assuming a 90-degree yaw head orientation) as described previously. It should be noted that when the voice signal is active and results in a larger absolute value of the ILD, the monaural factor dominates in the softer (right) channel where the voice signal does not initially reside.

[0269] The fourth row shows the spectrograms of the adapted time-frequency transmitted signals 1009 and 1011S adapted (b,n,i) that are processed as described previously. Thus, it is shown that the processing provides the voice sound to the channels of the two adapted time-frequency transmitted signals. However, as shown in the third row, in the time-frequency regions where the voice is inactive, the monaural factor Ξ(b,t,i) is low or zero, which means that the noise / environment retains most of its incoherence at the adapted time-frequency transmitted signals. This is advantageous in that spatial processing based on these signals can render the environment with zero or a minimal amount of decorrelation, which is known to be important for the sound quality of certain sound types such as applause / cheers.

[0270] The proposed embodiments can be applied to any parametric spatial audio stream or audio signal. For example, the Directional Audio Coding (DirAC) method can be applied to Ambisonics signals, and similar spatial metadata (e.g., direction and diffuseness values in a frequency band) can be obtained. The transmitted audio signal can be determined, for example, from the W and Y components of the Ambisonics signal by computing a cardioid pointing to ±90 degrees. The above methods can be applied to such spatial metadata and transmitted audio signals.

[0271] The proposed method has been described as applicable to head-tracking binaural rendering. This is generally understood to enable tracking of the movement of the listener's head (for which the rendered binaural output is created). These movements typically include at least rotation, but may also include translation. While this is the main use case of the proposed method, however, it is not limited to the use case of listener head tracking. If such binaural rendering is implemented in any situation where the rendering is rotated, the same proposed method can be applied to other embodiments. For example, in immersive codecs such as IVAS or MPEG-I, viewport / viewpoint adjustments can be made that are independent of the listener's head orientation. Similarly, this can be provided to the binaural renderer and can affect the rendering in the same way as head tracking. Generally, directional parameters can be provided to the binaural renderer, and the renderer is configured to implement the same steps proposed in the proposed method, regardless of the source of the directional parameters.

[0272] The covariance matrix-based rendering scheme discussed above is only an example, and other configurations are also possible. For example, an audio signal can be divided into a directional part and a non-directional part in a frequency band based on a scaling parameter; further, amplitude translation can be used to position the directional part to a virtual loudspeaker; the non-directional part can be assigned to all loudspeakers and decorrelated, and then the processed directional part and non-directional part are added together, and finally each virtual loudspeaker is processed with HRTF to obtain a binaural output. This process is described in more detail in the DirAC rendering scheme as described in Laitinen, M.V. and Pulkki, V. ("Binaural reproduction for directional audio coding", in Proceedings of the 2009 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (pp. 337-340), October 2009). In this case, the use of a transmission signal adapter can be beneficial because signals for virtual loudspeakers can be generated such that the left virtual loudspeaker is synthesized based on the left channel of the adapted time-frequency transmission signal, and the same applies to the right virtual loudspeaker.

[0273] The exemplary embodiments presented above include encoding and decoding steps. However, in some embodiments, the processing can also be applied in systems that do not involve encoding and decoding. For example, regarding Figure 11 , another exemplary embodiment is shown. The input microphone array audio signal 1100 is forwarded to the microphone array front end 1101, which is implemented in a similar manner as discussed regarding Figure 1 . However, the resulting transmitted audio signal 1102 and the spatial metadata 1104 are directly forwarded to the spatial synthesizer 1103 together with the head orientation 1106 information. The spatial synthesizer 1103 is configured to operate in the same manner as the above-described spatial synthesizer. Thus, the proposed method can also be used, for example, to directly (i.e., without encoding / decoding) render the sound captured by the microphone array. It should be noted that in this case (and possibly also in some other embodiments), the transmitted audio signals 1102 are not necessarily transmitted anywhere; they are simply audio signals that are suitable for and used for rendering.

[0274] Furthermore, the exemplary embodiments presented above use the microphone array signal as an input for creating a parametric spatial audio stream (i.e., the transmitted audio signal and the spatial metadata). However, in some other embodiments, other types of inputs can be used to create a parametric spatial audio stream. If the audio signal and the parametric spatial metadata (together with the head orientation or similar information) are input to the spatial synthesizer, then the source of the transmitted audio signal and the spatial metadata is not important for using the above embodiments.

[0275] For example, a parametric spatial audio stream can be created from multi-channel audio signals (such as 5.1 or 7.1+4 multi-channel signals) as well as audio objects. For example, WO2019086757A1 discloses methods for determining a parametric spatial audio stream from these input formats. As another example, the DirAC method can be used to create a parametric spatial audio stream from an Ambisonic signal.

[0276] Thus, a parametric spatial audio stream (i.e., the transmitted audio signal and the spatial metadata) can originate from any source and the methods proposed herein can be used.

[0277] The example embodiments presented above use head orientation as input. However, in some alternative embodiments, head orientation and head position may be used. In other words, the head can be tracked in 6 degrees of freedom (6DoF). An example parametric 6DoF rendering system is proposed in GB2007710.8, which operates using, for example, Ambisonic signals. Similarly, like the example embodiments presented above, 6DoF rendering requires the creation of prototype signals (or similar signals used in rendering). Therefore, the method proposed above can also be applied in 6DoF rendering and is applicable to cases where audio signals are transmitted using stereo.

[0278] As presented above, the proposed method can be used with the IVAS codec. Additionally, they can also be used with any other suitable codec or system. For example, they can be used with the MPEG-I codec. As another example, the present invention can be used in the Nokia OZO audio system, for example, for rendering binaural audio captured using a microphone array (attached in, for example, a mobile device).

[0279] The example embodiments presented above perform transmission signal adapter processing in frequency bins. In some alternative embodiments, this processing can be performed in frequency bands, for example, to optimize the computational complexity of the processing.

[0280] In the example above, crossmixing is only performed for the softer channels (in the frequency band or bin) in the mixer. In some embodiments, crossmixing can be performed for both channels. For example, mixing can be performed by determining the same mono factor for both channels:

[0281] Ξ(b,t,1) = Ξ(b,t,2) = Ξ ILD (b,t)Ξ orient (n)

[0282] In summary, if the value of the mono factor Ξ(b,n,1) = Ξ(b,n,2) is large, the sum of the left and right transmission signals is mainly used for both channels; and if the value is small, the original transmission signals are typically used.

[0283] The example embodiments presented above use a dedicated processing block to perform the adaptation of the transmission signal, resulting in a modified audio signal, which is then fed to a subsequent processing block. In some alternative embodiments, the adaptation of the transmission signal can be performed as part of the processing. Additionally, in some cases, the rendering of any intermediate signal is optional, but the mixing information can be used to affect the processing values. For example, the prototype matrix used in rendering can be modified. More specifically, in the foregoing formula, it is stipulated that the prototype matrix can be, for example, or However, in some alternative embodiments, this matrix is adaptive based on head orientation and inter-channel information. For example, a prototype matrix (denoted as Q(b,t)) can be determined as:

[0284]

[0285] In this example, in addition to the block transfer channel switch, a transmission signal adapter is not implemented either. Alternatively, in some embodiments, the transmission channel switching is included in the matrix Q(b,t) such that when the operation mode is to switch channels, then:

[0286]

[0287] In some embodiments, when decorrelated sound is needed, it is generated based on the signal Q(b,t)x(b,t).

[0288] The above example uses the inter-channel level difference (ILD) as the inter-channel information and determines the mixing information for transmitting the audio signal based on this information and the head orientation. However, in some embodiments, as a supplement or alternative to ILD, the inter-channel information can use the inter-channel correlation (IC) and the inter-channel phase difference (IPD). For example, if the IC value is very high (close to 1), then even if the ILD value will be greater than 1 dB as shown in the example, it is assumed that both channels have all the relevant signal content. Therefore, in this case, the thresholds L limit1 and L limit0 can be adapted to higher values in this case, for example, twice the values shown in the above embodiments. In another example using other values in addition to the ILD value, if the IC value is high and the IPD value is not zero, this means that the two transmitted audio signals contain delayed or otherwise out-of-phase signals. Therefore, when cross-mixing the signals between channels, the phase can be matched based on the IPD value during this mixing process to avoid frequency-dependent effects where some frequencies are amplified or attenuated more due to the phase difference.

[0289] In some alternative embodiments, the equalization gain g eq (b,t) can be restricted in an alternative or additional way rather than just restricting it to a certain value. For example, the average equalization factor on the frequency bin b can be calculated and the value g eq (b,t) can be restricted such that they can be no greater than c limit times the average value (e.g., c limit is 1, 1.125, or 2, or any suitable value). This or any other suitable restriction of the equalization value can be used to prevent the signal from being over-enhanced (so as to avoid generating audible noise).

[0290] In general, various embodiments of the present invention can be implemented using hardware or dedicated circuits, software, logic, or any combination thereof. For example, some aspects can be implemented using hardware, while other aspects can be implemented using firmware or software executable by a controller, microprocessor, or other computing device, but the present invention is not limited thereto. Although the various aspects of the present invention can be illustrated and described as block diagrams, flowcharts, or using some other graphical representation, it is well known that the blocks, devices, systems, techniques, or methods described herein can be implemented as non-limiting examples using hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers, or other computing devices, or some combination thereof.

[0291] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. In addition, in this regard, it should be noted that any block of the logical flow in the accompanying drawings can represent a program step, or interconnected logical circuits, blocks, and functions, or a combination of program steps and logical circuits, blocks, and functions. The software can be stored on a physical medium such as a memory chip or a memory block implemented within a processor, on a magnetic medium such as a hard disk or a floppy disk, and on an optical medium such as a DVD and its data variants, a CD.

[0292] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and, by way of non-limiting example, can include one or more of a general-purpose computer, a dedicated computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a gate-level circuit based on a multi-core processor architecture, and a processor.

[0293] Embodiments of the present invention can be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools can be used to convert a logic-level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0294] Programs, such as those provided by Synopsys, Inc., in Mountain View, California and Cadence Design in San Jose, California, use well-established design rules and a library of pre-stored design modules to automatically route conductors and place components on a semiconductor chip. Once the design of a semiconductor circuit is complete, the resulting design in a standardized electronic format (e.g., Opus, GDSII, etc.) can be transferred to a semiconductor manufacturing facility or "fab" for fabrication.

[0295] As used in this application, the term "circuit" may refer to one or more or all of the following:

[0296] (a) Only hardware circuit implementations (such as only analog and / or digital circuit implementations); and

[0297] (b) Combinations of hardware circuits and software, such as (if applicable):

[0298] (i) Combinations of analog and / or digital hardware circuits and software / firmware; and

[0299] (ii) Any part of a hardware processor with software (including a digital signal processor, software, and memory, which work together to enable a device such as a mobile phone or server to perform various functions); and

[0300] Hardware circuits and / or processors, such as a microprocessor or part of a microprocessor, which require software (e.g., firmware) to operate, but may not have software when not required to operate.

[0301] The above definition of "circuit" applies to all uses of the term in this application, including its use in any claims. As another example, as used in this application, the term "circuit" also encompasses implementations of only hardware circuits or processors (or multiple processors) or a part of a hardware circuit or processor and its accompanying software and / or firmware. The term "circuit" also encompasses (e.g., and if applicable to a particular claimed element) a baseband integrated circuit or a processor integrated circuit for a mobile device, or a similar integrated circuit in a server, a cellular network device, or other computing or network device.

[0302] As used herein, the term "non-transitory" is a limitation on the medium itself (i.e., tangible, rather than a signal), rather than a limitation on data storage persistence (e.g., RAM vs. ROM).

[0303] As used herein, "at least one of the following: <list of two or more components / elements>" and "at least one of <list of two or more components / elements>" and similar phrases (where the list of two or more components / elements is joined by "and" or "or") mean at least any one component / element, or at least any two or more components / elements, or at least all of the components / elements.

[0304] The foregoing description has provided a complete and useful description of exemplary embodiments of the invention by way of example and not of limitation. However, various modifications and adaptations will become apparent to those skilled in the relevant art in view of the foregoing description when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of the invention will still fall within the scope of the invention as defined by the appended claims.

Claims

1. A method for generating a spatial output audio signal, the method comprising: obtaining a spatial audio signal, the spatial audio signal including: at least two-channel audio signals; and at least one spatial parameter associated with the at least two-channel audio signals; analyzing the at least two-channel audio signals to determine at least one inter-channel characteristic; obtaining orientation and / or position parameters; determining mixing information based on the at least one inter-channel characteristic and the orientation and / or position parameters; and generating at least two-channel output audio signals based on the at least two-channel audio signals, the orientation and / or position parameters, and the mixing information.

2. The method according to claim 1, wherein Generating at least two-channel output audio signals further includes: generating the at least two-channel output audio signals based on the at least one spatial parameter associated with the at least two-channel audio signals.

3. The method according to any one of claims 1 or 2, wherein Determining the mixing information further includes: further determining the mixing information based on the at least one spatial parameter.

4. The method according to any one of claims 1 or 2, wherein Analyzing the at least two-channel audio signals to determine the at least one inter-channel characteristic includes: generating the inter-channel characteristic based on the at least one spatial parameter associated with the at least two-channel audio signals.

5. The method according to any one of claims 1 to 4, wherein, The at least one spatial parameter associated with the at least two-channel audio signals includes: spatial parameters associated with respective audio channel audio signals of the at least two audio channel audio signals; and spatial parameters associated with the at least two audio channel audio signals.

6. The method according to any one of claims 2 to 5 or any one of the claims dependent on claim 2, wherein, Generating at least two-channel output audio signals based on the at least two-channel audio signals, the at least one spatial parameter associated with the at least two-channel audio signals, the orientation parameter, and the mixing information includes: generating at least one prototype matrix based on the mixing information; rendering the at least two-channel output audio signals from the at least two-channel audio signals based on: the at least one spatial parameter associated with the at least two-channel audio signals, the orientation parameter, the orientation parameter, and the at least one prototype matrix.

7. The method according to any one of claims 2 to 6 or any one of the claims dependent on claim 2, wherein, Generating at least two-channel output audio signals based on the at least two-channel audio signals, the at least one spatial parameter associated with the at least two-channel audio signals, the orientation parameter, and the mixing information includes: processing the at least two-channel audio signals based on the mixing information to generate at least two-channel adapted audio signals; rendering the at least two-channel output audio signals from the at least two-channel adapted audio signals based on: the at least one spatial parameter associated with the at least two-channel audio signals; and the orientation parameter.

8. The method according to claim 7, wherein Processing the at least two-channel audio signals based on the mixing information to generate at least two-channel adapted audio signals includes: adapting the at least two-channel audio signals based on a current orientation and the inter-channel characteristic.

9. The method according to claim 8, wherein Adapting the at least two channel audio signals based on the current orientation and the inter-channel characteristics includes: determining a mono factor based on the current orientation and the inter-channel characteristics, the mono factor being configured to indicate how the at least two channel audio signals should be mixed to avoid negative artifacts in the at least two channel output audio signals.

10. The method according to any one of claims 1 to 9, wherein, Analyzing the at least two channel audio signals to determine at least one inter-channel characteristic includes: analyzing the at least two channel audio signals to determine at least one of the following: An inter-channel level difference between the at least two channel audio signals; The modified inter-channel level difference between the at least two channel audio signals, the modification being based on the orientation and / or position parameter; An inter-channel phase difference between the at least two channel audio signals; An inter-channel time difference between the at least two channel audio signals; An inter-channel similarity measure between the at least two channel audio signals; and An inter-channel correlation between the at least two channel audio signals.

11. The method according to any one of claims 7 to 10, wherein, Processing the at least two channel audio signals based on the mixing information to generate at least two channel adapted audio signals includes: mixing the at least two channel audio signals based on the inter-channel differences such that an audio component in substantially one of the at least two channel audio signals is mixed into a corresponding one of the at least two channel adapted audio signals and further at least partially cross-mixed into another one of the at least two channel adapted audio signals.

12. The method according to claim 11, wherein, Processing the at least two channel audio signals based on the mixing information to generate at least two channel adapted audio signals further includes: switching at least two of the at least two channel adapted audio signals based on the orientation and / or position parameter indicating an orientation towards a rear direction.

13. The method according to any one of claims 1 to 12, wherein The at least two channel output audio signals are binaural audio signals.

14. The method according to any one of claims 1 to 13, further comprising: Obtaining a user head orientation and / or position, and wherein obtaining the orientation and / or position parameter includes: processing the user head orientation and / or position to generate the orientation and / or position parameter.

15. An apparatus comprising components for performing the method according to any one of claims 1 to 14.

16. A computer program comprising instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 14.

17. A device for generating a spatial output audio signal, the device comprising components configured to perform the following operations: Obtain a spatial audio signal, the spatial audio signal including: At least two channel audio signals; And at least one spatial parameter associated with the at least two channel audio signals; Analyzing the at least two channel audio signals to determine at least one inter-channel characteristic; Obtaining an orientation and / or position parameter; Determining mixing information based on the at least one inter-channel characteristic and the orientation and / or position parameter; And Generate at least two-channel output audio signals based on the at least two-channel audio signals, the orientation and / or position parameters, and the mixing information.

18. An apparatus for generating spatial output audio signals, the apparatus comprising at least one processor and at least one memory storing instructions which, when executed by the at least one processor, cause the system to at least: Obtain a spatial audio signal, the spatial audio signal including: At least two-channel audio signals; And at least one spatial parameter associated with the at least two-channel audio signals; Analyze the at least two-channel audio signals to determine at least one inter-channel characteristic; Obtain orientation and / or position parameters; Determine mixing information based on the at least one inter-channel characteristic and the orientation and / or position parameters; And Generate at least two-channel output audio signals based on the at least two-channel audio signals, the orientation and / or position parameters, and the mixing information.

19. The apparatus according to claim 18, wherein, Causing the apparatus to generate the at least two-channel output audio signals causes the apparatus to: Generate the at least two-channel output audio signals based on the at least one spatial parameter associated with the at least two-channel audio signals.

20. The apparatus according to claim 18 or 19, wherein Causing the apparatus to determine the mixing information causes the apparatus to: Further determine the mixing information based on the at least one spatial parameter.

21. The device according to claim 18 or 19, wherein Causing the apparatus to analyze the at least two-channel audio signals to determine the at least one inter-channel characteristic causes the apparatus to: Generate the inter-channel characteristic based on the at least one spatial parameter associated with the at least two-channel audio signals.

22. The device according to any one of claims 18 to 21, wherein, The at least one spatial parameter associated with the at least two-channel audio signals includes: Spatial parameters associated with respective audio channel audio signals among the at least two audio channel audio signals; and Spatial parameters associated with the at least two audio channel audio signals.

23. The device according to any one of claims 18 to 22, wherein, The at least two-channel output audio signals are binaural audio signals.

24. The apparatus according to any one of claims 18 to 23, further causing the apparatus to: obtain a user's head orientation and / or position, and wherein, Obtaining the orientation and / or position parameters causes the apparatus to: Process the user's head orientation and / or position to generate the orientation and / or position parameters.

Citation Information

Patent Citations

  • Analysis of spatial metadata from multi-microphones having asymmetric geometry in devices

    GB201619573D0

  • Hydraulic drive system

    GB202007710D0

  • Spatial audio representation and rendering

    GB202007904D0

  • Spatial audio parameters and associated spatial audio playback

    GB2572650A

  • Determination of spatial audio parameter encoding and associated decoding

    GB2575305A